Repository navigation
Conversation
denodell
force-pushed
the
claude/soak-leak-diagnosis-a73cxf
branch
3 times, most recently
from
September 29, 2026 21:50
4f060ce to
da9be53
Compare
A failing soak test used to say that something leaked, but not what.
Now each run takes a heap snapshot at the baseline pass and, when it
fails, another after the last reading. It compares the two and reports
what leaked and what still references it, in a sentence and a chain:
A listener on window is still registered. Its callback `onResize`
still references the <section class="report-drawer"> element.
window → EventListener → onResize() → <section class="report-drawer">
- src/heap-snapshot.ts reads V8's heap snapshot format and finds the
shortest path from any object back to the root.
- src/heap-analysis.ts compares the two snapshots, and trims each chain
down to your own code: it skips V8's own roots, Blink's listener
plumbing, the virtual clock, code-split module objects and plain
objects, and folds elements that share a holder into one leak.
- src/heap-capture.ts takes the snapshots over CDP, streams them to
disk, and keeps within time, memory and disk limits. Running out
leaves a note on the result, never a failed test, and the work stops
5 seconds before the test's own timeout.
- src/diagnose.ts writes the sentences, and adds formatSoakReport().
It names the function nearest the leaked object, says when that
function is still in a set or array, trims long element markup to
the tag, id, role and first class, and shortens a long chain from the
root end. All of that came from running it against Floating UI with
one of its unsubscribes removed. When only Chrome's own objects keep
the elements, the report says your code isn't keeping them, and names
the two known causes: the undo history Chrome keeps for typing, found
on Quill, and console.log of an event while Playwright is connected,
found on Ionic. Chrome's own objects are marked kind: 'browser'.
- src/stats.ts finds the halfway point of a run by pass, not by
reading, so a long run that climbed and then stopped reads as
settled instead of noisy.
- src/cdp.ts has Chrome update the page's layout before each reading.
Until then Chrome keeps references to elements a flow just removed,
so the odd reading could count a whole component that was gone.
New options: diagnose ('on-failure', 'always' or 'off'), keepSnapshots
and diagnoseTimeout. New types: SoakDiagnosis, SoakDiagnoseMode,
SoakDetachedClass, SoakRetainerHop and SoakGrowth.
The demo app gains a timer leak, a leak that stays out of the DOM, and a
React 19 panel whose registry keeps its sections after they unmount.
Unit tests run the parser and the comparison against hand-built
snapshots; soak tests check the exact chain for each real leak.
README and CHANGELOG cover the release as 0.2.0. The README intro now
explains the virtual clock and network, a new section compares the
library with fuite and memlab, and the clock section says how to move
the clock inside a flow that waits on a timer.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01AWXrLaizb9XYxvkcxxrAzV
denodell
force-pushed
the
claude/soak-leak-diagnosis-a73cxf
branch
from
September 30, 2026 00:06
da9be53 to
2bad96c
Compare
`npm run test:coverage` runs the whole suite under c8 and fails if lines, statements, functions or branches drop under 99%. CI runs it on the Node 22 job. `src/types.ts` is left out, since it holds only types and compiles to nothing. Coverage was 95% of lines and 87% of branches. It is now 100% of lines, statements and functions, and 99.4% of branches. The new unit tests cover the reporter, the report's wording, the snapshot parser and retainer chains on small built graphs, and the CDP, clock, snapshot capture and runner code against fake pages and sessions, so the failures a real browser rarely produces can be made on purpose. `formatSignedBytes` had no callers, so it's gone. Three fallbacks in diagnose.ts that their callers made impossible are now stated as invariants. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AWXrLaizb9XYxvkcxxrAzV
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
When a soak test failed, it told you something leaked but not what. This PR takes a heap snapshot at the baseline pass and another after the last reading, compares the two, and reports what leaked and what is still referencing it. A failing run now prints something like this under the box:
The baseline snapshot has to be taken before anyone knows whether the run will fail, so every run takes one and a passing run deletes it. On a real app that adds a few seconds to each run, and
diagnose: 'off'turns it off.Files
src/heap-snapshot.tsreads V8's heap snapshot format. The file stores nodes and edges as flat arrays of integers that point into a table of strings. This decodes them, builds a reverse index so it can look up what references any node, and finds the shortest path from a node back to the root.src/heap-analysis.tscompares the two snapshots. It counts detached DOM nodes by tag, and JS objects and closures by name, then traces a path back to your code for each kind that went up. It doesn't touch the disk or the browser, so the unit tests run it against small hand-built snapshots.src/heap-capture.tsdoes the parts that do touch them: the Chrome DevTools Protocol calls, writing each snapshot to disk as it streams in, the time limit, and deleting the files afterwards.src/diagnose.tsturns the comparison into sentences like the one above. When the snapshots name a cause, the report no longer adds its own guess. It also addsformatSoakReport(result), which returns the report as text, so asoak.measure()result can be logged the same way aSoakLeakErrormessage reads.src/soak.tswires all of this into the run, andsrc/types.tsadds theSoakDiagnosis,SoakDetachedClass,SoakRetainerHopandSoakGrowthtypes to the result.src/stats.tsfinds the halfway point of a run by pass instead of by reading, so a long run that climbed and then stopped is called settled instead of noisy.src/cdp.tshas Chrome update the page's layout before each reading. See "Readings" below.README.mdandCHANGELOG.mdcover the new options and output. The README's older report examples are also brought up to date with what the tool prints. The intro now explains the virtual clock and network, a new "Other leak finders" section says how this differs from fuite and memlab, the clock section says how to move the clock inside a flow that waits on a timer, and the limitations cover typing into an editable area and logging events withconsole.log.Chains that point at your code
The shortest path from a leaked object back to the root is often useless, so most of the work went into steering the chain toward something you can fix.
(Global handles)or the list DevTools keeps of detached nodes. That's only two steps, so it beats any path through your code, and it tells you nothing you can act on. The walk looks for a route through the page first, and only falls back to those when there isn't one.InternalNodeandV8EventListenerobjects between an event listener and its callback. They're left out of the chain, and the variable name they carried moves up to the closure, so the closure still shows the name of the variable it captured.installSoakClockstores pending timers inside the fake clock Playwright adds to the page, so a leaked timer would otherwise be blamed on Playwright. Those steps become a single step calleda pending timer, marked withkind: 'timer'. A step that's a function is markedkind: 'closure', with just the function's name as itsnode, and a step that's one of Chrome's own objects, likeStyleEngine, is markedkind: 'browser'.ModuleandGeneratorobjects between the file that imports it and whatever it references. These are left out too, so a lazy-loaded panel reads the same as one bundled into the main file.storethenitemsreadsstore.items, and the demo's drawer readswindow.__drawer.open, which is whereopenDrawerlives.<li>and<section>elements in one list, are reported as one leak rather than a paragraph each.windowis named as one: "An array atwindow.cachekeeps growing".id,roleand first class, so a long list of utility classes or an inline style doesn't swamp the sentence. A chain longer than eight steps loses its middle near the root, and keeps the steps nearest the leak.console.logwhile Playwright is connected.Bundlers merge your modules into one scope, and V8 labels that scope with whichever function it picks. That means the chain can name a function from a different file than the one with the leak. The sentence names the variable rather than the function for that reason, and the README explains it.
Tried on real libraries
Floating UI. Its React test app, the one its own Playwright tests drive, ran clean through a menu, a popover and a tooltip, 200 passes each, with no growth in nodes or listeners. With one of its unsubscribes removed (
events.off('openchange', …)inFloatingFocusManager), the popover run failed at 19 nodes and 2 listeners a pass, and the report said:The first version of the report named the emitter's
mapinstead ofonOpenChange, printed every attribute of the dialog, and hidSet → onOpenChange()behind the ellipsis. That run is where the function, element name and ellipsis points under "Chains that point at your code" came from, and the chain is now a unit test.Quill. On its e2e page, creating and removing an editor with a toolbar ran clean over 200 passes, and so did opening and cancelling the link tooltip. Typing a line, making it bold and deleting it grew by 3 nodes a pass. The chain behind each of those nodes held nothing but Chrome's own objects, and a plain
contenteditablewith no Quill grew the same way, then stopped after about 330 passes. That's Chrome's undo history, which keeps up to 1,000 steps. The long plain run was also called noisy when it had climbed and then stopped. That came from the halfway point being found by reading instead of by pass, which dates back to 0.1.0 and is fixed here.Ionic. Its component test pages leaked every overlay they opened: alert, action sheet, loading and toast, 13 to 26 nodes a pass. The same action sheet opened from the select page was clean. The difference was one line on the leaking pages,
console.log('didDismiss', event). Adding that line to the clean page made it leak, and loggingevent.typeinstead didn't. With Playwright connected, Chrome builds a preview of anything logged, and itsStyleEnginethen keeps the overlay the event points to. For 24 of the 25 leaked alerts in one run,StyleEnginewas the only thing holding them. The report now says:Before, it said "Something still references" the alert, which pointed at your code.
Memory and time limits
A real app's snapshot can be hundreds of megabytes, and
JSON.parseneeds about four times the file size in memory to read one.--max-old-space-sizeraises the limit too, up to about 512MB, the most Node can read into one string. Past that limit the snapshots aren't read, and the result gets a note saying which limit it hit.diagnoseTimeout, which is 60 seconds by default. It also stops 5 seconds before the test's own timeout, so a slow diagnosis doesn't turn a leak report into a timeout. Either way, the result gets a note, and the test still passes or fails on its counts as it did before.keepSnapshotsis on. With it on, they're attached to the test result so you can open them in DevTools → Memory.Readings
Each reading now has Chrome update the page's layout before it collects garbage. Until Chrome does that, it keeps some references to elements that were just removed, so a reading taken the moment a flow ended could count a component that was already gone. On the React demo's fixed build, with passes back to back and one collection per reading, 16 of 180 readings counted the closed panel before this change and 0 of 180 after. It was what made the React test fail now and then on CI. Raising
gcPassesonly made it rarer, so the React test is back on the default and the README no longer suggests raising it for frameworks.Options
diagnose'on-failure''always'reports on passing runs too, and'off'skips the snapshotskeepSnapshotsfalsediagnoseTimeout60000Tests
The unit tests check the parser and the comparison against small hand-built snapshots, including the
(Global handles)shortcut and the fake clock.tests/fixtures/README.mddescribes each snapshot. The soak tests run real leaks in the browser: a leaked listener, a leaked timer, a leak that never touches the DOM, and a React 19 app whose component registry keeps references to sections after they unmount. Each one checks the exact chain that comes back. Other soak tests cover both time limits, a snapshot folder that can't be written, and a flow that throws partway through. The Floating UI, Quill and Ionic chains are unit tests too.Upgrading
A test suite on 0.1.0 gets the diagnosis on failing runs without any changes.
diagnose: 'off'goes back to the 0.1.0 behavior. One line of report output changes spelling:climbed early, then levelled offis nowclimbed early, then leveled off. Readings can come out slightly lower than before, where a pass used to count elements a flow had just removed.https://claude.ai/code/session_01AWXrLaizb9XYxvkcxxrAzV