Follow-up to the 2026-09-12 stutter work (partly fixed; this is the residue).
Fixed already: the 9.0 ms/apply diagnostic-string tax (audio summaries now on
a 1 s throttle) and the spike RATE (the structural rebuild is coalesced, 150 ms
leading+trailing). See CLAUDE.md "The operator stutter is the SNAPSHOT APPLY".
Two things remain unknown, and one log line now answers both.
1. Which step costs 197 ms
ApplyLiveParticipants's structural rebuild ran at up to 202 ms on the
owner's machine, 93 times in twelve minutes. On this dev box the same rebuild is
~4.7 ms (largest step multiviewGrid 1.3 ms), so it could not be profiled
here. The rebuild now times each of its ten steps and logs the breakdown whenever
it exceeds ONE FRAME:
perf: structural participant rebuild <total>ms participants=N ::
roomLists=… gallery=… audioRows=… multiviewTiles=… participantList=…
showInputEditors=… multiviewGrid=… previewRouting=… productionReadouts=…
Not gated on verbose diagnostics — the original incident was only diagnosable
because verbose happened to be on.
2. What actually flips the signature
The obvious answer is wrong and must not be acted on. The signature buckets
each participant as video-on/off, so "camera flicker during resubscribe churn"
reads as the explanation — the comment above it says as much. Measured instead:
207 engine participants payloads across that exact 14:00-14:12 window carry
ZERO video-on/off transitions, ZERO screen-share transitions and ZERO roster
id-set changes. Something else flips it 93 times.
The other half of the signature is the CAPTURE DEVICE set
(CaptureDevices.Select(d => d.Id)), which is the obvious next place to look —
ApplyDiscoveredCaptureDevices clears and refills that collection, and the
screen-discovery retry at the top of ApplySnapshot can call it, though that is
bounded to 10 attempts and cannot explain 93.
Hypotheses already killed by experiment (each its own build, test meeting)
| Hypothesis |
Result |
| Verbose diagnostics (slow applies begin the minute it was enabled) |
forced on — no effect, 1.4 ms median |
| Engine off vs on |
@full path reproduced, 1.3 ms |
| Sources assigned (9 sources / 18 subscriptions) |
1.5 ms |
| Tab realization (Audio tab via UIA) |
audioReadouts 0.1 -> 0.4 ms — right mechanism, 22x short of 9.0 ms |
Acceptance
- A
structural participant rebuild line from a real stuttering session naming
the >16.7 ms step.
- The signature trigger identified from evidence, not from the comment.
- Then fix THAT step (diff-update is the likely shape — spec P4 — but it is a
guess until the line names it, and it touches the 0xc000027b crash class).
Follow-up to the 2026-09-12 stutter work (partly fixed; this is the residue).
Fixed already: the 9.0 ms/apply diagnostic-string tax (audio summaries now on
a 1 s throttle) and the spike RATE (the structural rebuild is coalesced, 150 ms
leading+trailing). See CLAUDE.md "The operator stutter is the SNAPSHOT APPLY".
Two things remain unknown, and one log line now answers both.
1. Which step costs 197 ms
ApplyLiveParticipants's structural rebuild ran at up to 202 ms on theowner's machine, 93 times in twelve minutes. On this dev box the same rebuild is
~4.7 ms (largest step
multiviewGrid1.3 ms), so it could not be profiledhere. The rebuild now times each of its ten steps and logs the breakdown whenever
it exceeds ONE FRAME:
Not gated on verbose diagnostics — the original incident was only diagnosable
because verbose happened to be on.
2. What actually flips the signature
The obvious answer is wrong and must not be acted on. The signature buckets
each participant as video-on/off, so "camera flicker during resubscribe churn"
reads as the explanation — the comment above it says as much. Measured instead:
207 engine
participantspayloads across that exact 14:00-14:12 window carryZERO video-on/off transitions, ZERO screen-share transitions and ZERO roster
id-set changes. Something else flips it 93 times.
The other half of the signature is the CAPTURE DEVICE set
(
CaptureDevices.Select(d => d.Id)), which is the obvious next place to look —ApplyDiscoveredCaptureDevicesclears and refills that collection, and thescreen-discovery retry at the top of
ApplySnapshotcan call it, though that isbounded to 10 attempts and cannot explain 93.
Hypotheses already killed by experiment (each its own build, test meeting)
@fullpath reproduced, 1.3 msaudioReadouts0.1 -> 0.4 ms — right mechanism, 22x short of 9.0 msAcceptance
structural participant rebuildline from a real stuttering session namingthe >16.7 ms step.
guess until the line names it, and it touches the 0xc000027b crash class).