Skip to content

Operator stutter: what flips the participant structure signature, and which rebuild step costs 197ms #509

Description

@iamfatness

Follow-up to the 2026-09-12 stutter work (partly fixed; this is the residue).

Fixed already: the 9.0 ms/apply diagnostic-string tax (audio summaries now on
a 1 s throttle) and the spike RATE (the structural rebuild is coalesced, 150 ms
leading+trailing). See CLAUDE.md "The operator stutter is the SNAPSHOT APPLY".

Two things remain unknown, and one log line now answers both.

1. Which step costs 197 ms

ApplyLiveParticipants's structural rebuild ran at up to 202 ms on the
owner's machine, 93 times in twelve minutes. On this dev box the same rebuild is
~4.7 ms (largest step multiviewGrid 1.3 ms), so it could not be profiled
here. The rebuild now times each of its ten steps and logs the breakdown whenever
it exceeds ONE FRAME:

perf: structural participant rebuild <total>ms participants=N ::
  roomLists=… gallery=… audioRows=… multiviewTiles=… participantList=…
  showInputEditors=… multiviewGrid=… previewRouting=… productionReadouts=…

Not gated on verbose diagnostics — the original incident was only diagnosable
because verbose happened to be on.

2. What actually flips the signature

The obvious answer is wrong and must not be acted on. The signature buckets
each participant as video-on/off, so "camera flicker during resubscribe churn"
reads as the explanation — the comment above it says as much. Measured instead:
207 engine participants payloads across that exact 14:00-14:12 window carry
ZERO video-on/off transitions, ZERO screen-share transitions and ZERO roster
id-set changes.
Something else flips it 93 times.

The other half of the signature is the CAPTURE DEVICE set
(CaptureDevices.Select(d => d.Id)), which is the obvious next place to look —
ApplyDiscoveredCaptureDevices clears and refills that collection, and the
screen-discovery retry at the top of ApplySnapshot can call it, though that is
bounded to 10 attempts and cannot explain 93.

Hypotheses already killed by experiment (each its own build, test meeting)

Hypothesis Result
Verbose diagnostics (slow applies begin the minute it was enabled) forced on — no effect, 1.4 ms median
Engine off vs on @full path reproduced, 1.3 ms
Sources assigned (9 sources / 18 subscriptions) 1.5 ms
Tab realization (Audio tab via UIA) audioReadouts 0.1 -> 0.4 ms — right mechanism, 22x short of 9.0 ms

Acceptance

  • A structural participant rebuild line from a real stuttering session naming
    the >16.7 ms step.
  • The signature trigger identified from evidence, not from the comment.
  • Then fix THAT step (diff-update is the likely shape — spec P4 — but it is a
    guess until the line names it, and it touches the 0xc000027b crash class).

Activity

  1. iamfatness commented on Sep 13, 2026

    @iamfatness
    OwnerAuthor

    Reproduced and measured on the owner's machine (live session 2026-09-13, 13–16 participants)

    The perf: structural participant rebuild instrument fired 5× this session, all over the 16.7 ms frame budget, clustered when participants LEAVE:

    10:03  rebuild 17.3ms participants=16
    11:38  rebuild 18.7ms participants=16
    11:39  rebuild 21.1ms participants=15
    11:39  rebuild 20.3ms participants=14
    11:39  rebuild 19.7ms participants=13
    

    Full step breakdown of the worst (21.1 ms):

    roomLists=0.1  gallery=4.3  audioRows=4.0  multiviewTiles=0.3  participantList=0.2
    showInputEditors=4.4  multiviewGrid=0.6  previewRouting=0.3  productionReadouts=6.8
    

    What this settles

    It is not one giant 197 ms step (that was a different past occurrence, likely a GC or contention spike). On this machine it is the sum of four moderate UI projections — productionReadouts 6.8 + showInputEditors 4.4 + gallery 4.3 + audioRows 4.0 ≈ 20 ms — each rebuilt wholesale on every roster change. And the trigger this time is confirmed: participants leaving (15→14→13 in the 11:39 cluster) flips the structure signature.

    Direction

    Make those four projections incremental / diff-based instead of full rebuilds on the structural path (same treatment already applied elsewhere for the 0xc000027b collection-rebuild class), so a roster change doesn't cost ~20 ms on the UI thread. The 197 ms outlier remains worth catching if it recurs — the instrument now names the step when it does.

    Found via the live-session metrics sweep 2026-09-13.

  2. iamfatness commented on Sep 24, 2026

    @iamfatness
    OwnerAuthor

    Architecture review at 6f4f025 (2026-09-24): this stays the 197 ms participant-rebuild stutter. Silent drop of a fire-and-forget command on the single sync slot is #622.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    backlogRanked in docs/BACKLOG.md

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions