Repository navigation
Complete bounded per-frame delivery evidence and latency attribution (#517) #823
Description
Activity
- added 9 commits that reference this issue
on Oct 7, 2026 Trace-enabled snapshots currently expose counters but no bounded stage distributions or cached progress observation. Add exporter-owned aggregates for ten numeric stages, fixed reason/interval histograms, exact last frame identities and raw-clock progress ages. The worker publishes once per second; media append remains unchanged. Unobserved stages retain unknown identity/age, and combined-source inter-event intervals are explicitly not source cadence or content latency.
Connect the first real consumers: native deliveryEvidence snapshots, generated native/C#/Swift DTOs and validators, the shell qualification projection and retained QA snapshots. Extend the schema compiler only for bounded arrays and local definition references, with fixed-size and malformed/null/fractional fixtures. Preserve older peers that omit aggregates.
Validation: Release native build; 10 focused trace tests including concurrent diagnostic readers through finalization; 1,716 targeted C# lifecycle-contract/observer tests; 29 portable QA tests; generation/diff checks passed. The first aggregate build also passed the complete 1,516-test native suite (external network/NDI opt-ins were not executed). That first build exposed a genuine macOS libc++ compile failure for atomic<shared_ptr>; retained logs led to the portable short cache-pointer handoff, which is used only by exporter/snapshot threads and never media append. All 17 exact-head checks pass on 0e7c706, including native production/shell Windows, macOS Metal/stub, TSAN and CodeQL.
Two final-binary mixed-source GPU trials (45-second measurement, fixed 15-second warmup, 1080p60 buffer 2, eight fake 1080p30 I420 feeds plus 1080p/1440p60 BGRA inputs and recording) passed in inline/isolated monitor modes. All 5,419 Program deliveries had exact raw source/Program readiness evidence with zero trace/output/recording loss. All 358 sampled aggregate observations, across 46 cache revisions per trial, matched their raw exported event-prefix counts and identities; cached revisions remained immutable and camera/display/acquisition stages remained unknown. The isolated trial recorded zero render overruns, inline recorded two absorbed overruns; this is not parent release qualification. Artifacts/delivery-aggregate-823-mixed-final, -cache-evidence.json, -portable-focused.log, -contracts.log and -portable.log retain the evidence.
Final source 0e7c706; native tree 4ea6a17c7a34ef92e5fe4210d69368baec3d9f4c. Release /MP8, parallel 4. Core SHA256 0A2AA1824D5C2656A054D9A56260538F20F3768930CD5D67F75D16F55935578D. Source archive and exact binary hashes are retained in artifacts/delivery-aggregate-823-final-{source.zip,manifest.json}.
Continues #823 / parent #517; does not close either. Shared camera diagnostics, monitor/shell boundaries, per-source ages, interactive capture, actual content-latency qualification, instrumentation overhead, installed rehearsal and fleet evidence remain required. No performance default, installer or camera ABI change. Capture remains opt-in; aggregate counts cannot certify missing physical boundaries or a beta.
- added a commit that references this issue
on Oct 8, 2026 #823 / #517 qualification finding, retained without a PASS waiver
Matched same-binary timing experiment: source 879a287, native tree 33b8f0a96035f59ddbd4bfc314991855548e90e8, core SHA256 6C3715576855E064D3B056042ED415D774916D4227BF16DCC677371107BF7838. Timing collector enabled in both conditions; CPU preparation 1, GPU capture 0, buffer 2, eight fake 1080p30 I420 feeds, 1080p/1440p60 BGRA mappings, mixed Program at 1080p60, local recording; each inline/isolated trial measured 45 seconds after fixed 15-second warmup. Defaults/installed beta unchanged.
Trace-off trials passed output/recording/sample admission and timing-accounting checks. However the control already had substantial render tails: isolated mean 0.882 ms, p99 interval 11.749–12.730 ms, p99.9 interval 12.747–17.816 ms; two render completion overruns. Inline control had 64 overruns. These percentile intervals retain histogram/endpoint uncertainty and are not rounded point estimates.
Trace-on inline passed delivery/recording/raw trace. Trace-on isolated FAILED: five buffer underruns and five missing scheduled recording frames, four skipped CPU slots, nine render overruns; raw per-frame trace reports five misses/one gap with zero trace loss or export failures. Its high-percentile finite range was exceeded (>32.768 ms), so the timing distribution is INVALID rather than clamped to a healthy value. The matched cost comparison is INVALID because candidate delivery/recording failed; <1% trace overhead is not established.
The first divergence occurs 30.125–30.391 seconds into measurement. A 95.7085 ms CPU render-work interval precedes five misses for targets 2729–2733. Last timely Program submission 2728 occurred at 30.202053 seconds; next submission 2730 at 30.317059 seconds. Existing stage logs attribute this exceptional tick to Program 77.84 ms and monitor-request/build work 17.66 ms, not ingress (0.10 ms) or event drain (<0.114 ms maximum). The D3D renderer reports frame 2729: total 77.627 ms, encoder/virtual-camera/Program-buffer region 77.498 ms; drawing 0.040 ms, upload 0.009 ms, readback zero. Independently scheduled shell and multiview exports simultaneously reported keyed-mutex AcquireSync scopes around 77.022/77.004 ms despite zero requested timeout. This locates a GPU-handoff/API-stall region; it does NOT prove which producer call, driver scheduling, OS preemption or tracing caused it. Some other periodic tails are attributed to ingress fetch and require separate attribution.
No beta/default promotion or issue closure. Retain artifacts/render-work-823-first-{off,on}, raw captures, snapshots, stderr, exact manifest/source archive and artifacts/render-work-823-first-cost.json. Next focused diagnostic subdivision must distinguish producer mutex acquire, copy, release and remaining encoder-export work before changing the handoff. This evidence remains under the approved #823/#517 scope; it does not establish #825's separate wait-predicate amplification as the initiating cause.
Refined attribution from the failed run: the historical
vcam=77498stage label covers encoder-export prewarm/readiness/submit work in this buffered workload. The Program-buffer submission itself occurs later, in the reported 40 microsecond flush-stage region. The 77.498 ms initiating region is therefore specifically the live encoder-export handoff, not the Program-buffer producer submission. Source copy/draw were tiny and ingress was 0.10 ms in the exceptional tick. The particular encoder call and underlying driver/GPU-versus-OS scheduling cause remain unproven. The next #823 subdivision targets that encoder producer reservation/acquire/copy/release/enqueue path; it does not change media behavior or qualify a beta.Merged default-off GPU handoff CPU-scope attribution in #830 after all 17 exact-head checks passed (head 884c94f; merge c7521ab). Release build, 10 real GPU export tests each with profiling off/on, 10 raw-trace focused regressions and 24 portable tests pass. Frozen source/binary manifest and captures are retained.
Eight two-minute trials delivered 57,644 measured Program packets with zero native-buffer/recording/raw-trace losses, but the earlier five-loss/77 ms failure remains unresolved. No >=8 ms GPU handoff/readiness record recurred. The earlier harness omitted sampled logger-refusal counters, so absent slow records are limited evidence; #831 now retains those counters. The six-repeat CPU timing is also unqualified because brief portable Python tests ran during that set. Raw captures and this procedure deviation are retained. The uncontaminated first isolated trial had 17 render deadline misses and CPU p99/p99.9 over the parent limits; no strict performance/overhead/latency pass is asserted.
New opt-in Zoom handoff scopes in #831 have already reproduced a separate bottleneck with ordinary verbose logging off: publication body under runtime mutex 12–13 ms (six thumbnail encodes 6.5–7.1 ms), while Program configured()/fetch path waits 10–12 ms on that same mutex. Source confirms thumbnail base64 and large JSON construction remain under the publication lock. Filed new intake #832 with evidence and an off-lock packaging proposal; owner ranking pending. This does not establish the cause of the distinct 77 ms encoder stall.
#823/#517 remain open. No beta, installation, registration or default flag changed.
Both default-off CPU attribution slices are merged and parent acceptance remains open: #830 at c7521ab and #831 at 0db5714. Each has 17 successful exact-head checks. Final #831 source f5987f2 passed all 1,520 native tests, 59 enabled runtime/state tests, and 24 portable checks; external network/NDI opt-ins were not executed. Frozen source/binary manifests and complete raw captures are retained locally.
Final clean normal-logging trials: 14,410 measured Program deliveries, zero buffer/recording/raw-trace losses. CPU-tail gates still fail: inline p99 9.792–10.065 ms; isolated 8.198–8.465 ms. p99.9 upper uncertainty bounds exceed 12.5 ms in both. The earlier normal-logging pair retained 34/27 render deadline misses. Post-window 7/5 Program misses remain in the complete final captures and are excluded only from the explicitly pinned measured-window verdict, not deleted.
290 retained slow publication records and 85 slow fetch waits were observed; 76 waits overlapped publication bodies for at least 95% of their duration in the same steady-clock domain. Large optional thumbnail encoding/JSON packaging still occurs under the runtime mutex shared by Program configured()/fetch. The CPU record set includes warmup/shutdown and both processes report 7 logger drops (0 truncated/sink failures), so absence is not complete negative evidence. This supports the reproducible serialization finding #832 and its focused off-lock packaging proposal; owner ranking is pending. It does not attribute the separate earlier 77 ms encoder stall/five losses or prove GPU duration/kernel wait owner.
The #830 merge inadvertently triggered GitHub issue closure from a negated closing-keyword phrase in its body. #823 has been reopened and both PR bodies corrected. Parent gates remain unmet: camera/shared/monitor/display trace coverage, calibrated content latency, <1% p95 trace cost, installed/real-meeting/A/V/resource/hardware qualification. No beta/default/settings/registration/installation change. Installed October 5 core hash reverified unchanged as A0F1E91D9D3041F01975CC6DE72A69DBEA7BEE8BC3916E961F58764771F7F5F0.
- added a commit that references this issue
on Oct 10, 2026
Parent #517, owner-approved render-delivery specification/completion plan. This tracks the remaining diagnostics implementation, without changing the parent's backlog rank or authorizing default flags.
Current counters and sampled source admissions cannot identify the first divergent per-frame boundary or establish actual added content latency. Camera reader delivery logs still perform synchronous file logging. The approved spec requires fixed numeric events, bounded storage/export, and native/C#/Swift mixed-version contracts.
Implement preallocated 16 MiB trace storage with nonwaiting media-thread append, explicit lost/overwritten-event counters and a bounded background export on explicit diagnostic capture. Carry session/source epochs, Program sequence, selected source frame identity, stage, monotonic timestamp/frequency and refusal reason; include layout attribution. Connect actual producer GPU completion, Program scheduling/readiness/delivery, camera publication/reader emission, monitor completion and shell submission. Submission is not display proof. Preserve existing camera pixel/correlation ABIs; separate optional diagnostics channel and least-privilege access. No meeting names, secrets, URLs or media payloads.
Version deliveryEvidence and fixtures across native/C#/Swift, preserve unknown older peers, and connect the first real diagnostics/support/QA consumers in each implementation slice. Reject partial/malformed/lossy trace for boundary-dependent PASS. Trace captures must not cause unsafe camera-DLL background threads across unload.
Done when meaningful fault/identity/wrap/concurrency tests pass; actual-source/Program/independent OS receiver correlation locates injected divergences; real workload timing compares instrumentation on/off with added p95 render cost below 1%; explicit export/shutdown bounds and resource limits hold. Complete trace and actual content-latency median <=5 ms / p95 <=one 60fps frame versus matched baseline remain parent release gates. Physical display, real SDK, A/V and fleet qualification cannot be inferred from trace alone. Defaults stay off pending #517 acceptance.