Repository navigation
fix(testdata): make the Netflix benchmark harness honest and portable (ADR-1192) - #1334
Conversation
|
The It is not The real mechanism. Instrumenting A detail that matters for your throughput table. The duplicate write is not the only consequence. Draining the tail early also emits the last batch-boundary frame without the Two things from your findings I would keep exactly as they are:
Also confirming your refusal to regenerate |
|
Follow-up measurement on Current master 2/80 = 2.5%, against your 10/40 = 25% on A correction to my own work, so it is not repeated. I ran what looked like a clean A/B — 0/80 wrong with #1343's tree versus 11/80 on a "master" tree, p≈0.0006 — and it was confounded: the control worktree had drifted to master+#1325+#1321 while my branch sat at merge-base It could not have supported it anyway, which is the useful part: the So: #1343 fixes the filter's explicit |
ce8159e to
568e512
Compare
… (ADR-1192) Re-ran the Netflix benchmark suite on cd52f26 for epic #1245 items 1 and 5. All three fixtures reproduce on CPU, CUDA and SYCL through the FFmpeg filter path against a container-built current-master libvmaf, but every backend's pooled score has drifted from testdata/netflix_benchmark_results.json (recorded by PR #309 on 2026-05-02): CPU +2.83e-06, CUDA -1.07e-03, SYCL -1.40e-03 on the 576x324 pair. A rebuild of 5a08030 — the commit before the 2026-09-06 GPU merges #1307/#1312/#1324 — shows the same drift, so none of it comes from today's merges. The snapshot is deliberately NOT regenerated (ADR-1192) and no throughput baseline is recorded, because the run also reproduced two pre-existing GPU defects: - vmaf --threads N aborts on every GPU backend (exit 234, "context could not be synchronized"); without --threads both CUDA and SYCL score correctly and are bit-stable over 10 runs. bench_all.sh hard-codes --threads 1. - The libvmaf_cuda FFmpeg filter returns a wrong pooled score in 10 of 40 runs on master and 8 of 40 on 5a08030 — inside binomial noise of each other. Harness fixes in the same change: - bench_all.sh kept its stderr on /dev/null and relabelled every non-zero exit as "backend likely unavailable", which is how a hard abort passed for a missing device for months. It now captures stderr per row and prints FAIL with the exit code and the real last line. Its flag sets also drop --no_vulkan, unrecognized since ADR-0726 removed the Vulkan backend. - benchmark_netflix.py hard-coded /home/kilian/dev/ffmpeg-8/ffmpeg (gone) and /dev/dri/renderD130 for the SYCL/QSV import (now the AMD iGPU on the bench host, so the SYCL rows failed outright). Both are environment overrides now, VMAF_FFMPEG and the new VMAF_SYCL_RENDER_NODE, per the ADR-0792 pattern. No golden assertions touched; no snapshot regenerated. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Each dropped row restates one origin/master already carries; master is the authoritative record. Verified with scripts/ci/check-state-md-rows.sh. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
568e512 to
67fcc8c
Compare
Summary
Re-ran the Netflix benchmark suite on current master (
cd52f2670) for epic #1245 items 1 and 5, and made the harness honest and portable so the next person can reproduce it. The scores do not match the recorded snapshot on any backend, so per the epic's own gate no fresh throughput baseline is recorded andtestdata/netflix_benchmark_results.jsonis deliberately left untouched (ADR-1192). The run also reproduced two GPU defects — both confirmed pre-existing by rebuilding the commit before today's GPU merges.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificWhat was measured
Host
ryzen-4090-arc: AMD Ryzen 9 9950X3D (32 threads), 60 GiB, Linux 7.2.3-1-cachyos, RTX 4090 (driver 610.57.04), Arc A380. All runs insidevmaf-dev-mcp(Ubuntu 26.04, oneAPI DPC++ 2026.1.1, CUDA 13.3.73, FFmpeg n9.0.1-17-gfde3691) against an out-of-source libvmaf built--buildtype=release -Db_lto=false -Dc_args=-march=nativeand injected viaLD_LIBRARY_PATH. The workstation was not idle — 1-minute load average 7–19 throughout (a container image build plus other agent sessions); the timing sweep specifically ran at load 9.6–9.8.Scores versus
testdata/netflix_benchmark_results.json(CUDA at its modal value)src01_576x324src01_576x324src01_576x324checker_1080p_mildchecker_1080p_mildchecker_1080p_mildchecker_1080p_heavyNone of this is today's merges. A rebuild of
5a080300e(immediately before #1307, #1312, #1324) with identical flags gives CUDA 76.667830 and SYCL 76.667745 on the 576x324 pair — the same drift versus a snapshot last written by PR #309 on 2026-05-02. What the three merges did change: CPU per-frame values moved by up to ~8e-6 (pooled unchanged at six decimals) and the SYCLframes[].metricskey set shrank 35 → 24. The direction of the four-month drift is CUDA and SYCL converging towards CPU (recorded CUDA sat 1.07e-3 above CPU; it now tracks CPU to 2e-5).Throughput — observation only, not a recorded baseline
Median of 5, whole-process wall time, load 9.6–9.8. Published in the new docs page; not written into
docs/benchmarks.md, because a timing table next to a CUDA score that is wrong a quarter of the time would be misleading.Two GPU defects reproduced (both pre-existing)
vmaf --threads Naborts on every GPU backend —T-GPU-CLI-THREADS-CTX-SYNC-2026-09-06. Exit 234,context could not be synchronized, no output. Drop--threadsand both backends succeed and are bit-stable 10/10 (CUDA 76.667830, SYCL 76.667746).bench_all.shhard-codes--threads 1, so its GPU rows were masked failures.libvmaf_cudaFFmpeg filter is non-deterministic —T-CUDA-FFMPEG-FILTER-NONDETERMINISM-2026-09-06. 10 of 40 runs oncd52f2670and 8 of 40 on5a080300ereturned a non-modal pooled score; bad runs corrupt one or two individual frames (frame 1 = 0.0 vs CPU 82.639803). CPU 10/10 and SYCL 10/10 through the same FFmpeg are bit-stable, and CUDA through the CLI with no thread pool is 10/10 stable. 20 % vs 25 % over n=40 each is inside binomial noise, so the merges neither caused nor measurably worsened it.Both match the statically-derived
T-UPSTREAM-1305-CUDA-DRAIN-BATCH-THREAD-GLOBAL-2026-09-03row; these runs are the empirical reproducer that row said the fix PR would need. The mechanism is not proven here — only the symptom, the thread-pool dependency and the pre-existence.Checklist
make format && make lintis green locally —pre-commit run --fileson all 14 touched files passes (shellcheck, shfmt, markdownlint, black, ruff, semgrep, ADR + venv + model-single-source gates).meson test -C build. — no C/C++ source touched; the changed files are twotestdata/harness scripts and documentation./cross-backend-diffand the worst ULP is ≤ 2. — no GPU source touched; the cross-backend numbers this PR reports are measurement output, not a code change..c/.cpp/.cu/.h/.hpp, it has the appropriate license header. — none added.!orBREAKING CHANGE:. — not breaking.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt.Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR with a row in the appropriate section — two new Open-bug rows plus a dated update note.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.testdata/netflix_benchmark_results.jsonis untouched by design (ADR-1192).Cross-backend numerical results
Performance (if
perforfeat)See "Throughput" above. No performance change is claimed or intended by this PR.
Deep-dive deliverables (ADR-0108)
docs/research/2026-09-06-netflix-benchmark-rerun.md.docs/adr/1192-netflix-bench-snapshot-drift-not-regenerated.md## Alternatives considered(regenerate now / delete the CUDA rows / loosen tolerance / keep and record).AGENTS.mdinvariant note —core/AGENTS.md§"Backend-engagement foot-guns" gains the--threads-aborts-GPU and never-discard-stderr invariants plus the updated key counts.changelog.d/fixed/netflix-bench-harness-honesty.md, rendered viascripts/release/concat-changelog-fragments.sh --write.docs/rebase-notes.mdRN-2026-09-06.Reproducer
Known follow-ups
testdata/netflix_benchmark_results.jsonis blocked onT-GPU-CLI-THREADS-CTX-SYNC-2026-09-06andT-CUDA-FFMPEG-FILTER-NONDETERMINISM-2026-09-06closing (ADR-1192). The regenerating PR must cite that ADR.docs/benchmarks.mdstill carries its 2026-05 throughput table and references amake benchtarget that does not exist in theMakefile. Left alone here to keep this PR scoped./run-netflix-benchskill points attestdata/compare_combined.py, which comparesscores_cpu_*.jsonagainstscores_sycl_a380_*.jsonand never reads the Netflix snapshot. The new docs page records the mismatch; fixing the skill is a separate change.