Repository navigation
perf(cli): read each input on its own thread, two frames ahead of scoring - #1635
Merged
Merged
Conversation
…ring The vmaf CLI now reads the reference and the distorted input on two reader threads, each up to two frames ahead of the scoring loop. At 3840x2160 the serial read of both frames on the scoring thread cost about 7 ms per frame for every backend; `--feature psnr` now runs at 3.4 ms per frame on 16 CPU threads (6.9 on master) and the psnr, motion and adm twins on an Arc B580 at 3.8 to 4.6 ms (7.4 to 8.6 on master). Runs limited by extraction, such as the default model on 16 threads, keep their speed. Each input gets a FrameReader. Its thread runs fetch_picture() into a two-slot ring, reserving a slot before it takes a pool picture, so it never holds more than two pictures; the picture pool grows by four. score_frames() pops one frame per reader per step and still makes every vmaf_read_pictures() call in order on the main thread, so the frames scored, their pairing, the progress line and the ADR-1262 exit codes are unchanged. The readers stop after --frame_cnt frames or the first frame that ends or fails their stream; shutdown drains both rings before joining either reader, which wakes a reader blocked on the pool. Handles that may share a read position (one file or pipe on both sides and --no-reference on POSIX, anything but two regular files on Windows) and a failed thread creation keep the inline read. JSON is identical to master at --precision max on 22 CPU cases (Netflix 8/10-bit, Y4M, --frame_cnt, --frame_skip_*, --subsample, --threads 0/4/8/16, truncated and short inputs, same-file input, BBB 4K) and 6 SYCL cases on the B580. test_vmaf_frame_readahead pins order, pairing, limits and exits; it and the CLI runs are clean under ThreadSanitizer. Closes T-CLI-SYNC-FRAME-READ-FLOOR-2026-09-29. ADR-1366, Research-1366. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Session C measured the read-ahead on the i9-12900K's integrated UHD 770 and CPU motion on 16 threads, master and branch interleaved, 5 repetitions each. The branch was faster in every repetition: CPU --threads 16 motion 4.72 to 2.84 ms per 4K frame, UHD 770 psnr 24.84 to 20.63, motion 13.07 to 9.65, default model 70.67 to 67.29. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
lusoris
force-pushed
the
perf/cli-frame-readahead
branch
from
September 30, 2026 10:55
f401ee8 to
2a90e49
Compare
This was referenced Sep 30, 2026
7 of 13 tasks
lusoris
added a commit
that referenced
this pull request
Oct 8, 2026
…9cb9479f2 The fork's records of its Netflix/vmaf pull requests and of the upstream defects it tracks were last checked on 2026-10-01. Upstream master has moved to 9cb9479f2 since, with the MSVC series, the arm64 ADM kernels, the fused SpEED filter and a rewrite of integer_compute_adm(). This brings the records up to date. No fork code and no score changes. - known-upstream-bugs.md: 42 open pull requests instead of 15, which of them needed a rebase or a rework against 9cb9479f2, the #1494 and best15 status, the integer AIM change upstream (cffd5b77d), the arm64 gap in #1602, and what is new upstream since the parity pin. The pin heading is unchanged. - state.md: dated updates on the rows for #1422, #1551 (both halves), #1602, #1605, #1635, #1636 and #955. - upstream_parity.d: the APSNR zero-error row also names #1618, which changes the same cap upstream; the allowlist page is regenerated. - rebase-notes.d: what a port of upstream's arm64 ADM code (8bc5a5c6a, b41d2340a) must not copy, and the closed state of #1551. Signed-off-by: Lusoris <lusoris@proton.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The
vmafCLI now reads the reference and the distorted input on two reader threads, each up to two frames ahead of scoring. Before,run_frame_loop()read the reference frame, then the distorted frame, then scored the pair, all on the main thread; at 3840x2160 those two reads cost about 7 ms per frame whatever the backend did, so the GPU twins and cheap CPU extractors sat on that floor. Now--feature psnrat 4K takes 3.4 ms per frame on 16 CPU threads (6.9 on master) and thepsnr,motionandadmtwins on an Arc B580 take 3.8 to 4.6 ms (7.4 to 8.6 on master). Scores, frame order,--frame_cnt,--frame_skip_*, the progress line and the exit codes are unchanged; the JSON is identical to master at--precision max.Design (ADR-1366)
FrameReader. Itsstd::threadruns the existingfetch_picture()(pool picture, file read, plane copy) into a two-slot ring and reserves a slot before it takes a pool picture, so one reader never holds more than two pictures.preallocate_cli_pictures()adds2 * kReadaheadDepth(four) pictures to the pool, so the readers never draw on the pictures the pool was sized for.score_frames()pops one frame per reader per step and still makes everyvmaf_read_pictures()call, in order, on the main thread.classify_frame_fetch(),release_unpaired_pictures()and the progress line are unchanged, so the ADR-1262 exits (0 withended before, 102 with no report) are unchanged.--frame_cntframes, after the first frame that ends or fails its stream, or when stopped. Shutdown callsrequest_stop()on both readers (set the flag, drain the ring back to the pool, which wakes a reader blocked invmaf_fetch_preallocated_picture()) before joining either.st_dev/st_ino(one pipe or file on both sides, and--no-reference, which opens the distorted file twice); on Windows anything but two regular files. A failed thread creation also falls back to inline reading.vmaf.cppusesstd::thread/std::mutex/std::condition_variable; the CLI targets gain thethreadsdependency.Type
feat— new featurefix— bug fixperf— performance improvementrefactor— no behavior changedocs— documentation onlytest— test-onlybuild/ci— tooling / infraport— cherry-pick from upstream Netflix/vmafsycl/cuda/simd— backend-specificChecklist
make format && make lintis green locally. Run piecewise, see "Verification": clang-format 23.1.1, clang-tidy 22.1.2 and cppcheck 2.19 onvmaf.cpp,make docs-fragments-check, and the state, citation, ADR-link, ADR-numbering, research-ID and assertion-density gates.python3 scripts/ci/run_meson_test.py -- -C build. The fast suite passes excepttest_gpu_picture_pool_uaf, which times out identically on the master build (see "Verification")./cross-backend-diffand the worst ULP is ≤ 2. No kernel changed; SYCL runs are bit-identical to master (below)..c/.cpp/.cu/.h/.hpp, it has the appropriate license header (seeCONTRIBUTING.md). No new C/C++ file; the new test script carries the fork header.!orBREAKING CHANGE:and the migration path is documented below. Not breaking.docs/adr/_index_fragments/<NNNN-slug>.mdand the slug is appended todocs/adr/_index_fragments/_order.txt— do not editdocs/adr/README.mddirectly (regenerated byscripts/docs/concat-adr-index.sh; see ADR-0221).Bug-status hygiene (ADR-0165)
docs/state.mdupdated in this PR with a row in the appropriate section.T-CLI-SYNC-FRAME-READ-FLOOR-2026-09-29is filed under Recently closed with the evidence below.Netflix golden-data gate (ADR-0024)
assertAlmostEqual(...)score in the Netflix golden Python tests.Cross-backend numerical results
No kernel changed. Master
2d9d5b069against this branch, JSON compared with the wall-clockfpsfield removed:Stderr was byte-identical in every CPU case except the SpEED
est_params: covariance matrix was singular on N of M solvesline of the 4K--threads 16run, which varies between runs of master itself (12 or 16 solves: it sums per-thread contexts).Performance (if
perforfeat)BBB 3840x2160 8-bit 4:2:0, i9-12900K + Arc B580, Docker Desktop WSL2 (22 CPUs) shared with other agents' builds. Master and branch run back to back in each repetition; median of 5 with [min..max], ms per frame. The in-loop method reads the last FPS of the progress line under a pseudo-terminal (frames / wall time since the frame loop started, so start-up is excluded); GPU runs hold
flock /f/gpu.lockfor the whole run. Session B (quieter) shown, plus session C (load below 2) for CPUmotionand the integrated UHD 770 (level_zero:1); session A (load 6 to 14) gave the same ratios and is in Research-1366.--threadspsnrfloat_psnrpsnrfloat_psnrpsnrfloat_psnrmotionvifadmmotionpsnrmotionThe requested startup-difference method,
(t(--frame_cnt 22) - t(--frame_cnt 2)) / 20, 3 repetitions, agrees on the CPU:--threads 16psnr6.99 [6.91..7.17] to 3.31 [2.61..3.37],float_psnr8.13 [7.70..9.03] to 6.01 [5.53..6.17]; serialpsnr7.27 [6.97..7.45] to 3.55 [3.36..3.69],float_psnr16.04 [15.87..16.54] to 11.62 [11.60..11.81]. For the CPU default model and every GPU row it is dominated by start-up variance of hundreds of ms (some repetitions came out negative), so the in-loop numbers are the result there; the digest lists the raw values.Queue depth (branch built with depth 1, 2, 3, interleaved): depth 1 was slower than depth 2 in four of five serial
psnrrepetitions (3.40 against 3.10 ms), depth 3 gained nothing; depth 2 is kept.Deep-dive deliverables (ADR-0108)
## Alternatives considered: a single producer thread,USE_DIRECT_READ, depth 1 and depth 3, an opt-out flag, and detaching readers on error.AGENTS.mdinvariant note —core/tools/AGENTS.md"Frame read-ahead (ADR-1366)": readers only callfetch_picture()/vmaf_picture_unref(), slot reservation before the pool fetch, stop both readers before joining either, the--frame_cntbound, the inline path for shared handles, and how to resolve an upstreamvmaf.cloop conflict.changelog.d/changed/perf-cli-frame-readahead.md;CHANGELOG.mdre-rendered.docs/rebase-notes.md"perf/cli-frame-readahead — per-input reader threads in thevmafCLI (ADR-1366)".Reproducer
Verification
In
vmaf-dev-mcp:oclocon the WSL2 workstation (CPU build gcc 15, release,-Denable_dnn=disabled; SYCL build icx with AOTbmg-g21,adl-s, B580 aslevel_zero:0):test_vmaf_frame_readahead(new, 8 cases: order and pairing against known per-frame psnr_y,--threads 8,--frame_cnt, no reader past--frame_cnt,--frame_skip_dist, same-file inline path, early end, failed read),test_vmaf_read_error_exit,test_cli_parse,test_cli_parse_long_only_args,test_cli_feature_backend,test_vmaf_close_retryandtest_spinnerpass. The new test fails on three injected mutants (ring overwrite, dropped frame, ignored--frame_cnt).test_gpu_picture_pool_uaftimes out at 30 s on the master build as well.vmafexec_test.pyandvmafexec_feature_extractor_test.py, 141 passed against the branch CPU build.-Db_sanitize=thread, run in a seccomp-unconfined container so TSan can disable ASLR): Netflix 576x324 serial and--threads 4default model,--frame_cnt 7, Y4M--threads 2, truncated Y4M (exit 102), short distorted, same file, and the new test: 0 reports.-Db_lto=false): 0 findings invmaf.cpp, the same header findings as master (vmaf.cpp is at its baseline of 0). cppcheck 2.19 with the CI flags: 0 findings. clang-format 23.1.1: clean.make docs-fragments-checkinpython:3.14-slim,check-source-adr-citations.py,check-state-md-rows.sh,check-adr-links.py,check-research-digest-ids.py,check-adr-numbering.shandassertion-density.shpass.b20472dcb(fix(sycl): honour the CPU psnr, ssim and float_motion options on the SYCL twins #1624 SYCL option parity, fix(sycl): make motion_sycl bit-exact with the CPU motion #1628 motion_sycl; no overlap with the CLI loop) after the measurements; the CLI tests were re-run after the fix(sycl): honour the CPU psnr, ssim and float_motion options on the SYCL twins #1624 rebase. The SYCL timings use one libvmaf built from2d9d5b069for both CLIs.Known follow-ups
vmaf_read_pictures()and the device wait). Overlapping that needs double-buffered device pictures in the backend. Not profiled here.USE_DIRECT_READon top of read-ahead would remove one plane copy per frame on the reader threads; not measured.fstatand Windows_fstat64branches are compiled only on their platforms; CI covers MinGW and MSVC).🤖 Generated with Claude Code
no ffmpeg-patches update needed: the change is internal to the vmaf CLI's frame loop; no libvmaf API, header, CLI flag or meson option changes, and the FFmpeg filters do not use the CLI.