Skip to content

perf(cli): read each input on its own thread, two frames ahead of scoring - #1635

Merged
lusoris merged 3 commits into
masterfrom
perf/cli-frame-readahead
Sep 30, 2026
Merged

lusoris merged 3 commits into
masterfrom
perf/cli-frame-readahead

Conversation

@lusoris

@lusoris lusoris commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The vmaf CLI now reads the reference and the distorted input on two reader threads, each up to two frames ahead of scoring. Before, run_frame_loop() read the reference frame, then the distorted frame, then scored the pair, all on the main thread; at 3840x2160 those two reads cost about 7 ms per frame whatever the backend did, so the GPU twins and cheap CPU extractors sat on that floor. Now --feature psnr at 4K takes 3.4 ms per frame on 16 CPU threads (6.9 on master) and the psnr, motion and adm twins on an Arc B580 take 3.8 to 4.6 ms (7.4 to 8.6 on master). Scores, frame order, --frame_cnt, --frame_skip_*, the progress line and the exit codes are unchanged; the JSON is identical to master at --precision max.

Design (ADR-1366)

  • Each input gets a FrameReader. Its std::thread runs the existing fetch_picture() (pool picture, file read, plane copy) into a two-slot ring and reserves a slot before it takes a pool picture, so one reader never holds more than two pictures. preallocate_cli_pictures() adds 2 * kReadaheadDepth (four) pictures to the pool, so the readers never draw on the pictures the pool was sized for.
  • score_frames() pops one frame per reader per step and still makes every vmaf_read_pictures() call, in order, on the main thread. classify_frame_fetch(), release_unpaired_pictures() and the progress line are unchanged, so the ADR-1262 exits (0 with ended before, 102 with no report) are unchanged.
  • A reader stops after --frame_cnt frames, after the first frame that ends or fails its stream, or when stopped. Shutdown calls request_stop() on both readers (set the flag, drain the ring back to the pool, which wakes a reader blocked in vmaf_fetch_preallocated_picture()) before joining either.
  • Inputs that may share a read position keep the old inline read: on POSIX two handles with the same st_dev/st_ino (one pipe or file on both sides, and --no-reference, which opens the distorted file twice); on Windows anything but two regular files. A failed thread creation also falls back to inline reading.
  • No flag: outputs do not depend on the path taken. vmaf.cpp uses std::thread/std::mutex/std::condition_variable; the CLI targets gain the threads dependency.

Type

  • feat — new feature
  • fix — bug fix
  • perf — performance improvement
  • refactor — no behavior change
  • docs — documentation only
  • test — test-only
  • build / ci — tooling / infra
  • port — cherry-pick from upstream Netflix/vmaf
  • sycl / cuda / simd — backend-specific

Checklist

  • Commits follow Conventional Commits (the commit-msg hook enforces this).
  • make format && make lint is green locally. Run piecewise, see "Verification": clang-format 23.1.1, clang-tidy 22.1.2 and cppcheck 2.19 on vmaf.cpp, make docs-fragments-check, and the state, citation, ADR-link, ADR-numbering, research-ID and assertion-density gates.
  • Unit tests pass: python3 scripts/ci/run_meson_test.py -- -C build. The fast suite passes except test_gpu_picture_pool_uaf, which times out identically on the master build (see "Verification").
  • If I touched any SIMD/GPU code path, I ran /cross-backend-diff and the worst ULP is ≤ 2. No kernel changed; SYCL runs are bit-identical to master (below).
  • If I touched a feature extractor with SIMD/GPU twins, I either updated every twin or listed the gap under "Known follow-ups" below. No extractor changed.
  • If I added a new .c / .cpp / .cu / .h / .hpp, it has the appropriate license header (see CONTRIBUTING.md). No new C/C++ file; the new test script carries the fork header.
  • If this is a breaking change, the commit message uses ! or BREAKING CHANGE: and the migration path is documented below. Not breaking.
  • If this PR adds an ADR, the ADR row lives in docs/adr/_index_fragments/<NNNN-slug>.md and the slug is appended to docs/adr/_index_fragments/_order.txt — do not edit docs/adr/README.md directly (regenerated by scripts/docs/concat-adr-index.sh; see ADR-0221).

Bug-status hygiene (ADR-0165)

  • docs/state.md updated in this PR with a row in the appropriate section. T-CLI-SYNC-FRAME-READ-FLOOR-2026-09-29 is filed under Recently closed with the evidence below.

Netflix golden-data gate (ADR-0024)

  • I did not modify any assertAlmostEqual(...) score in the Netflix golden Python tests.
  • If I believe a golden value must change, I have explained why below AND pinged @lusoris for a CODEOWNERS exception. Not applicable.

Cross-backend numerical results

No kernel changed. Master 2d9d5b069 against this branch, JSON compared with the wall-clock fps field removed:

CPU build, 22 cases, all identical at --precision max:
  Netflix 576x324 default model (serial, --threads 4/16, + --feature psnr), psnr only,
  --frame_cnt 10 (serial, --threads 4) and 1, --frame_skip_ref 3 --frame_skip_dist 1,
  --subsample 3, 10-bit psnr, Y4M psnr and default model --threads 8,
  Y4M reference truncated mid-frame (exit 102, no report, both), raw truncated (exit 2, both),
  distorted / reference ending after 20 frames (exit 0, 20 frames), same file on both sides,
  BBB 3840x2160 50 frames default model + psnr --threads 16, float_psnr serial
SYCL build, Arc B580 (level_zero:0), 6 cases, all identical at --precision max:
  --backend sycl psnr / vif / default model, Netflix 576x324 (48 frames) and BBB 4K (50 frames)

Stderr was byte-identical in every CPU case except the SpEED est_params: covariance matrix was singular on N of M solves line of the 4K --threads 16 run, which varies between runs of master itself (12 or 16 solves: it sums per-thread contexts).

Performance (if perf or feat)

BBB 3840x2160 8-bit 4:2:0, i9-12900K + Arc B580, Docker Desktop WSL2 (22 CPUs) shared with other agents' builds. Master and branch run back to back in each repetition; median of 5 with [min..max], ms per frame. The in-loop method reads the last FPS of the progress line under a pseudo-terminal (frames / wall time since the frame loop started, so start-up is excluded); GPU runs hold flock /f/gpu.lock for the whole run. Session B (quieter) shown, plus session C (load below 2) for CPU motion and the integrated UHD 770 (level_zero:1); session A (load 6 to 14) gave the same ratios and is in Research-1366.

Backend --threads Feature master branch speedup
CPU 16 psnr 6.90 [6.70..7.38] 3.41 [3.23..3.75] 2.02x
CPU 16 float_psnr 7.44 [6.67..7.64] 5.11 [4.68..6.19] 1.46x
CPU 16 default model 21.45 [19.00..26.94] 20.51 [17.30..23.33] 1.05x
CPU 0 psnr 8.19 [7.65..8.88] 3.97 [3.64..4.25] 2.06x
CPU 0 float_psnr 17.59 [17.32..17.99] 12.01 [11.07..12.26] 1.46x
SYCL B580 0 psnr 7.81 [7.44..8.10] 4.13 [3.99..4.22] 1.89x
SYCL B580 0 float_psnr 7.71 [7.15..9.45] 4.45 [4.37..5.55] 1.73x
SYCL B580 0 motion 7.42 [7.10..8.04] 3.79 [3.29..4.14] 1.96x
SYCL B580 0 vif 8.49 [7.92..9.03] 6.79 [6.64..7.05] 1.25x
SYCL B580 0 adm 8.59 [7.50..9.59] 4.60 [4.29..5.27] 1.87x
SYCL B580 0 default model 63.73 [52.60..74.52] 59.88 [50.38..61.46] 1.06x
CPU (session C) 16 motion 4.72 [4.55..5.40] 2.84 [2.72..3.15] 1.66x
SYCL UHD 770 (session C) 0 psnr 24.84 [24.06..26.25] 20.63 [20.40..20.88] 1.20x
SYCL UHD 770 (session C) 0 motion 13.07 [12.56..14.17] 9.65 [9.57..10.52] 1.35x
SYCL UHD 770 (session C) 0 default model 70.67 [69.54..72.31] 67.29 [65.32..68.07] 1.05x

The requested startup-difference method, (t(--frame_cnt 22) - t(--frame_cnt 2)) / 20, 3 repetitions, agrees on the CPU: --threads 16 psnr 6.99 [6.91..7.17] to 3.31 [2.61..3.37], float_psnr 8.13 [7.70..9.03] to 6.01 [5.53..6.17]; serial psnr 7.27 [6.97..7.45] to 3.55 [3.36..3.69], float_psnr 16.04 [15.87..16.54] to 11.62 [11.60..11.81]. For the CPU default model and every GPU row it is dominated by start-up variance of hundreds of ms (some repetitions came out negative), so the in-loop numbers are the result there; the digest lists the raw values.

Queue depth (branch built with depth 1, 2, 3, interleaved): depth 1 was slower than depth 2 in four of five serial psnr repetitions (3.40 against 3.10 ms), depth 3 gained nothing; depth 2 is kept.

Deep-dive deliverables (ADR-0108)

  • Research digest — Research-1366: where the floor comes from, why the two streams can be read concurrently, the pool-sizing and shutdown argument against deadlock, the shared-object hazard, bit-exactness tables and both timing sessions.
  • Decision matrix — in ADR-1366 ## Alternatives considered: a single producer thread, USE_DIRECT_READ, depth 1 and depth 3, an opt-out flag, and detaching readers on error.
  • AGENTS.md invariant note — core/tools/AGENTS.md "Frame read-ahead (ADR-1366)": readers only call fetch_picture()/vmaf_picture_unref(), slot reservation before the pool fetch, stop both readers before joining either, the --frame_cnt bound, the inline path for shared handles, and how to resolve an upstream vmaf.c loop conflict.
  • Reproducer / smoke-test command — below.
  • CHANGELOG fragment — changelog.d/changed/perf-cli-frame-readahead.md; CHANGELOG.md re-rendered.
  • Rebase note — docs/rebase-notes.md "perf/cli-frame-readahead — per-input reader threads in the vmaf CLI (ADR-1366)".

Reproducer

# Scores are unchanged; the 4K per-frame cost drops (compare against master).
vmaf -r ref_3840x2160.yuv -d dis_3840x2160.yuv --width 3840 --height 2160 \
     --pixel_format 420 --bitdepth 8 --feature psnr --no_prediction \
     --frame_cnt 61 --threads 16 --json -o /tmp/out.json
# Order, pairing, limits and exits (fast suite, CPU, 64x48 x 40 frames):
python3 scripts/ci/run_meson_test.py -- -C build test_vmaf_frame_readahead test_vmaf_read_error_exit

Verification

In vmaf-dev-mcp:ocloc on the WSL2 workstation (CPU build gcc 15, release, -Denable_dnn=disabled; SYCL build icx with AOT bmg-g21,adl-s, B580 as level_zero:0):

  • test_vmaf_frame_readahead (new, 8 cases: order and pairing against known per-frame psnr_y, --threads 8, --frame_cnt, no reader past --frame_cnt, --frame_skip_dist, same-file inline path, early end, failed read), test_vmaf_read_error_exit, test_cli_parse, test_cli_parse_long_only_args, test_cli_feature_backend, test_vmaf_close_retry and test_spinner pass. The new test fails on three injected mutants (ring overwrite, dropped frame, ignored --frame_cnt).
  • Fast suite: 178 of 181 pass, 2 skip (no GPU configured); test_gpu_picture_pool_uaf times out at 30 s on the master build as well.
  • Netflix golden subset: vmafexec_test.py and vmafexec_feature_extractor_test.py, 141 passed against the branch CPU build.
  • ThreadSanitizer (gcc, -Db_sanitize=thread, run in a seccomp-unconfined container so TSan can disable ASLR): Netflix 576x324 serial and --threads 4 default model, --frame_cnt 7, Y4M --threads 2, truncated Y4M (exit 102), short distorted, same file, and the new test: 0 reports.
  • clang-tidy 22.1.2 (CI config, -Db_lto=false): 0 findings in vmaf.cpp, the same header findings as master (vmaf.cpp is at its baseline of 0). cppcheck 2.19 with the CI flags: 0 findings. clang-format 23.1.1: clean.
  • make docs-fragments-check in python:3.14-slim, check-source-adr-citations.py, check-state-md-rows.sh, check-adr-links.py, check-research-digest-ids.py, check-adr-numbering.sh and assertion-density.sh pass.
  • Rebased onto b20472dcb (fix(sycl): honour the CPU psnr, ssim and float_motion options on the SYCL twins #1624 SYCL option parity, fix(sycl): make motion_sycl bit-exact with the CPU motion #1628 motion_sycl; no overlap with the CLI loop) after the measurements; the CLI tests were re-run after the fix(sycl): honour the CPU psnr, ssim and float_motion options on the SYCL twins #1624 rebase. The SYCL timings use one libvmaf built from 2d9d5b069 for both CLIs.

Known follow-ups

  • On the B580 the single-feature runs now stop at about 4 ms per 4K frame; what remains is on the scoring thread (host-to-device upload of the pair inside vmaf_read_pictures() and the device wait). Overlapping that needs double-buffered device pictures in the backend. Not profiled here.
  • USE_DIRECT_READ on top of read-ahead would remove one plane copy per frame on the reader threads; not measured.
  • Not run: CUDA, HIP and Metal builds (the change is backend-independent host code), and a Windows / macOS build (the POSIX fstat and Windows _fstat64 branches are compiled only on their platforms; CI covers MinGW and MSVC).

🤖 Generated with Claude Code

no ffmpeg-patches update needed: the change is internal to the vmaf CLI's frame loop; no libvmaf API, header, CLI flag or meson option changes, and the FFmpeg filters do not use the CLI.

@github-actions github-actions Bot added the type:perf Performance improvement label Sep 29, 2026
lusoris and others added 3 commits September 30, 2026 12:53
…ring

The vmaf CLI now reads the reference and the distorted input on two
reader threads, each up to two frames ahead of the scoring loop. At
3840x2160 the serial read of both frames on the scoring thread cost
about 7 ms per frame for every backend; `--feature psnr` now runs at
3.4 ms per frame on 16 CPU threads (6.9 on master) and the psnr, motion
and adm twins on an Arc B580 at 3.8 to 4.6 ms (7.4 to 8.6 on master).
Runs limited by extraction, such as the default model on 16 threads,
keep their speed.

Each input gets a FrameReader. Its thread runs fetch_picture() into a
two-slot ring, reserving a slot before it takes a pool picture, so it
never holds more than two pictures; the picture pool grows by four.
score_frames() pops one frame per reader per step and still makes every
vmaf_read_pictures() call in order on the main thread, so the frames
scored, their pairing, the progress line and the ADR-1262 exit codes are
unchanged. The readers stop after --frame_cnt frames or the first frame
that ends or fails their stream; shutdown drains both rings before
joining either reader, which wakes a reader blocked on the pool.
Handles that may share a read position (one file or pipe on both sides
and --no-reference on POSIX, anything but two regular files on Windows)
and a failed thread creation keep the inline read.

JSON is identical to master at --precision max on 22 CPU cases (Netflix
8/10-bit, Y4M, --frame_cnt, --frame_skip_*, --subsample, --threads
0/4/8/16, truncated and short inputs, same-file input, BBB 4K) and 6
SYCL cases on the B580. test_vmaf_frame_readahead pins order, pairing,
limits and exits; it and the CLI runs are clean under ThreadSanitizer.

Closes T-CLI-SYNC-FRAME-READ-FLOOR-2026-09-29. ADR-1366, Research-1366.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Session C measured the read-ahead on the i9-12900K's integrated UHD 770
and CPU motion on 16 threads, master and branch interleaved, 5
repetitions each. The branch was faster in every repetition: CPU
--threads 16 motion 4.72 to 2.84 ms per 4K frame, UHD 770 psnr 24.84 to
20.63, motion 13.07 to 9.65, default model 70.67 to 67.29.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the perf/cli-frame-readahead branch from f401ee8 to 2a90e49 Compare September 30, 2026 10:55
@lusoris
lusoris merged commit 49de09b into master Sep 30, 2026
106 checks passed
@lusoris
lusoris deleted the perf/cli-frame-readahead branch September 30, 2026 11:40
lusoris added a commit that referenced this pull request Oct 8, 2026
…9cb9479f2

The fork's records of its Netflix/vmaf pull requests and of the upstream
defects it tracks were last checked on 2026-10-01. Upstream master has
moved to 9cb9479f2 since, with the MSVC series, the arm64 ADM kernels,
the fused SpEED filter and a rewrite of integer_compute_adm(). This
brings the records up to date. No fork code and no score changes.

- known-upstream-bugs.md: 42 open pull requests instead of 15, which of
  them needed a rebase or a rework against 9cb9479f2, the #1494 and
  best15 status, the integer AIM change upstream (cffd5b77d), the arm64
  gap in #1602, and what is new upstream since the parity pin. The pin
  heading is unchanged.
- state.md: dated updates on the rows for #1422, #1551 (both halves),
  #1602, #1605, #1635, #1636 and #955.
- upstream_parity.d: the APSNR zero-error row also names #1618, which
  changes the same cap upstream; the allowlist page is regenerated.
- rebase-notes.d: what a port of upstream's arm64 ADM code (8bc5a5c6a,
  b41d2340a) must not copy, and the closed state of #1551.

Signed-off-by: Lusoris <lusoris@proton.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:perf Performance improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant