Skip to content

docs: record the fork's check of the upstream defects verified on 6ec23e8f2 - #1761

Merged
lusoris merged 5 commits into
masterfrom
docs/known-upstream-bugs-2026-10-01
Oct 1, 2026
Merged

lusoris merged 5 commits into
masterfrom
docs/known-upstream-bugs-2026-10-01

Conversation

@lusoris

@lusoris lusoris commented Oct 1, 2026

Copy link
Copy Markdown
Contributor

Summary

Records the result of checking the fork against the fifteen upstream defects verified on Netflix master 6ec23e8f2: a dated section in docs/development/known-upstream-bugs.md (upstream issue, defect, what the fork does and the evidence, or the fork PR that fixed it) and the not-affected results in docs/state.md ("Confirmed not-affected"). No code changes.

Outcome: three reproduced and are fixed (#1305 vif accumulator stream by #1750, #1420 as a hang in vmaf_close() by #1752, #1613 chroma download by #1754), two are documented (#910, #755 and #1180 by #1753), ten are not affected. Also records that Netflix master 8e7a1ac4e (revert of #1476) needs nothing from the fork.

Type

  • docs — documentation only

Deep-dive deliverables (ADR-0108)

  • Research digest — no digest needed: documentation of measurements; the numbers are in the rows.
  • Decision matrix — no alternatives: only-one-way fix (a record of results).
  • AGENTS.md invariant note — no rebase-sensitive invariants: documentation only.
  • Reproducer / smoke-test command — the reproducers are named in the section; the directories are on the host.
  • CHANGELOG fragment — no changelog needed: documents existing behaviour and the fixes carry their own fragments.
  • Rebase note — no rebase impact: documentation only.

Reproducer

~/.cache/vmafx-upstream-rebase/evidence/<issue>/ holds the upstream reproducers; each row says what was run.

Notes

lusoris and others added 5 commits October 1, 2026 22:43
…it-identical (#1734)

* fix(cuda): compute float_adm in the CPU's arithmetic so the twin is bit-identical

float_adm_cuda matched the CPU extractor on 144 of 791 measured scores
and was up to 1.3e-5 from it (adm_scale0 at 3840x2160). An exact twin
was built first and the old constructs were put back one at a time; all
of them together give the old twin's output bit for bit. Sizes alone:

- the angle test's threshold as cos^2 * (|o|^2 * |t|^2), where
  adm_angle_flag_s() evaluates (cos^2 * |o|^2) * |t|^2 (1.3e-5);
- each row reduced per tile and the rows added in fp64, where
  adm_csf_den_scale_s() and adm_cm_s() add a row into one float and the
  rows into another (4.4e-7);
- CSF weights from a host copy of dwt_quant_step() with fp32
  intermediates, four of eight default weights 1 to 3 ulp off (2.1e-7);
- t / (o + eps), where the x86 CPU multiplies by rcp_s(), a Newton step
  on the processor's RCPSS estimate (1.3e-7);
- the masking threshold's centre taps after its neighbours, where
  adm_cm_thresh3x3_s() adds the centre fifth in one sum per band
  (9.4e-8);
- fp32 1/15 and 1/30 where the CPU's are double literals (7.2e-8,
  1.5e-10), and an fp32 gain limit (1.0e-7 at 1.2);
- the frame sums floored at 1e-2 * area / 1080p where compute_adm()
  uses 1e-10: adm2 = 1 instead of the CPU's 0 on a flat 16-bit frame
  with one raised sample and adm_noise_weight=0.

ADR-1420:

- cuda/float_adm/float_adm_device.h holds the decouple, the CSF, the
  threshold and the reduction terms of adm_tools.c operation for
  operation, with explicit round-to-nearest intrinsics on the device;
  the host compiles it for a test.
- RCPSS is specified by an error bound, so the reference's quotient
  belongs to the host processor. adm_reciprocal_model_probe() fills a
  4096-entry table from the instruction and proves the model against it
  (9.4 ms at extractor start); the decouple kernel evaluates it in
  integer arithmetic. A CPU build that divides (MSVC, ARM) makes the
  twin divide.
- float_adm_terms stores nine terms per sample of the reduced region,
  float_adm_row_sums adds each row in one thread, the host adds the
  rows.
- adm_tools.c exports its reduced region, CSF weights, angle constant
  and reciprocal estimate (adm_float_reference.h), and four of its
  reductions call the exported adm_pool_bands_s() instead of repeating
  it. CPU scores are unchanged (4 752 outputs).
- EXACT_TWINS lists float_adm / cuda: the gate cell is an equality.

Measured on an RTX 4090 at --precision max against master 5c8b9e9,
identical scores before and after: Netflix 576x324 8-bit 66/336 and
336/336, 10-, 12- and 16-bit 2/21 and 21/21, checkerboard 1 px 2/21 and
21/21, 10 px 4/21 and 21/21, BBB 3840x2160 66/350 and 350/350; 658 and
2034 of 2034 with debug=true, also with clang's CUDA driver and with
non-default gain limit, bypass, noise weight and viewing geometry. Not
identical: adm_p_norm other than 1 or 3, within 1.1e-7 (device powf).
A run of the twin alone takes 1.87 and 1.98 ms per 4K frame; its
kernels take 0.76 and 1.11 ms
(T-CUDA-FLOAT-ADM-EXACT-THROUGHPUT-2026-10-01).

Found on the way and recorded: the SYCL, HIP and Metal twins carry the
same arithmetic and floor (T-GPU-FLOAT-ADM-CPU-ARITHMETIC-2026-10-01,
T-GPU-FLOAT-ADM-FRAME-SUM-FLOOR-2026-10-01); the CPU's float ADM
depends on the processor through RCPSS
(T-FLOAT-ADM-RECIPROCAL-ESTIMATE-HOST-DEPENDENT-2026-10-01), reads
outside its bands below 17x17
(T-FLOAT-ADM-TINY-FRAME-BAND-READS-2026-10-01) and files its debug
ratio under an unsuffixed key
(T-FLOAT-ADM-DEBUG-KEY-UNSUFFIXED-2026-10-01).

Tests: test_cuda_float_adm_parity asserts equality over 15 cases, each
of which fails on the old twin; test_float_adm_device_math compares the
header with the CPU routines and the reciprocal model with the
instruction on the host; test_cuda_float_adm_exact_contract.py pins the
design.

* docs: regenerate the indexes and the citation map after rebasing
* chore(deps): Update dependency openai to >=3.22.1

* chore(deps): regenerate dev-llm lock for openai >=3.22.1

* docs: regenerate the indexes and the citation map after rebasing
…nd HISS standard (ADR-1142) (#1767)

* test(core): refactor test_barten_csf for standards and HISS compliance

* test(core): refactor test_psnr_hvs_simd for HISS modularity

* test(core): refactor test_propagate_metadata for standards compliance

* test(core): refactor test_thread_pool for standards compliance

* test(core): refactor test_context for standards compliance

* test(core): refactor test_predict for standards compliance

* test(core): refactor test_ciede for standards compliance

* test(core): refactor test_pic_preallocation for standards compliance

* test(core): refactor test_locale_handling for branch and locale hygiene

Add SPDX license identifier, wrap test TU in ADR-0141 and ADR-1138 NOLINT
brackets, extract output verification and tmpfile helpers to satisfy
branch complexity limits, and replace rewind with checked fseek.

* test(core): refactor test_dict for anonymous namespaces and branch hygiene

Add SPDX license identifier, replace NULL with nullptr, move test functions
into discrete anonymous namespaces to stay within HISS-04 60-line block budgets,
and simplify assertions into modular helpers to satisfy branch complexity limits.

* test(core): flatten inner 8x8 block loop in test_psnr_hvs_simd

Flatten nested i and j block loops into a single k loop to remain within
the nesting depth threshold 4 without altering floating-point accumulation.

* chore(ci): tighten CPU tidy and HISS debt baselines for batch A

Record updated debt ratchets for 10 core test TUs in tidy-baseline-cpu.json
and decrement active infractions by 2 in .standards-baseline.json and README.md.
Add changelog fragment for std-core-tests-a and update CHANGELOG.md.

* docs: regenerate the indexes and the citation map after rebasing
…s bit-identical (#1759)

* fix(cuda): return ssimulacra2's sums in the CPU's order so the twin is bit-identical

ssimulacra2_cuda computed the CPU's planes and the CPU's fp64 per-pixel
terms and added the terms in a fixed tree. ssim_map() and
edge_diff_map() add each of a channel's six terms pixel after pixel
into one double; every add rounds, so the order changes the last
digits. 8 of 113 measured frames matched and the rest were up to
7.3e-11 away.

ADR-1433: the device forms the sums of those loops.

- feature/ordered_sum.h: while a running sum of non-negative doubles
  stays in one binade, adding a term adds an integer. A run of terms
  carries two integer increments (even and odd start, because a tie
  rounds to even), runs compose in order, and a walk adds each chunk's
  increment after checking it against the exact sum; a chunk that
  crosses a binade is added term by term. Compiled into the kernels and
  into a host test.
- ssimulacra2_device.cu: chunk_sums, chunk_plan, chunk_units and
  ordered_totals replace combine_partials and combine_final. A chunk is
  1024 pixels in raster order; 9 to 27 of 8100 chunks per sum take the
  term-by-term path on a 4K frame. The readback stays 864 bytes.
- The gate lists ssimulacra2 / cuda as an exact twin.

Measured on an RTX 4090 at --precision max against master 5c8b9e9,
identical frames and largest difference, before and after: Netflix
576x324 8-bit 8/48, 1.3e-13 and 48/48; 10-, 12- and 16-bit 0/3 and 3/3;
checkerboard 1 px 0/3, 3.4e-13 and 3/3; checkerboard 10 px 0/3, 7.3e-11
and 3/3; BBB 3840x2160 0/50, 1.5e-12 and 50/50 (200 frames through the
gate: 0).

It costs time: a run of the twin alone takes 15.6 ms per 3840x2160
frame instead of 7.8 ms and 1.7 ms instead of 0.4 ms at 576x324; the
CPU extractor takes 126 ms per 4K frame on sixteen threads
(T-CUDA-SSIMULACRA2-EXACT-THROUGHPUT-2026-10-01). The SYCL and HIP
twins keep their tree of fp32 pairs
(T-GPU-SSIMULACRA2-SUM-ORDER-2026-10-01).

* docs: regenerate the indexes and the citation map after rebasing
…23e8f2 (#1761)

* docs: record the fork's check of the upstream defects verified on 6ec23e8f2

Fifteen defects reproduced on Netflix master were run against the fork:
three reproduced and are fixed (#1305, #1420 as a hang, #1613), two are
documented (#910, #755 and #1180), ten are not affected. The dated section
in known-upstream-bugs.md and the Confirmed not-affected rows of the state
ledger carry the evidence; Netflix 8e7a1ac4e (revert of #1476) needs
nothing from the fork.

* docs: regenerate the indexes and the citation map after rebasing
@lusoris
lusoris force-pushed the docs/known-upstream-bugs-2026-10-01 branch from 29cf75f to 4a0e1d8 Compare October 1, 2026 20:57
@lusoris
lusoris merged commit 4a0e1d8 into master Oct 1, 2026
58 of 71 checks passed
@lusoris
lusoris deleted the docs/known-upstream-bugs-2026-10-01 branch October 1, 2026 20:57
@github-actions github-actions Bot added the type:docs Documentation updates label Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:docs Documentation updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant