Skip to content

fix(feature-cpu): CIEDE 4:2:2 chroma swap (OOB) + cambi init leak + integer_ssim 16bpc overflow (bug-hunt r1+r2) - #1050

Merged
lusoris merged 1 commit into
masterfrom
fix/bughunt-feature-cpu
Jun 27, 2026
Merged

lusoris merged 1 commit into
masterfrom
fix/bughunt-feature-cpu

Conversation

@lusoris

@lusoris lusoris commented Jun 27, 2026 •

Copy link
Copy Markdown
Contributor

What

Confirmed, adversarially-verified fixes from the codebase bug-hunt sweep. See commit body for per-finding detail + verification. Golden-safe: no Netflix assertAlmostEqual modified.

Reproducer / smoke

See commit body — each fix lists its build + test evidence (golden gate / parity / regression tests run locally).

Deliverables (ADR-0108)

  • Research digest — no digest needed: verified bug fixes, root cause in commit body + state.md row.
  • Decision matrix — no alternatives: only-one-way correctness fixes matching the CPU/upstream reference.
  • AGENTS.md invariant note — core/src/feature/AGENTS.md.
  • Reproducer / smoke-test command — see above + commit body.
  • CHANGELOG fragment — changelog.d/fixed/bughunt-feature-cpu-ciede-422-cambi-leak.md.
  • Rebase note — docs/rebase-notes.md entry.
  • state.md — T-BUGHUNT-FEATURE-CPU-2026-06-27 (Recently closed).

I did not modify any Netflix golden assertAlmostEqual value.

no docs needed: bug fixes to existing feature extractors (CIEDE2000 4:2:2 chroma swap, cambi init-OOM leak, integer_ssim 16-bpc samplemax² overflow) — correctness only, no new or changed user-discoverable surface or output schema; ADR-0100 / CLAUDE §12 r10 bug-fix exclusion.

@lusoris
lusoris force-pushed the fix/bughunt-feature-cpu branch 2 times, most recently from c69a11d to dfa581c Compare June 27, 2026 14:45
@lusoris lusoris changed the title fix(ciede,cambi): YUV422P chroma-upsample subsample swap (OOB) + cambi init-OOM leak (bug-hunt) fix(feature-cpu): CIEDE 4:2:2 chroma swap (OOB) + cambi init leak + integer_ssim 16bpc overflow (bug-hunt r1+r2) Jun 27, 2026
@lusoris
lusoris marked this pull request as ready for review June 27, 2026 14:54
@lusoris
lusoris enabled auto-merge (squash) June 27, 2026 14:54
…nteger_ssim 16bpc overflow

Bug-hunt feature-cpu batch (rounds 1+2):
- CIEDE2000 4:2:2 chroma-upsample ss_hor/ss_ver flag swap (heap OOB read
  + wrong scores on YUV422P); fix + load-bearing 422 regression tests.
- cambi init() partial-failure leak: route error paths through
  null-tolerant close_cambi().
- integer_ssim 16-bpc samplemax*samplemax signed-int overflow (round-2
  R2-1): 65535*65535 > INT_MAX wraps to -131071 in int, corrupting the
  SSIM c1/c2 stability constants for 16-bit input and diverging from the
  CUDA/HIP/SYCL twins (which use int64_t/double). Hoist
  const double sm = (double)samplemax. 8/10/12-bpc bit-unchanged
  (product fits int). Adds load-bearing test_ssim_16bit_distorted_in_range
  (0.999992 with fix vs 1.185 out-of-[0,1] with the overflow).

Round-2 R2-8 (cambi histogram alloc int-overflow) was already fixed in
tree (c_values_histograms already casts (size_t)alloc_w) -- skipped.

Golden-safe: 8-bit golden SSIM/CIEDE unchanged; no Netflix assertions touched.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
@lusoris
lusoris force-pushed the fix/bughunt-feature-cpu branch from dfa581c to 15ea919 Compare June 27, 2026 14:55
@lusoris
lusoris merged commit 8af3cf3 into master Jun 27, 2026
70 of 81 checks passed
@lusoris
lusoris deleted the fix/bughunt-feature-cpu branch June 27, 2026 15:39
@lusoris lusoris added this to the 1.0.0 — First release milestone Sep 4, 2026
lusoris added a commit that referenced this pull request Oct 2, 2026
…each with its size and the upstream change that ends it (ADR-1479 to ADR-1486)

The reference for code inherited from Netflix/vmaf is Netflix's source;
a difference needs an ADR. The upstream parity audit of 2026-10-02 found
deliberate differences that had none of their own, or whose ADR
(ADR-1033) names neither upstream's behaviour nor the size:

- ADR-1479 ciede on 4:2:2: chroma flags (fork PR #1050); 0.153 on 48 of
  48 frames; Netflix/vmaf#1611.
- ADR-1480 speed_temporal buffers at speed_prescale above 1 (#1643);
  up to 195, upstream segfaults on two fixtures; Netflix/vmaf#1627.
- ADR-1481 a failing extractor fails the run (#871); status only, 78
  probe runs where upstream is silent and 88 where it crashes.
- ADR-1482 integer adm on frames of 17 to 32 pixels (#1473, #1507);
  scale 3 up to 0.23; Netflix/vmaf#1599, #1600.
- ADR-1483 odd-sized chroma planes round up (4f08d32); psnr_cb / cr
  up to 0.684 / 0.826 dB, ciede 0.198.
- ADR-1484 float_ms_ssim magnitude before pow() (#641, ADR-1033 item 2);
  NaN upstream on the 10 px checkerboard; Netflix/vmaf#1665.
- ADR-1485 apsnr of a plane without error (#641, item 1); 114 against
  60 dB; Netflix/vmaf#1666.
- ADR-1486 float_motion scale-1 stride (#641, item 9); up to 25.1;
  Netflix/vmaf#1667.

Each ADR gives upstream's file and line at Netflix 9e48141b, the fork's
lines, the reason found in the fork's pull request, commit or code, and
the measured size from the audit. Documentation only.
lusoris added a commit that referenced this pull request Oct 2, 2026
… and mirror them in the CUDA, SYCL and HIP twins (ADR-1476) (#1892)

* fix(ciede): form ciede2000()'s two products in float as upstream does and mirror them in the CUDA, SYCL and HIP twins (ADR-1476)

Netflix's ciede2000() writes sqrt(c_prime_1 * c_prime_2) and
+ r_sub_t * chroma * hue with float operands, so both products are
rounded to float before the double expression widens them
(libvmaf/src/feature/ciede.c:224-225, :235-236 at 9e48141b). PR #552, a
CodeQL sweep, cast the first operand of each to double, which keeps the
products exact and changes the score. The upstream parity audit of
2026-10-02 found it (U4); the golden gate does not see it.

ciede.c carries upstream's expressions again; the squares stay products
(ADR-1467). cuda/integer_ciede/ciede_device.h and ciede_ff_math.h (SYCL
and HIP) form the same float products in place of the exact ones.

Against Netflix master (C API, %.17g, 327 frames from 8x8 to 3840x2160,
GCC 16.2.1, glibc 2.44), scalar and default dispatch: 153 frames
identical, 7 before. The rest: 119 frames up to 2.2e-11 are ADR-1467's
product where upstream calls powf(x, 2) (272 of 327 with that call put
back), 48 are 4:2:2 input (chroma flag fix, PR #1050), 7 are odd sizes
(chroma planes rounded up). The CPU score moves by up to 1.3e-9 on
frames of 160x90 and larger.

Twins, 153 frames at --precision max, largest difference from the CPU
before and after: CUDA (RTX 4090) 8.4e-13, 2.1e-12; SYCL (Arc A380)
7.3e-13, 2.1e-12; HIP (gfx1036) 7.3e-13, 2.1e-12; bound 1e-9.

test_ciede_upstream_products replays the CUDA header's ciede_delta_e()
with float and with widened products over 20 000 colour pairs (193 tell
them apart); test_ciede_device_math holds the header to the CPU
extractor bit for bit. The SYCL and CUDA source contracts pin upstream's
statements.
Netflix golden gate: 271 passed, 12 skipped on x86-64 and aarch64.

* docs: regenerate the indexes and the citation map after rebasing
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant