Skip to content

cuda: fix row stride in calculate_motion_score_kernel_16bpc (wrong scores for >8-bit input) - #1552

Merged
kylophone merged 1 commit into
Netflix:masterfrom
BardieJoensen:cuda-fix-motion-16bpc-stride
Jul 31, 2026
Merged

kylophone merged 1 commit into
Netflix:masterfrom
BardieJoensen:cuda-fix-motion-16bpc-stride

Conversation

@BardieJoensen

Copy link
Copy Markdown
Contributor

The CUDA motion feature extractor computes wrong scores for any input with bpc > 8.

In calculate_motion_score_kernel_16bpc the source rows are addressed like this:

blurred_y += filter_d[yf] * (reinterpret_cast<const uint16_t*>(src.data[0])
    + mirror(y-radius+yf, height) * src.stride[0])[mirror(x-radius+xf, width)];

src.stride[0] is in bytes, but the pointer arithmetic happens on a uint16_t*, so the row offset is doubled. Every row reads from row 2*y instead of y, and everything from height/2 down reads past the end of the plane. The 8bpc kernel just above does the same addressing correctly through a uint8_t*.

Whether this crashes or silently corrupts the result depends on what happens to be mapped after the allocation. On my setup (RTX 5060 Ti, driver 595.71.05, reproduced with CUDA 12.9 and 13.3 builds) it usually doesn't crash, which is probably why it hasn't been reported. On a 1080p 10-bit test pair, master @ 0f9912e:

before (GPU) after (GPU) = CPU --gpumask 0
vmaf mean 62.6037 63.5439
integer_motion2 mean 0.7869 1.5742
max per-frame vmaf diff 1.0769 0

compute-sanitizer attributes it deterministically:

========= Invalid __global__ read of size 2 bytes
=========     at calculate_motion_score_kernel_16bpc+0xf90
=========     by thread (0,14,0) in block (6,33,0)
=========     Access to 0x100120380bc is out of bounds
=========     and is 189 bytes after the nearest allocation at 0x10011c00000 of size 4423680 bytes

The numbers check out: the faulting thread is (x,y) = (96,542), whose first filter tap reads row 540. With the doubled stride that's byte offset 540 * 4096 * 2 = 4,423,680 — exactly the size of the 1080-row, 4096-byte-pitch source plane — plus 2 * (96-2) = 188 bytes of column offset.

Repro:

ffmpeg -f lavfi -i testsrc2=size=1920x1080:rate=24:duration=2 -pix_fmt yuv420p10le -strict -1 ref.y4m
ffmpeg -i ref.y4m -vf gblur=sigma=1.5 -pix_fmt yuv420p10le -strict -1 dist.y4m
vmaf -r ref.y4m -d dist.y4m -m version=vmaf_v0.6.1 -o gpu.json --json
vmaf -r ref.y4m -d dist.y4m -m version=vmaf_v0.6.1 -o cpu.json --json --gpumask 0
# per-frame vmaf / integer_motion2 differ between the two

An 8-bit pair produces identical GPU/CPU scores, so only >8bpc input is affected.

With this one-line change the GPU path matches --gpumask 0 exactly on both pairs and the sanitizer is clean.

src.stride[0] is a byte stride, but the 16bpc kernel adds it to a
uint16_t pointer, doubling the effective row offset. Every row reads
from 2*y instead of y, and rows >= height/2 read past the end of the
source plane, so the motion feature is computed from wrong or
out-of-bounds data for any input with bpc > 8. Depending on what is
mapped after the allocation this yields silently wrong scores
(integer_motion2 came out at roughly half its true value on a 1080p10
test pair, dragging fused VMAF down about 1 point) or kills the CUDA
context with CUDA_ERROR_ILLEGAL_ADDRESS.

Do the row arithmetic on a byte pointer before casting, as the 8bpc
kernel already does. With this change the GPU scores match the CPU
path exactly on 8-bit and 10-bit test content, and compute-sanitizer
reports no errors.
@kylophone
kylophone merged commit 4991d2b into Netflix:master Jul 31, 2026
11 checks passed
social4hyq pushed a commit to social4hyq/homebrew-core that referenced this pull request Sep 20, 2026
libvmaf 3.2.1

Created-by: HarmonybrewBot
Commit-by: HarmonybrewBot
Merged-by: HarmonybrewBot
Description: Created by `brew bump`

---

Created with `brew bump-formula-pr`.<details>
  <summary>release notes</summary>
  <pre>Small release with a fix for #1587.

## What's Changed
* doc/vmaf_v1: update Netflix techblog url by @kylophone in Netflix/vmaf#1545
* build(deps): bump actions/checkout from 6 to 7 by @dependabot[bot] in Netflix/vmaf#1546
* build(deps): bump actions/cache from 4 to 6 by @dependabot[bot] in Netflix/vmaf#1550
* libvmaf/thread_pool: fix thread pool queue depth to restore backpressure by @kylophone in Netflix/vmaf#1557
* libvmaf/feature: enable speed_chroma/speed_temporal without enable_float by @kylophone in Netflix/vmaf#1558
* .github/ci: fix release asset name collision across build matrix by @kylophone in Netflix/vmaf#1559
* libvmaf/speed_chroma: remove bilinear prescale index/weight computation from per-pixel loop by @kylophone in Netflix/vmaf#1565
* build(deps): bump actions/setup-python from 6 to 7 by @dependabot[bot] in Netflix/vmaf#1569
* cuda: fix row stride in calculate_motion_score_kernel_16bpc (wrong scores for >8-bit input) by @BardieJoensen in Netflix/vmaf#1552
* Fix race condition with generated header under parallel builds by @StormBytePP in Netflix/vmaf#1578
* Add ARM NEON implementation for 8-bit integer motion feature by @kjg0724 in Netflix/vmaf#1570
* AVX512: Drop -mavx512vbmi from generic AVX-512 compile flags by @StormBytePP in Netflix/vmaf#1577
* Add support for Intel CET (SHSTK) by @StormBytePP in Netflix/vmaf#1575

## New Contributors
* @BardieJoensen made their first contribution in Netflix/vmaf#1552
* @kjg0724 made their first contribution in Netflix/vmaf#1570

**Full Changelog**: https://github.com/Netflix/vmaf/compare/v3.2.0...v3.2.1</pre>
  <p>View the full release notes at <a href="https://github.com/Netflix/vmaf/releases/tag/v3.2.1">https://github.com/Netflix/vmaf/releases/tag/v3.2.1</a>.</p>
</details>
<hr>

See merge request: Harmonybrew/homebrew-core!20180
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants