Skip to content

CUDA accelerated PSNR - #1175

Open
gedoensmax wants to merge 5 commits into
Netflix:masterfrom
gedoensmax:psnr
Open

gedoensmax wants to merge 5 commits into
Netflix:masterfrom
gedoensmax:psnr

Conversation

@gedoensmax

Copy link
Copy Markdown
Contributor

The speedup that we see is very significant GPU compared to CPU, this scales well for higher resolutions.
When used with FFmpeg this is especially important as also omits a needed PCI copy when using the hardware decoders. When i find more time i will do the same for SSIM but this is a little more work.

./libvmaf/build/tools/vmaf --reference ../data/reference_1080p_yuv420p.yuv --distorted ../data/distorted_1080p_yuv420p.yuv --width 1920 --height 1080 --pixel_format 420 --bitdepth 8 -o res/test_gpu.json --json --feature psnr_cuda
>>> VMAF version f52a8d72
>>> 128 frames ⠋⠉ 303.32 FPS
>>>  vmaf_v0.6.1: 99.867883

./libvmaf/build/tools/vmaf --reference ../data/reference_1080p_yuv420p.yuv --distorted ../data/distorted_1080p_yuv420p.yuv --width 1920 --height 1080 --pixel_format 420 --bitdepth 8 -o res/test.json --json --feature psnr
>>> VMAF version f52a8d72
>>> 128 frames ⠋⠉ 204.50 FPS
>>> vmaf_v0.6.1: 99.867883

@gedoensmax

Copy link
Copy Markdown
Contributor Author

Oh this will also contribute to ffmpeg as a colleague of mine has been experimenting with 8K footage and saw that there is no GPU accelerated PSNR as of now in ffmpeg. (At least not to our knowledge)

@gedoensmax

Copy link
Copy Markdown
Contributor Author

Based on #1174

@BlueSwordM

Copy link
Copy Markdown

This looks interesting, but this doesn't have a lot of value considering it's still PSNR at the end of the day.

Instead, I believe some focus should be on GPU accelerating much more powerful metrics like butteraugli and ssimulacra2 respectively:
https://github.com/cloudinary/ssimulacra2

@gedoensmax

Copy link
Copy Markdown
Contributor Author

The motivation behind this is to not hold CUDA VMAF backe because of PSNR. If video is decoded accelerated it is already in GPU memory and would have to be downloaded to CPU just to calculate PSNR.

@gedoensmax

Copy link
Copy Markdown
Contributor Author

@kylophone could you give this a review/test ?

@kylophone

Copy link
Copy Markdown
Collaborator

I tested this and there was a speed regression for vmaf only with raw inputs, likely due to the chroma copy.

@gedoensmax

Copy link
Copy Markdown
Contributor Author

Yes that can be true, in ffmpeg that should not be happening. Can you put any numbers behind that speed regression?

@gedoensmax

Copy link
Copy Markdown
Contributor Author

@kylophone any update on this ? As said the big benefit comes from using this with ffmpeg: GPU decode + GPU filter. If PSNR has to be calculated on the CPU the GPU data has to be downloaded and blocks processing a lot.

@gedoensmax

Copy link
Copy Markdown
Contributor Author

@kylophone Do you see the speed regression on the standalone tool as a blocker ? In ffmpeg this would not lead to a compression due to either using HW decode or overlapping with the kernels which the standalone tool cannot do (blocking fread in the main thread).

@lusoris

lusoris commented Oct 1, 2026

Copy link
Copy Markdown

Checked against master 6ec23e8f2 (gcc 16.2.1, CUDA 13.4.92, RTX 4090; built with -Wno-error=incompatible-pointer-types, which that revision needed with GCC 16; it is no longer needed since master fixed the cause in 8e7a1ac4e, the only upstream change since). I could not run the extractor, because the branch is too far behind to build or merge:

  • git merge of the PR head (4e889f95, merge base 4c08f00bd, 330 commits behind master) conflicts in integer_adm_cuda.c, integer_motion_cuda.c, integer_vif_cuda.c and src/meson.build.
  • Built as it is, it fails at configure: src/meson.build:275 and :279 add /usr/local/cuda/include (Include dir /usr/local/cuda/include does not exist here). With that path pointed at /opt/cuda/include, the first failure is in a file the PR does not touch: src/cuda/picture_cuda.c:44 and :71 (CUdeviceptr assigned from void *), :148 and :216 (implicit declaration of vmaf_ref_init / vmaf_ref_load), :226 (cuMemFreeAsync argument). GCC 14+ treats these as errors; that file is different on master.
  • cuda: add psnr_cuda, ssim_cuda and ciede_cuda feature extractors #1563 adds a psnr_cuda extractor too (same name, integer_psnr_cuda.c, integer_psnr/psnr.cu). git merge-tree of the two PRs conflicts in 8 files, so only one of them can go in as it is.

On the standalone-CLI speed regression reported in July 2023: on master translate_picture_host() uploads plane mask 0x1, luma only (libvmaf.c:757), and translate_picture_device() downloads the same (:787). This PR changes both to 0xF for every CUDA run, so the CLI uploads the chroma planes whether or not any extractor reads them. #1563 and #1574 make the upload mask depend on a flag the extractor sets (VMAF_FEATURE_EXTRACTOR_CUDA_CHROMA and VMAF_FEATURE_EXTRACTOR_CHROMA, both 1 << 4 in feature_extractor.h, so those two conflict with each other as well).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants