Skip to content

fix: SG-43367: eliminate EXR+WAV playback stuttering - #1350

Draft
bernie-laberge wants to merge 2 commits into
AcademySoftwareFoundation:mainfrom
bernie-laberge:fix_playback_stuttering
Draft

bernie-laberge wants to merge 2 commits into
AcademySoftwareFoundation:mainfrom
bernie-laberge:fix_playback_stuttering

Conversation

@bernie-laberge

@bernie-laberge bernie-laberge commented Jul 16, 2026 •

Copy link
Copy Markdown
Contributor

fix: SG-43367: eliminate EXR+WAV playback stuttering

Linked issues

NA

Summarize your change.

Fixes stuttering during playback of high-resolution multi-part EXR sequences with WAV/MOV audio on Linux. The image was correct but playback dropped well below the target frame rate even with frames fully resident in cache. Root causes were in the GPU texture-upload path, the EXR decode scheduling, and an OpenEXR 3.3 I/O regression. Playback now holds 24 fps with ~100% of large frames on the DMA fast path and per-frame present stalls effectively gone.

GPU texture upload (primary fix) -- src/lib/ip/IPCore/ImageRenderer.{cpp,h}:

  • Make the staging PBO transient. It was allocated in initializeTexture() and pinned to the cached TextureDescription for the texture's entire cache life (~10+ frames), so with a large look-ahead cache the small fixed PBO pool was permanently occupied and most large frames fell back to the slow synchronous glTexImage2D-from-client-memory path (~72 ms, ~1.4 GB/s). The PBO is now acquired inside uploadPlane() for a single upload and released back to the pool immediately after the transfer is issued, so a handful of buffers serve an unbounded stream. Result: PBO fast-path use went ~64% -> 99.9% and CPU submit ~72 ms -> ~14 ms (~6.8 GB/s) for 4546x2864 frames. Since a PBO no longer lives as long as its texture, the per-resident-texture cap (RV_RENDERING_MAX_CONCURRENT_PBOS) is obsolete and is removed; the pool's own soft limits and free-memory check bound the buffers (verified on a 24-source session: the pool stayed at its 8 pre-allocated buffers with one in use at a time).
  • Upload 3-channel RGB half/float frames via a native RGBA path: allocate the texture as GL_RGBA16F/RGBA32F and expand RGB->RGBA on the CPU during upload (GL_TEXTURE_RECTANGLE only). GL_RGB16F/RGB32F have no native DMA path; the driver was expanding per-pixel on upload, which is far slower.
  • Fix a texture cross-reuse corruption: an expanded 3ch->RGBA texture and a genuine 4ch RGBA texture produce identical destination geometry/format, so the getTexture() Compatible-reuse path could recycle one for the other and run the wrong upload path (wrong channel stride), causing image artifacts. compatible() now also compares expandRGBToRGBA so the two are never reused across each other.

PBO pool sizing -- src/lib/graphics/TwkGLF/GLPixelBufferObjectPool.cpp:

  • Raise the fixed-size PBO size constants from x6 to x8 bytes/px (MAX_SIZE 4032*4536*8, MIN_SIZE 1920*1080*8). Both express resolution x bytes-per-pixel, and the RGB->RGBA expansion above changes the uploaded format from 3-channel half (6 B/px) to 4-channel half (8 B/px). Left at 6, MAX_SIZE gave an effective ceiling of 13.7 Mpx at RGBA16 -- only 5.1% above the 4546x2864 frames in the repro -- and a frame over the max makes pop() return NULL, silently falling back to the slow non-PBO upload, i.e. reinstating the very stall the transient-PBO change removes. Measured on 4546x3400 (117.9 MiB) frames, same media and window: old ceiling 0.0% PBO / 1.97 GB/s / 414 uploads; new ceiling 99.9% PBO / 5.23 GB/s / 739 uploads. The pool pre-allocates ceil(7*1.05)=8 buffers of MAX_SIZE, so the pre-allocated footprint goes 837 -> 1116 MiB. MIN_SIZE is consistency only: it is read solely inside initPBOPool() branches gated on a poolUpperLimit that is always assigned poolSize, so nothing reachable consumes it today.

RGB->RGBA expanders -- src/lib/image/TwkFB/FastMemcpy.{cpp,h}:

  • Add parallel (task-pool) expand_rgb_to_rgba_16bit/32bit(_MP) helpers used by the upload path.

EXR decode scheduling -- src/lib/ip/IPBaseNodes/StackIPNode.cpp, src/lib/ip/IPBaseNodes/FileSourceIPNode.cpp:

  • RV forces single-threaded block caching for a frame when testEvaluate() reports slow-random-access video (e.g. a long-GOP MOV). testEvaluate() also visited inputs and media that evaluate() never decodes, so a hidden slow MOV serialized fast EXR decode onto a single caching thread.
  • StackIPNode: in "topmost" mode evaluate() renders only the first in-range input, but testEvaluate() tested every input. It now applies the same input selection as evaluate() (topmost, dissolve, strict frame ranges) and passes each input its mapped frame. This is the escalation layout: EXR shots stacked over the edit MOV.
  • FileSourceIPNode: test only the media that supplies the image, not e.g. a MOV attached for its audio track.
  • Frames that really decode slow media are still serialized as before.

EXR decode thread count -- src/lib/app/RvApp/Options.{cpp,h}, src/bin/apps/rv/main.cpp, src/bin/nsapps/RV/main.cpp, src/lib/app/RvCommon/RvPreferences.cpp (+ user manual):

  • The Automatic EXR thread count was logicalCores - 1, which oversizes the single OpenEXR pool shared by all decodes on high-core machines and starves the UI, audio and caching threads. New Rv::automaticExrThreadCount(): cores - 1 up to 16 cores (unchanged), then 15 + (cores - 16) / 4 (24 -> 17, 64 -> 27, 128 -> 43). On a 64-logical-core system, 12-32 threads played best and 56-64 dropped frames. Overridable with RV_EXR_AUTO_MAX_THREADS; manual counts and rvio are unchanged.

OpenEXR 3.3 I/O regression -- src/lib/image/IOexr/FileStreamIStream.{cpp,h}:

  • OpenEXR 3.3.x is much slower for custom Imf::IStream subclasses that do not implement size()/stateless read(). RV's default memory-mapped I/O uses such a stream. Implement size(), isStatelessRead() and the stateless read() overload (guarded by IOEXR_HAS_STATELESS_ISTREAM) so 3.3+ takes the fast, concurrent read path from RV's mapped buffers.

Playback diagnostics (opt-in, env-gated) -- src/lib/base/TwkUtil/ PlaybackDiagnostics.{cpp,h} (+ CMakeLists), and instrumentation in Session.{cpp,h}, GLWindow.cpp, FBCache.{cpp,h}, IPGraph.cpp, FileSourceIPNode.cpp, ImageRenderer.cpp, ALSASafeAudioModule/ALSASafeAudioRenderer.cpp (+ CMakeLists):

  • Add a thread-safe, RV_PLAYBACK_DIAG-gated CSV logger capturing background decode times/concurrency/compression, audio cache-miss/underrun events, buffering pauses and frame skips, display pacing (interval/dframe/refreshes), cache hit/miss and runway, per-plane GPU upload cost/PBO usage, and a stall anatomy split (present/swap vs in-graph work). All timing is behind the env gate and has no cost when disabled.
  • tools/analyze_playback_diag.py: analyzer that summarizes the log and reports a bottleneck verdict (decode vs audio vs cache vs present).

Describe the reason for the change.

Playback of high-resolution multi-part EXR sequences with WAV/MOV audio on Linux could potentially stutter

Describe what you have tested and on which operating system.

Successfully tested on Rocky Linux 9

Add a list of changes, and note any that might need special attention during the review.

If possible, provide screenshots.

@bernie-laberge
bernie-laberge force-pushed the fix_playback_stuttering branch 3 times, most recently from 0048fe4 to 94521f7 Compare August 28, 2026 19:21
@KevinJW

KevinJW commented Aug 30, 2026

Copy link
Copy Markdown

Not wanting to derail anything here, but something that used to be beneficial in the old days of 10bit DPX, was performing the GPU upload with the 'wrong' pixel type and then running a shader to fix up the pixels after it was uploaded. In the old case you upload 10bit RGB packed (32 bits) as 8 bit RGBA (same 32 bits size) and then do the unpack on the GPU. This was a similar case where the 8bit RGBA upload case was optimised.

In this case the EXR originating as 16 bit RGB but pretending to be 16bit RGBA the bandwidth used to to the upload would be 25% less at the cost of a more complex RGB -> RGBA buffer conversion on the GPU. though I'm guessing we are not hitting any bandwidth limits on PCIe4 or later devices, even if you account for reading things back to feed to SDI presentation devices.

It was a problem on older gen 3 PCIe GFX cards at these kinds of resolutions.

Comment thread src/lib/app/RvApp/Options.cpp Outdated
// numCPUs() is the logical count, so on hyper-threaded systems this is
// roughly the physical core count. Smaller machines keep the previous
// behavior (all but one core).
if (cores > 16)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if there is a more graceful way of scaling these rather than having a potentially confusing result when you move between a 16 core and say an 18 core processor in a machine.

I'm avoiding the obvious issues of what hyperthreading contributes to the decode performance, or the different between P and E cores, multiple CPU packages in a machine etc..

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I totally agree @KevinJW, I don't like this new heuristic in its current form. The only thing I know for sure is that oversubscription leads to poorer performances here. I am currently running some performance tests using various decoding threads on my 2-socket 64 cores system. I will publish the results here once I am done. Thanks !

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @KevinJW, good point. The step from 15 threads at 16 cores to 9 at 18 was confusing. I replaced it in the latest push with a continuous, monotonic curve:

  • up to 16 logical cores: cores - 1 (unchanged)
  • above 16: 15 + (cores - 16) / 4 -> RV uses 15 threads plus one more for every 4 additional logical processors

So moving to a machine with more cores never lowers the count.
On my 64-logical-core system, 12-32 threads played back best, while 56-64 threads dropped frames.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This results in quite similar numbers to how we have historically configured our dedicated RV playback machines if not exactly, then directionally - so plus one from me.

@bernie-laberge
bernie-laberge force-pushed the fix_playback_stuttering branch 3 times, most recently from 5d89579 to c446832 Compare October 7, 2026 00:30
bernie-laberge and others added 2 commits October 6, 2026 23:22
… OpenEXR)

Fixes stuttering during playback of high-resolution multi-part EXR sequences
with WAV/MOV audio on Linux. The image was correct but playback dropped well
below the target frame rate even with frames fully resident in cache. Root
causes were in the GPU texture-upload path, the EXR decode scheduling, and an
OpenEXR 3.3 I/O regression. Playback now holds 24 fps with ~100% of large
frames on the DMA fast path and per-frame present stalls effectively gone.

GPU texture upload (primary fix) -- src/lib/ip/IPCore/ImageRenderer.{cpp,h}:
- Make the staging PBO transient. It was allocated in initializeTexture() and
  pinned to the cached TextureDescription for the texture's entire cache life
  (~10+ frames), so with a large look-ahead cache the small fixed PBO pool was
  permanently occupied and most large frames fell back to the slow synchronous
  glTexImage2D-from-client-memory path (~72 ms, ~1.4 GB/s). The PBO is now
  acquired inside uploadPlane() for a single upload and released back to the
  pool immediately after the transfer is issued, so a handful of buffers serve
  an unbounded stream. Result: PBO fast-path use went ~64% -> 99.9% and CPU
  submit ~72 ms -> ~14 ms (~6.8 GB/s) for 4546x2864 frames.
  Since a PBO no longer lives as long as its texture, the per-resident-texture
  cap (RV_RENDERING_MAX_CONCURRENT_PBOS) is obsolete and is removed; the
  pool's own soft limits and free-memory check bound the buffers.
- Upload 3-channel RGB half/float frames via a native RGBA path: allocate the
  texture as GL_RGBA16F/RGBA32F and expand RGB->RGBA on the CPU during upload
  (GL_TEXTURE_RECTANGLE only). GL_RGB16F/RGB32F have no native DMA path; the
  driver was expanding per-pixel on upload, which is far slower.
- Fix a texture cross-reuse corruption: an expanded 3ch->RGBA texture and a
  genuine 4ch RGBA texture produce identical destination geometry/format, so
  the getTexture() Compatible-reuse path could recycle one for the other and
  run the wrong upload path (wrong channel stride), causing image artifacts.
  compatible() now also compares expandRGBToRGBA so the two are never reused
  across each other.

RGB->RGBA expanders -- src/lib/image/TwkFB/FastMemcpy.{cpp,h}:
- Add parallel (task-pool) expand_rgb_to_rgba_16bit/32bit(_MP) helpers used by
  the upload path.

EXR decode scheduling -- src/lib/ip/IPBaseNodes/StackIPNode.cpp,
src/lib/ip/IPBaseNodes/FileSourceIPNode.cpp:
- RV forces single-threaded block caching for a frame when testEvaluate()
  reports slow-random-access video (e.g. a long-GOP MOV). testEvaluate() also
  visited inputs and media that evaluate() never decodes, so a hidden slow MOV
  serialized fast EXR decode onto a single caching thread.
- StackIPNode: in "topmost" mode evaluate() renders only the first in-range
  input, but testEvaluate() tested every input. It now applies the same input
  selection as evaluate() (topmost, dissolve, strict frame ranges) and passes
  each input its mapped frame (it previously passed the stack's context). This
  is the escalation layout: EXR shots stacked over the edit MOV.
- FileSourceIPNode: test only the media that supplies the image, not e.g. a MOV
  attached to the source for its audio track.
- Frames that really decode slow media are still serialized as before.

OpenEXR 3.3 I/O regression -- src/lib/image/IOexr/FileStreamIStream.{cpp,h}:
- OpenEXR 3.3.x is much slower for custom Imf::IStream subclasses that do not
  implement size()/stateless read(). RV's default memory-mapped I/O uses such
  a stream. Implement size(), isStatelessRead() and the stateless read()
  overload (guarded by IOEXR_HAS_STATELESS_ISTREAM) so 3.3+ takes the fast,
  concurrent read path from RV's mapped buffers.

Playback diagnostics (opt-in, env-gated) -- src/lib/base/TwkUtil/
PlaybackDiagnostics.{cpp,h} (+ CMakeLists), and instrumentation in
Session.{cpp,h}, GLWindow.cpp, FBCache.{cpp,h},
IPGraph.cpp, FileSourceIPNode.cpp, ImageRenderer.cpp,
ALSASafeAudioModule/ALSASafeAudioRenderer.cpp (+ CMakeLists):
- Add a thread-safe, RV_PLAYBACK_DIAG-gated CSV logger capturing background
  decode times/concurrency/compression, audio cache-miss/underrun events,
  buffering pauses and frame skips, display pacing (interval/dframe/refreshes),
  cache hit/miss and runway, per-plane GPU upload cost/PBO usage, and a stall
  anatomy split (present/composite/swap vs in-graph work). All timing is behind
  the env gate and has no cost when disabled.
- tools/analyze_playback_diag.py: analyzer that summarizes the log and reports
  a bottleneck verdict (decode vs audio vs cache vs present).

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Bernard Laberge <bernard.laberge@autodesk.com>
When the OpenEXR "Automatic Threads" preference is on (exrcpus == 0), RV set
the shared OpenEXR global thread pool to numCPUs-1 (e.g. 63 on a 64-core box).
Every EXR decode draws from that single global pool, and the decode work
needed for playback is bounded by the media (resolution, frame rate,
compression), not by the core count. On high-core machines the oversized pool
only adds contention with RV's playback/UI, audio and caching threads and
causes dropped frames -- worst with DWA/DWAB-compressed EXRs.

Add Rv::automaticExrThreadCount() and use it wherever the Automatic value is
applied. It returns a continuous, monotonic curve of the logical core count:
- logicalCores <= 1:  1
- logicalCores <= 16: logicalCores - 1 (unchanged from before)
- logicalCores  > 16: 15 + (logicalCores - 16) / 4
  e.g. 24 -> 17, 32 -> 19, 48 -> 23, 64 -> 27, 128 -> 43

Moving to a machine with more cores never lowers the count. The curve keeps
growing on larger machines without reaching the oversubscribed range seen in
testing: on a 64-logical-core system (2x16 cores + HT) with 4546x2864
half-float ZIPS/DWAB EXR sequences, 12-32 decode threads gave the best decode
times and steady 24 fps, while 40+ threads slowed decode and 56-64 threads
dropped frames.

The value can be overridden at runtime with the RV_EXR_AUTO_MAX_THREADS
environment variable. Manual thread counts (exrcpus > 0) and rvio (batch, no
interactive UI to starve) are unchanged.

Call sites updated: src/bin/apps/rv/main.cpp, src/bin/nsapps/RV/main.cpp, and
RvPreferences::exrNumThreadsFinished()/exrAutoThreads().

Docs: describe the Automatic heuristic (with example values) and
RV_EXR_AUTO_MAX_THREADS in the RV user manual EXR decoding threads section, and
refresh the -exrcpus entries in the RV and RVIO command-line reference tables.

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Signed-off-by: Bernard Laberge <bernard.laberge@autodesk.com>
@bernie-laberge
bernie-laberge force-pushed the fix_playback_stuttering branch from c446832 to 1f46509 Compare October 7, 2026 04:05
@bernie-laberge

Copy link
Copy Markdown
Contributor Author

About your comment @KevinJW :

Not wanting to derail anything here, but something that used to be beneficial in the old days of 10bit DPX, was performing the GPU upload with the 'wrong' pixel type and then running a shader to fix up the pixels after it was uploaded. In the old case you upload 10bit RGB packed (32 bits) as 8 bit RGBA (same 32 bits size) and then do the unpack on the GPU. This was a similar case where the 8bit RGBA upload case was optimised.

In this case the EXR originating as 16 bit RGB but pretending to be 16bit RGBA the bandwidth used to to the upload would be 25% less at the cost of a more complex RGB -> RGBA buffer conversion on the GPU. though I'm guessing we are not hitting any bandwidth limits on PCIe4 or later devices, even if you account for reading things back to feed to SDI presentation devices.

It was a problem on older gen 3 PCIe GFX cards at these kinds of resolutions.

Thanks Kevin, that's a neat trick. I'd rather keep this PR simple and avoid a dedicated unpack shader, because upload bandwidth no longer looks like the bottleneck here:

  • On Rocky 9 the 4546x2864 half-float frames now upload through the PBO fast path 99.9% of the time, at about 14 ms per frame (~6.6 GB/s CPU submit), well inside the 41.7 ms budget at 24 fps.
  • The RGB -> RGBA expansion is what puts these uploads on the driver's native fast path; uploading as GL_RGB16F was much slower. On an Apple M1 Pro, turning the expansion off doubled the upload time (25 -> 47 ms) and dropped playback from ~32 to ~21 fps.
  • 6 B/px RGB doesn't map evenly onto RGBA16 texels, so the packed upload would need extra width/edge handling plus a shader unpack in the GL 2.1 path.

If we later see bandwidth limits on older PCIe gen 3 cards or with SDI readback, I'd be happy to revisit it as a separate change.

What would another good investment I think would be to restore the threaded GPU upload functionality that we used to have in RV as it seems to be broken now. I will look into it if I get the chance.

@bernie-laberge
bernie-laberge marked this pull request as ready for review October 7, 2026 04:25
@KevinJW

KevinJW commented Oct 7, 2026

Copy link
Copy Markdown

What would another good investment I think would be to restore the threaded GPU upload functionality that we used to have in RV as it seems to be broken now. I will look into it if I get the chance.

Yes, we definitely saw benefits to that feature in the past, especially when pushing more than 24 FPS, e.g. Stereo, theme park, VR and other higher framerate options.

@bernie-laberge
bernie-laberge marked this pull request as draft October 8, 2026 19:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants