Repository navigation
fix: SG-43367: eliminate EXR+WAV playback stuttering - #1350
bernie-laberge wants to merge 2 commits into
Conversation
0048fe4 to
94521f7
Compare
|
Not wanting to derail anything here, but something that used to be beneficial in the old days of 10bit DPX, was performing the GPU upload with the 'wrong' pixel type and then running a shader to fix up the pixels after it was uploaded. In the old case you upload 10bit RGB packed (32 bits) as 8 bit RGBA (same 32 bits size) and then do the unpack on the GPU. This was a similar case where the 8bit RGBA upload case was optimised. In this case the EXR originating as 16 bit RGB but pretending to be 16bit RGBA the bandwidth used to to the upload would be 25% less at the cost of a more complex RGB -> RGBA buffer conversion on the GPU. though I'm guessing we are not hitting any bandwidth limits on PCIe4 or later devices, even if you account for reading things back to feed to SDI presentation devices. It was a problem on older gen 3 PCIe GFX cards at these kinds of resolutions. |
94521f7 to
14a11e6
Compare
05ef0a5 to
6372f08
Compare
| // numCPUs() is the logical count, so on hyper-threaded systems this is | ||
| // roughly the physical core count. Smaller machines keep the previous | ||
| // behavior (all but one core). | ||
| if (cores > 16) |
There was a problem hiding this comment.
I wonder if there is a more graceful way of scaling these rather than having a potentially confusing result when you move between a 16 core and say an 18 core processor in a machine.
I'm avoiding the obvious issues of what hyperthreading contributes to the decode performance, or the different between P and E cores, multiple CPU packages in a machine etc..
There was a problem hiding this comment.
I totally agree @KevinJW, I don't like this new heuristic in its current form. The only thing I know for sure is that oversubscription leads to poorer performances here. I am currently running some performance tests using various decoding threads on my 2-socket 64 cores system. I will publish the results here once I am done. Thanks !
There was a problem hiding this comment.
Thanks @KevinJW, good point. The step from 15 threads at 16 cores to 9 at 18 was confusing. I replaced it in the latest push with a continuous, monotonic curve:
- up to 16 logical cores: cores - 1 (unchanged)
- above 16: 15 + (cores - 16) / 4 -> RV uses 15 threads plus one more for every 4 additional logical processors
So moving to a machine with more cores never lowers the count.
On my 64-logical-core system, 12-32 threads played back best, while 56-64 threads dropped frames.
There was a problem hiding this comment.
This results in quite similar numbers to how we have historically configured our dedicated RV playback machines if not exactly, then directionally - so plus one from me.
5d89579 to
c446832
Compare
… OpenEXR)
Fixes stuttering during playback of high-resolution multi-part EXR sequences
with WAV/MOV audio on Linux. The image was correct but playback dropped well
below the target frame rate even with frames fully resident in cache. Root
causes were in the GPU texture-upload path, the EXR decode scheduling, and an
OpenEXR 3.3 I/O regression. Playback now holds 24 fps with ~100% of large
frames on the DMA fast path and per-frame present stalls effectively gone.
GPU texture upload (primary fix) -- src/lib/ip/IPCore/ImageRenderer.{cpp,h}:
- Make the staging PBO transient. It was allocated in initializeTexture() and
pinned to the cached TextureDescription for the texture's entire cache life
(~10+ frames), so with a large look-ahead cache the small fixed PBO pool was
permanently occupied and most large frames fell back to the slow synchronous
glTexImage2D-from-client-memory path (~72 ms, ~1.4 GB/s). The PBO is now
acquired inside uploadPlane() for a single upload and released back to the
pool immediately after the transfer is issued, so a handful of buffers serve
an unbounded stream. Result: PBO fast-path use went ~64% -> 99.9% and CPU
submit ~72 ms -> ~14 ms (~6.8 GB/s) for 4546x2864 frames.
Since a PBO no longer lives as long as its texture, the per-resident-texture
cap (RV_RENDERING_MAX_CONCURRENT_PBOS) is obsolete and is removed; the
pool's own soft limits and free-memory check bound the buffers.
- Upload 3-channel RGB half/float frames via a native RGBA path: allocate the
texture as GL_RGBA16F/RGBA32F and expand RGB->RGBA on the CPU during upload
(GL_TEXTURE_RECTANGLE only). GL_RGB16F/RGB32F have no native DMA path; the
driver was expanding per-pixel on upload, which is far slower.
- Fix a texture cross-reuse corruption: an expanded 3ch->RGBA texture and a
genuine 4ch RGBA texture produce identical destination geometry/format, so
the getTexture() Compatible-reuse path could recycle one for the other and
run the wrong upload path (wrong channel stride), causing image artifacts.
compatible() now also compares expandRGBToRGBA so the two are never reused
across each other.
RGB->RGBA expanders -- src/lib/image/TwkFB/FastMemcpy.{cpp,h}:
- Add parallel (task-pool) expand_rgb_to_rgba_16bit/32bit(_MP) helpers used by
the upload path.
EXR decode scheduling -- src/lib/ip/IPBaseNodes/StackIPNode.cpp,
src/lib/ip/IPBaseNodes/FileSourceIPNode.cpp:
- RV forces single-threaded block caching for a frame when testEvaluate()
reports slow-random-access video (e.g. a long-GOP MOV). testEvaluate() also
visited inputs and media that evaluate() never decodes, so a hidden slow MOV
serialized fast EXR decode onto a single caching thread.
- StackIPNode: in "topmost" mode evaluate() renders only the first in-range
input, but testEvaluate() tested every input. It now applies the same input
selection as evaluate() (topmost, dissolve, strict frame ranges) and passes
each input its mapped frame (it previously passed the stack's context). This
is the escalation layout: EXR shots stacked over the edit MOV.
- FileSourceIPNode: test only the media that supplies the image, not e.g. a MOV
attached to the source for its audio track.
- Frames that really decode slow media are still serialized as before.
OpenEXR 3.3 I/O regression -- src/lib/image/IOexr/FileStreamIStream.{cpp,h}:
- OpenEXR 3.3.x is much slower for custom Imf::IStream subclasses that do not
implement size()/stateless read(). RV's default memory-mapped I/O uses such
a stream. Implement size(), isStatelessRead() and the stateless read()
overload (guarded by IOEXR_HAS_STATELESS_ISTREAM) so 3.3+ takes the fast,
concurrent read path from RV's mapped buffers.
Playback diagnostics (opt-in, env-gated) -- src/lib/base/TwkUtil/
PlaybackDiagnostics.{cpp,h} (+ CMakeLists), and instrumentation in
Session.{cpp,h}, GLWindow.cpp, FBCache.{cpp,h},
IPGraph.cpp, FileSourceIPNode.cpp, ImageRenderer.cpp,
ALSASafeAudioModule/ALSASafeAudioRenderer.cpp (+ CMakeLists):
- Add a thread-safe, RV_PLAYBACK_DIAG-gated CSV logger capturing background
decode times/concurrency/compression, audio cache-miss/underrun events,
buffering pauses and frame skips, display pacing (interval/dframe/refreshes),
cache hit/miss and runway, per-plane GPU upload cost/PBO usage, and a stall
anatomy split (present/composite/swap vs in-graph work). All timing is behind
the env gate and has no cost when disabled.
- tools/analyze_playback_diag.py: analyzer that summarizes the log and reports
a bottleneck verdict (decode vs audio vs cache vs present).
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Bernard Laberge <bernard.laberge@autodesk.com>
When the OpenEXR "Automatic Threads" preference is on (exrcpus == 0), RV set the shared OpenEXR global thread pool to numCPUs-1 (e.g. 63 on a 64-core box). Every EXR decode draws from that single global pool, and the decode work needed for playback is bounded by the media (resolution, frame rate, compression), not by the core count. On high-core machines the oversized pool only adds contention with RV's playback/UI, audio and caching threads and causes dropped frames -- worst with DWA/DWAB-compressed EXRs. Add Rv::automaticExrThreadCount() and use it wherever the Automatic value is applied. It returns a continuous, monotonic curve of the logical core count: - logicalCores <= 1: 1 - logicalCores <= 16: logicalCores - 1 (unchanged from before) - logicalCores > 16: 15 + (logicalCores - 16) / 4 e.g. 24 -> 17, 32 -> 19, 48 -> 23, 64 -> 27, 128 -> 43 Moving to a machine with more cores never lowers the count. The curve keeps growing on larger machines without reaching the oversubscribed range seen in testing: on a 64-logical-core system (2x16 cores + HT) with 4546x2864 half-float ZIPS/DWAB EXR sequences, 12-32 decode threads gave the best decode times and steady 24 fps, while 40+ threads slowed decode and 56-64 threads dropped frames. The value can be overridden at runtime with the RV_EXR_AUTO_MAX_THREADS environment variable. Manual thread counts (exrcpus > 0) and rvio (batch, no interactive UI to starve) are unchanged. Call sites updated: src/bin/apps/rv/main.cpp, src/bin/nsapps/RV/main.cpp, and RvPreferences::exrNumThreadsFinished()/exrAutoThreads(). Docs: describe the Automatic heuristic (with example values) and RV_EXR_AUTO_MAX_THREADS in the RV user manual EXR decoding threads section, and refresh the -exrcpus entries in the RV and RVIO command-line reference tables. Co-authored-by: Cursor <cursoragent@cursor.com> Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com> Signed-off-by: Bernard Laberge <bernard.laberge@autodesk.com>
c446832 to
1f46509
Compare
|
About your comment @KevinJW :
Thanks Kevin, that's a neat trick. I'd rather keep this PR simple and avoid a dedicated unpack shader, because upload bandwidth no longer looks like the bottleneck here:
If we later see bandwidth limits on older PCIe gen 3 cards or with SDI readback, I'd be happy to revisit it as a separate change. What would another good investment I think would be to restore the threaded GPU upload functionality that we used to have in RV as it seems to be broken now. I will look into it if I get the chance. |
Yes, we definitely saw benefits to that feature in the past, especially when pushing more than 24 FPS, e.g. Stereo, theme park, VR and other higher framerate options. |
fix: SG-43367: eliminate EXR+WAV playback stuttering
Linked issues
NA
Summarize your change.
Fixes stuttering during playback of high-resolution multi-part EXR sequences with WAV/MOV audio on Linux. The image was correct but playback dropped well below the target frame rate even with frames fully resident in cache. Root causes were in the GPU texture-upload path, the EXR decode scheduling, and an OpenEXR 3.3 I/O regression. Playback now holds 24 fps with ~100% of large frames on the DMA fast path and per-frame present stalls effectively gone.
GPU texture upload (primary fix) -- src/lib/ip/IPCore/ImageRenderer.{cpp,h}:
RV_RENDERING_MAX_CONCURRENT_PBOS) is obsolete and is removed; the pool's own soft limits and free-memory check bound the buffers (verified on a 24-source session: the pool stayed at its 8 pre-allocated buffers with one in use at a time).PBO pool sizing -- src/lib/graphics/TwkGLF/GLPixelBufferObjectPool.cpp:
x6tox8bytes/px (MAX_SIZE4032*4536*8,MIN_SIZE1920*1080*8). Both express resolution x bytes-per-pixel, and the RGB->RGBA expansion above changes the uploaded format from 3-channel half (6 B/px) to 4-channel half (8 B/px). Left at 6,MAX_SIZEgave an effective ceiling of 13.7 Mpx at RGBA16 -- only 5.1% above the 4546x2864 frames in the repro -- and a frame over the max makespop()return NULL, silently falling back to the slow non-PBO upload, i.e. reinstating the very stall the transient-PBO change removes. Measured on 4546x3400 (117.9 MiB) frames, same media and window: old ceiling 0.0% PBO / 1.97 GB/s / 414 uploads; new ceiling 99.9% PBO / 5.23 GB/s / 739 uploads. The pool pre-allocates ceil(7*1.05)=8 buffers ofMAX_SIZE, so the pre-allocated footprint goes 837 -> 1116 MiB.MIN_SIZEis consistency only: it is read solely insideinitPBOPool()branches gated on apoolUpperLimitthat is always assignedpoolSize, so nothing reachable consumes it today.RGB->RGBA expanders -- src/lib/image/TwkFB/FastMemcpy.{cpp,h}:
EXR decode scheduling -- src/lib/ip/IPBaseNodes/StackIPNode.cpp, src/lib/ip/IPBaseNodes/FileSourceIPNode.cpp:
EXR decode thread count -- src/lib/app/RvApp/Options.{cpp,h}, src/bin/apps/rv/main.cpp, src/bin/nsapps/RV/main.cpp, src/lib/app/RvCommon/RvPreferences.cpp (+ user manual):
logicalCores - 1, which oversizes the single OpenEXR pool shared by all decodes on high-core machines and starves the UI, audio and caching threads. NewRv::automaticExrThreadCount():cores - 1up to 16 cores (unchanged), then15 + (cores - 16) / 4(24 -> 17, 64 -> 27, 128 -> 43). On a 64-logical-core system, 12-32 threads played best and 56-64 dropped frames. Overridable withRV_EXR_AUTO_MAX_THREADS; manual counts and rvio are unchanged.OpenEXR 3.3 I/O regression -- src/lib/image/IOexr/FileStreamIStream.{cpp,h}:
Playback diagnostics (opt-in, env-gated) -- src/lib/base/TwkUtil/ PlaybackDiagnostics.{cpp,h} (+ CMakeLists), and instrumentation in Session.{cpp,h}, GLWindow.cpp, FBCache.{cpp,h}, IPGraph.cpp, FileSourceIPNode.cpp, ImageRenderer.cpp, ALSASafeAudioModule/ALSASafeAudioRenderer.cpp (+ CMakeLists):
Describe the reason for the change.
Playback of high-resolution multi-part EXR sequences with WAV/MOV audio on Linux could potentially stutter
Describe what you have tested and on which operating system.
Successfully tested on Rocky Linux 9
Add a list of changes, and note any that might need special attention during the review.
If possible, provide screenshots.