Skip to content

Fast head rotation causes visible splat gaps / stutter in large Gaussian splat scenes (visionOS) #1248

Description

@untoldengine

Summary

On visionOS, large Gaussian splat scenes (tested with the Mip-NeRF 360 "garden" capture,
~2M working-set splats) show visible gaps in the leading edge of a fast head rotation
that slowly fill back in over the following frames, instead of tracking the rotation
cleanly.

Environment

  • Scene: gardenscene.untoldgs, cooked via:
    untoldengine export \
        --input "gardenscene.ply" \
        --output "gardenscene.untoldgs" \
        --splat-up-axis=-y \
        --splat-environment \
        --splat-chunk-splats 4096 \
        --splat-sh-degree 2 \
        --splat-min-opacity 0.005
    
  • GaussianRuntimeLimits.workingSetSplatsOverride = 2_000_000 (raised from a previously
    validated 1.3M baseline on visionOS, which defaults to the 1,000,000 "mobile bucket" cap)
  • Loaded as a single whole-scene entity via plain setEntityGaussian(filename:) (not
    tile-streamed, not progressive-LOD)
  • Device profiling shows the app is CPU-bound during the rotation, not GPU-bound

Symptom

Rotating the headset about the Y axis (left to right) at speed: empty gaps appear on the
leading (far right) edge of the view, and gradually fill in over subsequent frames rather
than the rotation tracking cleanly. Consistent with the CPU/GPU work for a frame not
completing before the next drawable is presented, so the compositor shows a stale
drawable while newer data is still catching up.

Investigation so far

Ruled out: a per-frame stereo pose staleness in the XR frame loop (the Gaussian and
mesh cull read renderInfo.xrEye0/1View, only refreshed as a side effect of renderXR's
per-eye draw loop, one stage after cull/preprocess/sort run). A fix hoisting fresh
per-eye matrices ahead of cull was implemented and reverted — it made FPS measurably
worse, most likely because it also fed fresh view-projections into the general mesh HZB
occlusion test (CullingSystem.swift), which structurally depends on being consistently
one frame behind its own depth pyramid (buildHZBDepthPyramid can only ever reflect last
frame's rendered depth). Decoupling the frustum test's freshness from the occlusion test's
depth broke that internal consistency and likely caused more geometry to pass cull
(over-draw) than before.

Confirmed and fixed: GaussianChunkTreeCull.visibleChunkRanges (the CPU tree
pre-filter over a chunked .untoldgs asset's baked cluster tree) recomputed
maxLogScaleMax — an O(chunk count) scan — on every call, and it's called once per
chunked entity, every frame, unconditionally, in executeGaussianFrustumCulling
(GaussianSystem.swift:382). The function's own doc comment already flagged this as a
known risk ("a caller on a per-frame path should cache it once per load"), but the actual
per-frame caller wasn't doing so. Fixed by caching maxLogScaleMax once at
GaussianChunkTable construction (GaussianChunkLoader.swift) instead of recomputing it
every frame. Landed on bugfix/headset_rotation, all 102 relevant Gaussian tests passing,
builds clean on macOS and visionOS. Not yet validated on-device against the original
symptom.

Open questions / next steps

  • Validate the tree-cull caching fix on-device against the original fast-rotation
    symptom; check FrustumCull cpuEncodeMs in the [Gaussian] log before/after
  • If CPU cost is still high, profile GaussianPageManager.tick() specifically
    (topCandidates, selectVictims) — its own throughput caps
    (maxConcurrentReads = 8, maxCommitsPerTick = 4096, commitBudget = 0.5ms,
    GaussianPagingPolicy.swift) may also explain the gradual "fill back in" shape and
    could need retuning for a 2M-splat working set
  • Re-trace Preprocess/RadixSort cpuEncodeMs (not just FrustumCull) at 2M vs.
    the last validated 1.3M baseline
  • Decide whether workingSetSplatsOverride = 2_000_000 should stay, or whether the
    real fix is splitting the scene into tile-streamed Gaussian entities so distant
    regions are never resident/competing for budget at all (bigger project — no
    existing Gaussian tile-splitting tooling yet)

Related files

  • Sources/UntoldEngine/Systems/GaussianSystem.swift
  • Sources/UntoldEngine/Systems/GaussianChunkCull.swift
  • Sources/UntoldEngine/RuntimeAssets/GaussianChunkLoader.swift
  • Sources/UntoldEngine/RuntimeAssets/GaussianPageManager.swift
  • Sources/UntoldEngine/RuntimeAssets/GaussianPagingPolicy.swift
  • Sources/UntoldEngine/Utils/GaussianRuntimeLimits.swift
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions