Summary
On visionOS, large Gaussian splat scenes (tested with the Mip-NeRF 360 "garden" capture,
~2M working-set splats) show visible gaps in the leading edge of a fast head rotation
that slowly fill back in over the following frames, instead of tracking the rotation
cleanly.
Environment
- Scene:
gardenscene.untoldgs, cooked via:
untoldengine export \
--input "gardenscene.ply" \
--output "gardenscene.untoldgs" \
--splat-up-axis=-y \
--splat-environment \
--splat-chunk-splats 4096 \
--splat-sh-degree 2 \
--splat-min-opacity 0.005
GaussianRuntimeLimits.workingSetSplatsOverride = 2_000_000 (raised from a previously
validated 1.3M baseline on visionOS, which defaults to the 1,000,000 "mobile bucket" cap)
- Loaded as a single whole-scene entity via plain
setEntityGaussian(filename:) (not
tile-streamed, not progressive-LOD)
- Device profiling shows the app is CPU-bound during the rotation, not GPU-bound
Symptom
Rotating the headset about the Y axis (left to right) at speed: empty gaps appear on the
leading (far right) edge of the view, and gradually fill in over subsequent frames rather
than the rotation tracking cleanly. Consistent with the CPU/GPU work for a frame not
completing before the next drawable is presented, so the compositor shows a stale
drawable while newer data is still catching up.
Investigation so far
Ruled out: a per-frame stereo pose staleness in the XR frame loop (the Gaussian and
mesh cull read renderInfo.xrEye0/1View, only refreshed as a side effect of renderXR's
per-eye draw loop, one stage after cull/preprocess/sort run). A fix hoisting fresh
per-eye matrices ahead of cull was implemented and reverted — it made FPS measurably
worse, most likely because it also fed fresh view-projections into the general mesh HZB
occlusion test (CullingSystem.swift), which structurally depends on being consistently
one frame behind its own depth pyramid (buildHZBDepthPyramid can only ever reflect last
frame's rendered depth). Decoupling the frustum test's freshness from the occlusion test's
depth broke that internal consistency and likely caused more geometry to pass cull
(over-draw) than before.
Confirmed and fixed: GaussianChunkTreeCull.visibleChunkRanges (the CPU tree
pre-filter over a chunked .untoldgs asset's baked cluster tree) recomputed
maxLogScaleMax — an O(chunk count) scan — on every call, and it's called once per
chunked entity, every frame, unconditionally, in executeGaussianFrustumCulling
(GaussianSystem.swift:382). The function's own doc comment already flagged this as a
known risk ("a caller on a per-frame path should cache it once per load"), but the actual
per-frame caller wasn't doing so. Fixed by caching maxLogScaleMax once at
GaussianChunkTable construction (GaussianChunkLoader.swift) instead of recomputing it
every frame. Landed on bugfix/headset_rotation, all 102 relevant Gaussian tests passing,
builds clean on macOS and visionOS. Not yet validated on-device against the original
symptom.
Open questions / next steps
Related files
Sources/UntoldEngine/Systems/GaussianSystem.swift
Sources/UntoldEngine/Systems/GaussianChunkCull.swift
Sources/UntoldEngine/RuntimeAssets/GaussianChunkLoader.swift
Sources/UntoldEngine/RuntimeAssets/GaussianPageManager.swift
Sources/UntoldEngine/RuntimeAssets/GaussianPagingPolicy.swift
Sources/UntoldEngine/Utils/GaussianRuntimeLimits.swift
Summary
On visionOS, large Gaussian splat scenes (tested with the Mip-NeRF 360 "garden" capture,
~2M working-set splats) show visible gaps in the leading edge of a fast head rotation
that slowly fill back in over the following frames, instead of tracking the rotation
cleanly.
Environment
gardenscene.untoldgs, cooked via:GaussianRuntimeLimits.workingSetSplatsOverride = 2_000_000(raised from a previouslyvalidated 1.3M baseline on visionOS, which defaults to the 1,000,000 "mobile bucket" cap)
setEntityGaussian(filename:)(nottile-streamed, not progressive-LOD)
Symptom
Rotating the headset about the Y axis (left to right) at speed: empty gaps appear on the
leading (far right) edge of the view, and gradually fill in over subsequent frames rather
than the rotation tracking cleanly. Consistent with the CPU/GPU work for a frame not
completing before the next drawable is presented, so the compositor shows a stale
drawable while newer data is still catching up.
Investigation so far
Ruled out: a per-frame stereo pose staleness in the XR frame loop (the Gaussian and
mesh cull read
renderInfo.xrEye0/1View, only refreshed as a side effect ofrenderXR'sper-eye draw loop, one stage after cull/preprocess/sort run). A fix hoisting fresh
per-eye matrices ahead of cull was implemented and reverted — it made FPS measurably
worse, most likely because it also fed fresh view-projections into the general mesh HZB
occlusion test (
CullingSystem.swift), which structurally depends on being consistentlyone frame behind its own depth pyramid (
buildHZBDepthPyramidcan only ever reflect lastframe's rendered depth). Decoupling the frustum test's freshness from the occlusion test's
depth broke that internal consistency and likely caused more geometry to pass cull
(over-draw) than before.
Confirmed and fixed:
GaussianChunkTreeCull.visibleChunkRanges(the CPU treepre-filter over a chunked
.untoldgsasset's baked cluster tree) recomputedmaxLogScaleMax— an O(chunk count) scan — on every call, and it's called once perchunked entity, every frame, unconditionally, in
executeGaussianFrustumCulling(
GaussianSystem.swift:382). The function's own doc comment already flagged this as aknown risk ("a caller on a per-frame path should cache it once per load"), but the actual
per-frame caller wasn't doing so. Fixed by caching
maxLogScaleMaxonce atGaussianChunkTableconstruction (GaussianChunkLoader.swift) instead of recomputing itevery frame. Landed on
bugfix/headset_rotation, all 102 relevant Gaussian tests passing,builds clean on macOS and visionOS. Not yet validated on-device against the original
symptom.
Open questions / next steps
symptom; check
FrustumCull cpuEncodeMsin the[Gaussian]log before/afterGaussianPageManager.tick()specifically(
topCandidates,selectVictims) — its own throughput caps(
maxConcurrentReads = 8,maxCommitsPerTick = 4096,commitBudget = 0.5ms,GaussianPagingPolicy.swift) may also explain the gradual "fill back in" shape andcould need retuning for a 2M-splat working set
Preprocess/RadixSortcpuEncodeMs(not justFrustumCull) at 2M vs.the last validated 1.3M baseline
workingSetSplatsOverride = 2_000_000should stay, or whether thereal fix is splitting the scene into tile-streamed Gaussian entities so distant
regions are never resident/competing for budget at all (bigger project — no
existing Gaussian tile-splitting tooling yet)
Related files
Sources/UntoldEngine/Systems/GaussianSystem.swiftSources/UntoldEngine/Systems/GaussianChunkCull.swiftSources/UntoldEngine/RuntimeAssets/GaussianChunkLoader.swiftSources/UntoldEngine/RuntimeAssets/GaussianPageManager.swiftSources/UntoldEngine/RuntimeAssets/GaussianPagingPolicy.swiftSources/UntoldEngine/Utils/GaussianRuntimeLimits.swift