Summary
Stock PTOAS 0.65 and 0.66 serialize a vector operation behind an unrelated MTE2 load when two provably disjoint UB rows are selected through dynamic pto.subview offsets. Using constant row offsets preserves the intended independence.
This is a conservative synchronization / missed-optimization report. The reproducer does not demonstrate incorrect numerical results. It uses only standard PTO operations, explicit UB addresses, and level3; no memory-planning pass or cross-core communication is needed.
Related: #536 — subview alias provenance and correctness; this report concerns precision after preserving that provenance.
Related: #1569 — redundant multi-buffer synchronization and event exhaustion; this reproducer exposes an earlier alias-precision loss without a loop or event exhaustion.
Command line
Save the input below as dynamic_disjoint.pto:
ptoas --version
ptoas --pto-level=level3 --enable-insert-sync dynamic_disjoint.pto -o dynamic_disjoint.cpp
ptoas --pto-level=level3 --enable-insert-sync --emit-pto-ir --pto-insert-sync-debug=3 dynamic_disjoint.pto -o dynamic_disjoint.sync.pto 2> dynamic_disjoint.sync.log
Repeat the first compile with the static and same-slot controls described below.
Reproduction input
The parent allocation contains three non-overlapping rows of 64 FP32 elements: 256 bytes per slot, 768 bytes total. The modulo-64 normalization bounds ordinal + 2 to at most 65, so the disjointness proof does not assume absence of integer wraparound.
!ring = !pto.tile_buf<loc=vec, dtype=f32, rows=3, cols=64, v_row=3, v_col=64, blayout=row_major, slayout=none_box, fractal=512, pad=0>
!slot = !pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=64, v_row=1, v_col=64, blayout=row_major, slayout=none_box, fractal=512, pad=0>
!gm = !pto.partition_tensor_view<1x64xf32>
module {
func.func @probe(%i: index, %g0: !gm, %g1: !gm, %out0: !gm, %out1: !gm) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
%base = arith.constant 0 : i64
%c0 = arith.constant 0 : index
%c2 = arith.constant 2 : index
%c3 = arith.constant 3 : index
%c64 = arith.constant 64 : index
%one = arith.constant 1.0 : f32
%ordinal = arith.remui %i, %c64 : index
%cur = arith.remui %ordinal, %c3 : index
%future = arith.addi %ordinal, %c2 : index
%next = arith.remui %future, %c3 : index
%parent = pto.alloc_tile addr = %base : !ring
%current = pto.subview %parent[%cur, %c0] sizes [1, 64] : !ring -> !slot
%future_tile = pto.subview %parent[%next, %c0] sizes [1, 64] : !ring -> !slot
pto.tload ins(%g0 : !gm) outs(%current : !slot)
pto.tload ins(%g1 : !gm) outs(%future_tile : !slot)
pto.tadds ins(%current, %one : !slot, f32) outs(%current : !slot)
pto.tstore ins(%current : !slot) outs(%out0 : !gm)
pto.tstore ins(%future_tile : !slot) outs(%out1 : !gm)
return
}
}
The slot pairs are (0,2), (1,0), or (2,1), so the two views are always disjoint.
Static control: replace only the definitions of %cur and %next with constants 0 and 2.
Same-slot negative control: replace only %next with "arith.remui %ordinal, %c3 : index". The second load then overwrites the data used by the vector operation and must remain a dependency.
All three inputs compile with both tested releases.
Expected performance
The vector operation should depend on the load of its own slot, while the other slot's load may overlap it. A sufficient ordering is shown below with descriptive names; exact event IDs or C++ spelling are not requirements:
TLOAD(current, g0);
set_flag(PIPE_MTE2, PIPE_V, ready_current);
TLOAD(future_tile, g1);
wait_flag(PIPE_MTE2, PIPE_V, ready_current);
TADDS(current, current, one);
The future tile's subsequent store still requires its own load-completion dependency. The same-slot control must continue waiting for the second load.
Actual performance
Both releases place the current-tile readiness signal after the second load:
TLOAD(current, g0);
TLOAD(future_tile, g1);
set_flag(PIPE_MTE2, PIPE_V, ready_current);
wait_flag(PIPE_MTE2, PIPE_V, ready_current);
TADDS(current, current, one);
Consequently, TADDS cannot overlap the second load. In the constant-offset control, the MTE2-to-V signal is emitted between the two loads, and the future tile's store receives a separate MTE2-to-MTE3 dependency. The same-slot control correctly retains the wait for both loads.
| Input |
PTOAS 0.65 |
PTOAS 0.66 |
| Dynamic, disjoint slots |
Current vector operation waits for both loads |
Same |
| Constant slots 0 and 2 |
Current readiness signaled before the second load |
Same |
| Dynamic, same slot |
Required wait for the second load retained |
Same |
Profiling data (optional)
The evidence here is compiled synchronization placement and InsertSync debug output. No device timing or speedup is claimed for this reduced reproducer.
At the tested 0.66 revision:
- PTOIRTranslator.cpp: getPtoSubViewBaseAddresses requires constant offsets. UpdateTileSubViewAliasBufferInfo falls back to UpdateConservativeAliasBufferInfo for dynamic offsets, preserving the whole parent range.
- InsertSyncAnalysis.cpp: forward-dependency filtering already uses compareSlotSSA, but obtains slot expressions through findMultiTileSlotExpr.
- SlotAffineAnalysis.cpp: findMultiTileSlotExpr walks through subviews looking for a multi_tile_get. A subview of alloc_tile has no such ancestor, so it cannot enter that slot-aware path.
Git commit
PTOAS v0.66: 0a3e017.
The 0.65 binary reports "ptoas 0.65"; the 0.66 binary reports "ptoas 0.66".
No compiler patches were applied.
Suggested implementation
- Preserve an internal rotating-slot access descriptor for provable full-row subviews of a dense row-major parent: allocation identity, base/byte range, slot stride, slot count, and slot-index SSA. Start with static geometry, zero column offset, one complete row, and in-range indices. Retain the existing conservative fallback for other layouts or unknown geometry.
- Share this descriptor between multi_tile_get and supported subview accesses. Propagate it through byte-preserving reshape/bitcast aliases. Reuse compareSlotSSA for the relation proof; do not infer the slot count merely from the length of a conservative address list.
- Extend the existing forward-dependency filter to consume that descriptor. Drop a same-iteration edge only when physical-region compatibility and disjoint slots are both proven. Do not make MemoryDependentAnalyzer::MemAlias globally return false for these views: the same views can alias across iterations.
- For rotating loops, reuse the existing slot-keyed event machinery, including producer/consumer slot expressions, event allocation, and boundary handling. With three slots and preload two, COMPUTE(i-1) and LOAD(i+2) use the same physical slot; that reuse dependency must remain. If a store is the last reader, slot release must wait for MTE3. Preserve actual producer readiness, guarded-load behavior, and correct priming/draining for short and empty loops.
- Keep unknown relations and ambiguous view chains conservative. The optimization should work with one loop body and existing public PTO operations.
Acceptance checks: the three inputs above; stages 2/3/4; runtime trip counts 0, 1, S-1, S, S+1 and repeated wraps; guarded future loads; in-place vector chains; MTE3 readers; aliases through reshape/bitcast; and unknown or overlapping subviews. Verify synchronization structure and numerical results without pinning tests to particular event IDs.
Summary
Stock PTOAS 0.65 and 0.66 serialize a vector operation behind an unrelated MTE2 load when two provably disjoint UB rows are selected through dynamic pto.subview offsets. Using constant row offsets preserves the intended independence.
This is a conservative synchronization / missed-optimization report. The reproducer does not demonstrate incorrect numerical results. It uses only standard PTO operations, explicit UB addresses, and level3; no memory-planning pass or cross-core communication is needed.
Related: #536 — subview alias provenance and correctness; this report concerns precision after preserving that provenance.
Related: #1569 — redundant multi-buffer synchronization and event exhaustion; this reproducer exposes an earlier alias-precision loss without a loop or event exhaustion.
Command line
Save the input below as dynamic_disjoint.pto:
ptoas --version ptoas --pto-level=level3 --enable-insert-sync dynamic_disjoint.pto -o dynamic_disjoint.cpp ptoas --pto-level=level3 --enable-insert-sync --emit-pto-ir --pto-insert-sync-debug=3 dynamic_disjoint.pto -o dynamic_disjoint.sync.pto 2> dynamic_disjoint.sync.logRepeat the first compile with the static and same-slot controls described below.
Reproduction input
The parent allocation contains three non-overlapping rows of 64 FP32 elements: 256 bytes per slot, 768 bytes total. The modulo-64 normalization bounds ordinal + 2 to at most 65, so the disjointness proof does not assume absence of integer wraparound.
The slot pairs are (0,2), (1,0), or (2,1), so the two views are always disjoint.
Static control: replace only the definitions of %cur and %next with constants 0 and 2.
Same-slot negative control: replace only %next with "arith.remui %ordinal, %c3 : index". The second load then overwrites the data used by the vector operation and must remain a dependency.
All three inputs compile with both tested releases.
Expected performance
The vector operation should depend on the load of its own slot, while the other slot's load may overlap it. A sufficient ordering is shown below with descriptive names; exact event IDs or C++ spelling are not requirements:
The future tile's subsequent store still requires its own load-completion dependency. The same-slot control must continue waiting for the second load.
Actual performance
Both releases place the current-tile readiness signal after the second load:
Consequently, TADDS cannot overlap the second load. In the constant-offset control, the MTE2-to-V signal is emitted between the two loads, and the future tile's store receives a separate MTE2-to-MTE3 dependency. The same-slot control correctly retains the wait for both loads.
Profiling data (optional)
The evidence here is compiled synchronization placement and InsertSync debug output. No device timing or speedup is claimed for this reduced reproducer.
At the tested 0.66 revision:
Git commit
PTOAS v0.66: 0a3e017.
The 0.65 binary reports "ptoas 0.65"; the 0.66 binary reports "ptoas 0.66".
No compiler patches were applied.
Suggested implementation
Acceptance checks: the three inputs above; stages 2/3/4; runtime trip counts 0, 1, S-1, S, S+1 and repeated wraps; guarded future loads; in-place vector chains; MTE3 readers; aliases through reshape/bitcast; and unknown or overlapping subviews. Verify synchronization structure and numerical results without pinning tests to particular event IDs.