Summary
At --pto-level=level2 --enable-insert-sync, when a loop body performs two
pto.multi_tile_get on the same !pto.multi_tile_buf region, ptoas emits the
per-slot dynamic WAR pair (wait_flag/set_flag on a derived event_t) for the
first access only. The second slot's TLOAD is emitted with no preceding
wait_flag at all, so the next iteration's write into that slot races the
current iteration's read of it.
The result is a silently wrong kernel on device — not a missed optimization.
Related to but distinct from #1106 (that one is a level3 address-folding issue;
this one is a level2 sync-derivation gap).
Environment
- ptoas v0.54 (aarch64),
--enable-insert-sync --pto-level=level2 --pto-arch a3
- Device: Ascend a2a3
--enable-graph-sync-solver behaves the same.
Reproducer
Input .pto (trimmed — one region, two gets per iteration, both slot indices
affine in the loop variable):
%ub_mb = pto.alloc_multi_tile valid_row = %c64_index valid_col = %c64_index
: !pto.multi_tile_buf<!pto.tile_buf<loc=vec, dtype=f32, rows=64, cols=64,
v_row=?, v_col=?, blayout=row_major,
slayout=none_box, fractal=512, pad=0>, count=2>
scf.for %i = %c0_index to %c4_index step %c1_index {
%0 = arith.remsi %i, %c2_index : index // i % 2
%lo = pto.multi_tile_get %ub_mb[%0] : ... -> !pto.tile_buf<loc=vec, ...>
pto.tload ins(%a_view ...) outs(%lo : ...)
%3 = arith.addi %i, %c1_index : index
%4 = arith.remsi %3, %c2_index : index // (i + 1) % 2
%hi = pto.multi_tile_get %ub_mb[%4] : ... -> !pto.tile_buf<loc=vec, ...>
pto.tload ins(%b_view ...) outs(%hi : ...)
%s = pto.tadd ins(%lo, %hi : ...) outs(%s : ...) // both slots read here
pto.tstore ins(%s ...) outs(%out_view ...)
}
Emitted code (v0.54)
// prologue — BOTH event ids are primed, so the region's slot count is understood
set_flag(PIPE_V, PIPE_MTE2, EVENT_ID0);
set_flag(PIPE_V, PIPE_MTE2, EVENT_ID1);
for (int64_t i11 = v8; i11 < v5; i11 += v6) {
int64_t v12 = i11 % v4;
uint64_t v14 = (uint64_t)(v12 == v6 ? v9 : v8);
TASSIGN(v13, v14); // lo -> slot i % 2
int64_t v20 = (int64_t)((uint64_t)v12 % (uint64_t)v4) == v6 ? v6 : v8;
event_t v21 = (event_t) v20; // dynamic id derived from lo's slot
wait_flag(PIPE_V, PIPE_MTE2, v21); // <-- guards lo
TLOAD(v13, v19);
uint64_t v23 = (uint64_t)((int64_t)((uint64_t)i11 + (uint64_t)v6) % v4 == v6 ? v9 : v8);
TASSIGN(v22, v23); // hi -> slot (i + 1) % 2
TLOAD(v22, v26); // <-- NO wait_flag before this load
set_flag(PIPE_MTE2, PIPE_V, EVENT_ID0);
wait_flag(PIPE_MTE2, PIPE_V, EVENT_ID0);
TADD(v27, v13, v22); // reads BOTH slots
event_t v29 = (event_t) v20;
set_flag(PIPE_V, PIPE_MTE2, v29); // <-- releases lo's slot only
...
}
EVENT_ID1 is primed in the prologue and then never waited on or set inside the
loop: only one of the two slots is ever guarded.
Observed failure
out = a + b over 4 row-blocks of [64, 64] fp32, with distinct per-element
inputs. Expected out[k] = a[k] + b[k]; measured
out[k] = a[block i+1][k] + a[block i][k]
i.e. iteration i+1's lo load into slot 1 landed before iteration i's TADD
had read slot 1 as hi. Exactly the missing WAR edge.
What does work
- One
multi_tile_get per iteration (the ordinary ping-pong) is guarded
correctly — dynamic wait_flag/set_flag pair present, numerically correct on
device. This is the shape we ship.
- Straight-line code with two constant slots: no loop, no cross-iteration
reuse, correct.
- An Acc/L0C region with two co-live slots happens to be correct, but not via
per-slot ids — ptoas falls back to a single static wait_flag(PIPE_FIX, PIPE_M, EVENT_ID0) placed after both drains, which serializes conservatively.
Expected behaviour
Every multi_tile_get of a region in an iteration should receive its own
per-slot WAR guard (or, failing that, a conservative guard) — never no guard at
all.
Impact on PyPTO
PyPTO now refuses to lower a slotted allocation with two slots live inside one
loop body under memory_planner=PTOAS, and points the author at the PyPTO memory
planner instead. We would like to lift that restriction once this is fixed; on our
side it is a single condition in PTOCodegen::PlanMultiBufferRegions, with no
user-facing API change.
Summary
At
--pto-level=level2 --enable-insert-sync, when a loop body performs twopto.multi_tile_geton the same!pto.multi_tile_bufregion, ptoas emits theper-slot dynamic WAR pair (
wait_flag/set_flagon a derivedevent_t) for thefirst access only. The second slot's
TLOADis emitted with no precedingwait_flagat all, so the next iteration's write into that slot races thecurrent iteration's read of it.
The result is a silently wrong kernel on device — not a missed optimization.
Related to but distinct from #1106 (that one is a level3 address-folding issue;
this one is a level2 sync-derivation gap).
Environment
--enable-insert-sync --pto-level=level2 --pto-arch a3--enable-graph-sync-solverbehaves the same.Reproducer
Input
.pto(trimmed — one region, two gets per iteration, both slot indicesaffine in the loop variable):
Emitted code (v0.54)
EVENT_ID1is primed in the prologue and then never waited on or set inside theloop: only one of the two slots is ever guarded.
Observed failure
out = a + bover 4 row-blocks of[64, 64]fp32, with distinct per-elementinputs. Expected
out[k] = a[k] + b[k]; measuredi.e. iteration i+1's
loload into slot 1 landed before iteration i'sTADDhad read slot 1 as
hi. Exactly the missing WAR edge.What does work
multi_tile_getper iteration (the ordinary ping-pong) is guardedcorrectly — dynamic
wait_flag/set_flagpair present, numerically correct ondevice. This is the shape we ship.
reuse, correct.
per-slot ids — ptoas falls back to a single static
wait_flag(PIPE_FIX, PIPE_M, EVENT_ID0)placed after both drains, which serializes conservatively.Expected behaviour
Every
multi_tile_getof a region in an iteration should receive its ownper-slot WAR guard (or, failing that, a conservative guard) — never no guard at
all.
Impact on PyPTO
PyPTO now refuses to lower a slotted allocation with two slots live inside one
loop body under
memory_planner=PTOAS, and points the author at the PyPTO memoryplanner instead. We would like to lift that restriction once this is fixed; on our
side it is a single condition in
PTOCodegen::PlanMultiBufferRegions, with nouser-facing API change.