Skip to content

[Pass Bug] level2 insert-sync guards only the first multi_tile_get of an iteration; a second co-live slot is left unsynchronized #1118

Description

@lyfne123

Summary

At --pto-level=level2 --enable-insert-sync, when a loop body performs two
pto.multi_tile_get on the same !pto.multi_tile_buf region, ptoas emits the
per-slot dynamic WAR pair (wait_flag/set_flag on a derived event_t) for the
first access only. The second slot's TLOAD is emitted with no preceding
wait_flag at all
, so the next iteration's write into that slot races the
current iteration's read of it.

The result is a silently wrong kernel on device — not a missed optimization.

Related to but distinct from #1106 (that one is a level3 address-folding issue;
this one is a level2 sync-derivation gap).

Environment

  • ptoas v0.54 (aarch64), --enable-insert-sync --pto-level=level2 --pto-arch a3
  • Device: Ascend a2a3
  • --enable-graph-sync-solver behaves the same.

Reproducer

Input .pto (trimmed — one region, two gets per iteration, both slot indices
affine in the loop variable):

%ub_mb = pto.alloc_multi_tile valid_row = %c64_index valid_col = %c64_index
       : !pto.multi_tile_buf<!pto.tile_buf<loc=vec, dtype=f32, rows=64, cols=64,
                                           v_row=?, v_col=?, blayout=row_major,
                                           slayout=none_box, fractal=512, pad=0>, count=2>
scf.for %i = %c0_index to %c4_index step %c1_index {
  %0  = arith.remsi %i, %c2_index : index                 // i % 2
  %lo = pto.multi_tile_get %ub_mb[%0]  : ... -> !pto.tile_buf<loc=vec, ...>
  pto.tload ins(%a_view ...) outs(%lo : ...)
  %3  = arith.addi %i, %c1_index : index
  %4  = arith.remsi %3, %c2_index : index                 // (i + 1) % 2
  %hi = pto.multi_tile_get %ub_mb[%4]  : ... -> !pto.tile_buf<loc=vec, ...>
  pto.tload ins(%b_view ...) outs(%hi : ...)
  %s  = pto.tadd ins(%lo, %hi : ...) outs(%s : ...)       // both slots read here
  pto.tstore ins(%s ...) outs(%out_view ...)
}

Emitted code (v0.54)

// prologue — BOTH event ids are primed, so the region's slot count is understood
set_flag(PIPE_V, PIPE_MTE2, EVENT_ID0);
set_flag(PIPE_V, PIPE_MTE2, EVENT_ID1);

for (int64_t i11 = v8; i11 < v5; i11 += v6) {
  int64_t v12 = i11 % v4;
  uint64_t v14 = (uint64_t)(v12 == v6 ? v9 : v8);
  TASSIGN(v13, v14);                              // lo -> slot i % 2

  int64_t  v20 = (int64_t)((uint64_t)v12 % (uint64_t)v4) == v6 ? v6 : v8;
  event_t  v21 = (event_t) v20;                   // dynamic id derived from lo's slot
  wait_flag(PIPE_V, PIPE_MTE2, v21);              // <-- guards lo
  TLOAD(v13, v19);

  uint64_t v23 = (uint64_t)((int64_t)((uint64_t)i11 + (uint64_t)v6) % v4 == v6 ? v9 : v8);
  TASSIGN(v22, v23);                              // hi -> slot (i + 1) % 2
  TLOAD(v22, v26);                                // <-- NO wait_flag before this load

  set_flag(PIPE_MTE2, PIPE_V, EVENT_ID0);
  wait_flag(PIPE_MTE2, PIPE_V, EVENT_ID0);
  TADD(v27, v13, v22);                            // reads BOTH slots

  event_t v29 = (event_t) v20;
  set_flag(PIPE_V, PIPE_MTE2, v29);               // <-- releases lo's slot only
  ...
}

EVENT_ID1 is primed in the prologue and then never waited on or set inside the
loop: only one of the two slots is ever guarded.

Observed failure

out = a + b over 4 row-blocks of [64, 64] fp32, with distinct per-element
inputs. Expected out[k] = a[k] + b[k]; measured

out[k] = a[block i+1][k] + a[block i][k]

i.e. iteration i+1's lo load into slot 1 landed before iteration i's TADD
had read slot 1 as hi. Exactly the missing WAR edge.

What does work

  • One multi_tile_get per iteration (the ordinary ping-pong) is guarded
    correctly — dynamic wait_flag/set_flag pair present, numerically correct on
    device. This is the shape we ship.
  • Straight-line code with two constant slots: no loop, no cross-iteration
    reuse, correct.
  • An Acc/L0C region with two co-live slots happens to be correct, but not via
    per-slot ids — ptoas falls back to a single static wait_flag(PIPE_FIX, PIPE_M, EVENT_ID0) placed after both drains, which serializes conservatively.

Expected behaviour

Every multi_tile_get of a region in an iteration should receive its own
per-slot WAR guard (or, failing that, a conservative guard) — never no guard at
all.

Impact on PyPTO

PyPTO now refuses to lower a slotted allocation with two slots live inside one
loop body under memory_planner=PTOAS, and points the author at the PyPTO memory
planner instead. We would like to lift that restriction once this is fixed; on our
side it is a single condition in PTOCodegen::PlanMultiBufferRegions, with no
user-facing API change.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions