Summary
A real PyPTO-generated, 3-slot in-place vector software pipeline produces redundant synchronization groups in PTOAS v0.63. After Remove Redundant Sync, both MTE2 -> V and V -> MTE2 still require three groups of three event IDs. Event allocation falls back to two PIPE_ALL barriers per steady-state iteration.
For y = (x + z) + 1, the resulting kernel takes 74,909 simulator cycles, versus 19,064 for the existing 3-way unroll implementation. Both use 24 KiB UB and execute the same loads, vector operations, and stores. A manually synchronized C++ reference with the same rotating buffers and preload schedule passes correctness and takes 18,831 cycles.
The software input already selects each current slot once and reuses that tile handle for tadd, tadds, and tstore; repeated multi_tile_get calls are not required to reproduce the problem.
Related: #1519 — the earlier preload/slot correctness issue is fixed; the current report concerns correct execution with excessive synchronization.
Related: #233 — another auto-sync performance report, with a different workload and reproduction.
Command line
Save the two inputs below as unroll.pto and software_pipeline.pto, then run:
ptoas --version
ptoas --pto-arch=a3 --pto-level=level3 --enable-insert-sync \
--pto-insert-sync-debug=3 unroll.pto -o unroll.cpp 2> unroll.sync.log
ptoas --pto-arch=a3 --pto-level=level3 --enable-insert-sync \
--pto-insert-sync-debug=3 software_pipeline.pto \
-o software_pipeline.cpp 2> software_pipeline.sync.log
These inputs have explicit UB addresses and use level3. They do not depend on the PTOAS memory planner.
Reproduction input
- Function ABI:
(float* y, float* x, float* z); three distinct contiguous GM buffers.
- Shape:
[64, 1024], FP32; each tile is [1, 1024] / 4096 bytes.
- Computation:
x + z, then add 1.0f, then store to y.
- Unroll: three tile pairs per loop; the result reuses the corresponding X buffer.
- Software pipeline:
slots = 3, prefetch = slots - 1 = 2; preload tiles 0/1, load tile i + 2 and compute/store tile i, then drain tiles 62/63. X and Z each occupy three slots; the result reuses X.
Both are actual PTO inputs from the PyPTO experiment. Only source locations were removed and repeated types replaced with MLIR aliases for readability. Recompiling these compact inputs produces byte-identical C++ to the measured originals.
unroll.pto — complete baseline input
!tile = !pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=1024, v_row=?, v_col=?, blayout=row_major, slayout=none_box, fractal=512, pad=0>
!view = !pto.tensor_view<?x?xf32>
!part = !pto.partition_tensor_view<1x1024xf32>
module attributes {pto.target_arch = "a2a3"} {
func.func @vec_add_binary_incore_0(%arg0: !pto.ptr<f32>, %arg1: !pto.ptr<f32>, %arg2: !pto.ptr<f32>) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
%c0_i64 = arith.constant 0 : i64
%c4096_i64 = arith.constant 4096 : i64
%c8192_i64 = arith.constant 8192 : i64
%c12288_i64 = arith.constant 12288 : i64
%c16384_i64 = arith.constant 16384 : i64
%c20480_i64 = arith.constant 20480 : i64
%c64_index = arith.constant 64 : index
%c1024_index = arith.constant 1024 : index
%c1_index = arith.constant 1 : index
%c0_index = arith.constant 0 : index
%c63_index = arith.constant 63 : index
%c3_index = arith.constant 3 : index
%c2_index = arith.constant 2 : index
%cst_13 = arith.constant 1.00000000000000000e+00 : f32
%y__ssa_v0_view = pto.make_tensor_view %arg0, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
%x__ssa_v0_view = pto.make_tensor_view %arg1, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
%z__ssa_v0_view = pto.make_tensor_view %arg2, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
scf.for %i__idx_v0 = %c0_index to %c63_index step %c3_index {
%a__tile = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%12 = arith.maxsi %i__idx_v0, %c0_index : index
%x__ssa_v0_pview = pto.partition_view %x__ssa_v0_view, offsets = [%12, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%x__ssa_v0_pview : !part) outs(%a__tile : !tile)
%b__tile = pto.alloc_tile addr = %c4096_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%13 = arith.maxsi %i__idx_v0, %c0_index : index
%z__ssa_v0_pview = pto.partition_view %z__ssa_v0_view, offsets = [%13, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%z__ssa_v0_pview : !part) outs(%b__tile : !tile)
%0 = pto.alloc_tile addr = %c8192_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%14 = arith.addi %i__idx_v0, %c1_index : index
%15 = arith.maxsi %14, %c0_index : index
%16 = pto.partition_view %x__ssa_v0_view, offsets = [%15, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%16 : !part) outs(%0 : !tile)
%1 = pto.alloc_tile addr = %c12288_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%17 = arith.addi %i__idx_v0, %c1_index : index
%18 = arith.maxsi %17, %c0_index : index
%19 = pto.partition_view %z__ssa_v0_view, offsets = [%18, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%19 : !part) outs(%1 : !tile)
%2 = pto.alloc_tile addr = %c16384_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%20 = arith.addi %i__idx_v0, %c2_index : index
%21 = arith.maxsi %20, %c0_index : index
%22 = pto.partition_view %x__ssa_v0_view, offsets = [%21, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%22 : !part) outs(%2 : !tile)
%3 = pto.alloc_tile addr = %c20480_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%23 = arith.addi %i__idx_v0, %c2_index : index
%24 = arith.maxsi %23, %c0_index : index
%25 = pto.partition_view %z__ssa_v0_view, offsets = [%24, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%25 : !part) outs(%3 : !tile)
%c__tile = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadd ins(%a__tile, %b__tile : !tile, !tile) outs(%c__tile : !tile)
%t__tile = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadds ins(%c__tile, %cst_13 : !tile, f32) outs(%t__tile : !tile)
%26 = arith.maxsi %i__idx_v0, %c0_index : index
%y__iter_v1_pview = pto.partition_view %y__ssa_v0_view, offsets = [%26, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%t__tile : !tile) outs(%y__iter_v1_pview : !part)
%4 = pto.alloc_tile addr = %c8192_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadd ins(%0, %1 : !tile, !tile) outs(%4 : !tile)
%5 = pto.alloc_tile addr = %c8192_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadds ins(%4, %cst_13 : !tile, f32) outs(%5 : !tile)
%28 = arith.addi %i__idx_v0, %c1_index : index
%29 = arith.maxsi %28, %c0_index : index
%y__tile_pview = pto.partition_view %y__ssa_v0_view, offsets = [%29, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%5 : !tile) outs(%y__tile_pview : !part)
%6 = pto.alloc_tile addr = %c16384_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadd ins(%2, %3 : !tile, !tile) outs(%6 : !tile)
%7 = pto.alloc_tile addr = %c16384_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadds ins(%6, %cst_13 : !tile, f32) outs(%7 : !tile)
%31 = arith.addi %i__idx_v0, %c2_index : index
%32 = arith.maxsi %31, %c0_index : index
%33 = pto.partition_view %y__ssa_v0_view, offsets = [%32, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%7 : !tile) outs(%33 : !part)
}
%8 = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%34 = pto.partition_view %x__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%34 : !part) outs(%8 : !tile)
%9 = pto.alloc_tile addr = %c4096_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
%35 = pto.partition_view %z__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%35 : !part) outs(%9 : !tile)
%10 = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadd ins(%8, %9 : !tile, !tile) outs(%10 : !tile)
%11 = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
pto.tadds ins(%10, %cst_13 : !tile, f32) outs(%11 : !tile)
%y__rv_v2_main_pview = pto.partition_view %y__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%11 : !tile) outs(%y__rv_v2_main_pview : !part)
return
}
}
software_pipeline.pto — complete failing-performance input
!tile = !pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=1024, v_row=?, v_col=?, blayout=row_major, slayout=none_box, fractal=512, pad=0>
!view = !pto.tensor_view<?x?xf32>
!part = !pto.partition_tensor_view<1x1024xf32>
!ring = !pto.multi_tile_buf<!pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=1024, v_row=?, v_col=?, blayout=row_major, slayout=none_box, fractal=512, pad=0>, count=3>
module attributes {pto.target_arch = "a2a3"} {
func.func @vec_add_binary_incore_0(%arg0: !pto.ptr<f32>, %arg1: !pto.ptr<f32>, %arg2: !pto.ptr<f32>) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
%c0_i64 = arith.constant 0 : i64
%c12288_i64 = arith.constant 12288 : i64
%c4096_i64 = arith.constant 4096 : i64
%c16384_i64 = arith.constant 16384 : i64
%c8192_i64 = arith.constant 8192 : i64
%c20480_i64 = arith.constant 20480 : i64
%c1_index = arith.constant 1 : index
%c1024_index = arith.constant 1024 : index
%c64_index = arith.constant 64 : index
%c0_index = arith.constant 0 : index
%c62_index = arith.constant 62 : index
%c2_index = arith.constant 2 : index
%c3_index = arith.constant 3 : index
%cst_13 = arith.constant 1.00000000000000000e+00 : f32
%c63_index = arith.constant 63 : index
%y__ssa_v0_view = pto.make_tensor_view %arg0, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
%x__ssa_v0_view = pto.make_tensor_view %arg1, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
%z__ssa_v0_view = pto.make_tensor_view %arg2, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
%pipe_a__tile_0_mb = pto.alloc_multi_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !ring
%pipe_b__tile_1_mb = pto.alloc_multi_tile addr = %c12288_i64 valid_row = %c1_index valid_col = %c1024_index : !ring
%a__tile = pto.multi_tile_get %pipe_a__tile_0_mb[%c0_index] : !ring -> !tile
%x__ssa_v0_pview = pto.partition_view %x__ssa_v0_view, offsets = [%c0_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%x__ssa_v0_pview : !part) outs(%a__tile : !tile)
%b__tile = pto.multi_tile_get %pipe_b__tile_1_mb[%c0_index] : !ring -> !tile
%z__ssa_v0_pview = pto.partition_view %z__ssa_v0_view, offsets = [%c0_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%z__ssa_v0_pview : !part) outs(%b__tile : !tile)
%0 = pto.multi_tile_get %pipe_a__tile_0_mb[%c1_index] : !ring -> !tile
%14 = pto.partition_view %x__ssa_v0_view, offsets = [%c1_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%14 : !part) outs(%0 : !tile)
%1 = pto.multi_tile_get %pipe_b__tile_1_mb[%c1_index] : !ring -> !tile
%15 = pto.partition_view %z__ssa_v0_view, offsets = [%c1_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%15 : !part) outs(%1 : !tile)
scf.for %i__idx_v0_pipe = %c0_index to %c62_index step %c1_index {
%16 = arith.addi %i__idx_v0_pipe, %c2_index : index
%17 = arith.remui %16, %c3_index : index
%2 = pto.multi_tile_get %pipe_a__tile_0_mb[%17] : !ring -> !tile
%18 = arith.addi %i__idx_v0_pipe, %c2_index : index
%19 = arith.maxsi %18, %c0_index : index
%20 = pto.partition_view %x__ssa_v0_view, offsets = [%19, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%20 : !part) outs(%2 : !tile)
%21 = arith.addi %i__idx_v0_pipe, %c2_index : index
%22 = arith.remui %21, %c3_index : index
%3 = pto.multi_tile_get %pipe_b__tile_1_mb[%22] : !ring -> !tile
%23 = arith.addi %i__idx_v0_pipe, %c2_index : index
%24 = arith.maxsi %23, %c0_index : index
%25 = pto.partition_view %z__ssa_v0_view, offsets = [%24, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tload ins(%25 : !part) outs(%3 : !tile)
%26 = arith.remui %i__idx_v0_pipe, %c3_index : index
%4 = pto.multi_tile_get %pipe_a__tile_0_mb[%26] : !ring -> !tile
%27 = arith.remui %i__idx_v0_pipe, %c3_index : index
%5 = pto.multi_tile_get %pipe_b__tile_1_mb[%27] : !ring -> !tile
pto.tadd ins(%4, %5 : !tile, !tile) outs(%4 : !tile)
pto.tadds ins(%4, %cst_13 : !tile, f32) outs(%4 : !tile)
%28 = arith.maxsi %i__idx_v0_pipe, %c0_index : index
%y__iter_v1_pview = pto.partition_view %y__ssa_v0_view, offsets = [%28, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%4 : !tile) outs(%y__iter_v1_pview : !part)
}
%6 = pto.multi_tile_get %pipe_a__tile_0_mb[%c2_index] : !ring -> !tile
%7 = pto.multi_tile_get %pipe_b__tile_1_mb[%c2_index] : !ring -> !tile
pto.tadd ins(%6, %7 : !tile, !tile) outs(%6 : !tile)
pto.tadds ins(%6, %cst_13 : !tile, f32) outs(%6 : !tile)
%y__rv_v2_steady_pview = pto.partition_view %y__ssa_v0_view, offsets = [%c62_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%6 : !tile) outs(%y__rv_v2_steady_pview : !part)
%10 = pto.multi_tile_get %pipe_a__tile_0_mb[%c0_index] : !ring -> !tile
%11 = pto.multi_tile_get %pipe_b__tile_1_mb[%c0_index] : !ring -> !tile
pto.tadd ins(%10, %11 : !tile, !tile) outs(%10 : !tile)
pto.tadds ins(%10, %cst_13 : !tile, f32) outs(%10 : !tile)
%y__tile_pview = pto.partition_view %y__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
pto.tstore ins(%10 : !tile) outs(%y__tile_pview : !part)
return
}
}
Expected performance
Preserve the required RAW/WAR/WAW dependencies, including slot reuse, preload, and drain, without exhausting event IDs or inserting PIPE_ALL inside this loop. A final completion barrier is expected.
A sufficient dependency graph is:
LOAD_X(i) --------------------> ADD(i)
LOAD_Z(i) --------------------> ADD(i)
ADD(i) -> ADDS(i) -> STORE_X(i) -> LOAD_X(i + 3)
ADD(i) ------------------------> LOAD_Z(i + 3)
Consequently, several retained direct edges are transitively covered:
| Retained group |
Covering path for this input |
LOAD_X -> ADDS (idx=11) |
LOAD_X -> ADD -> ADDS |
ADD/ADDS -> next LOAD_X (idx=5/4) |
ADD -> ADDS -> STORE_X -> next LOAD_X |
Previous STORE_X -> ADD/ADDS (idx=9/12) |
Previous STORE_X -> LOAD_X -> ADD -> ADDS |
LOAD_X -> STORE_X (idx=13) |
LOAD_X -> ADD -> ADDS -> STORE_X |
The following complete, hand-written expected C++ reference demonstrates a sufficient event protocol. It is not output from a modified PTOAS compiler. It keeps the software input's 24 KiB UB layout, in-place operations, three slots, and two-tile preload, and uses:
| Dependency |
Pipe pair |
Event IDs |
| X load ready |
MTE2 -> V |
0–2 |
| Z load ready |
MTE2 -> V |
3–5 |
| Z slot free after ADD |
V -> MTE2 |
0–2 |
| Result ready for store |
V -> MTE3 |
3–5 |
| X slot free after STORE |
MTE3 -> MTE2 |
0–2 |
Only free-slot tokens are primed. Ready tokens come from real loads/computes; remaining free-slot tokens are drained at exit.
expected.cpp — complete reference, compiled and simulator-checked
// MANUAL EXPECTED REFERENCE: this file was not emitted by PTOAS.
// Workload: y = (x + z) + 1, 64 rows of 1024 FP32 values, slots=3, prefetch=2.
// Change kTrips, and the sibling .pto shape metadata, together to test another N.
// UB layout matches the actual software-pipeline case exactly: X=[0,12288),
// Z=[12288,24576). No extra output or intermediate buffer is introduced.
#include "pto/pto-inst.hpp"
using namespace pto;
namespace {
constexpr int64_t kTrips = 64;
constexpr int64_t kWidth = 1024;
constexpr int64_t kSlots = 3;
constexpr int64_t kPrefetch = kSlots - 1;
constexpr uint64_t kSlotBytes = kWidth * sizeof(float);
constexpr uint64_t kZBase = kSlots * kSlotBytes;
static_assert(kTrips >= kPrefetch, "This reference requires at least two tiles");
#ifdef __DAV_VEC__
// Dynamic valid extents match the actual emitted C++ tile type; their runtime
// values remain the same constants, 1 and 1024.
using VecTile = Tile<TileType::Vec, float, 1, kWidth, BLayout::RowMajor, -1, -1, SLayout::NoneBox, 512,
PadValue::Null, CompactMode::Null>;
using RowTensor =
GlobalTensor<float, Shape<1, 1, 1, 1, kWidth>, Stride<kWidth, kWidth, kWidth, kWidth, 1>, Layout::ND>;
static AICORE inline void LoadPair(__gm__ float* x, __gm__ float* z, int64_t tile) {
const int64_t slot = tile % kSlots;
const event_t slot_event = static_cast<event_t>(slot);
const event_t z_ready_event = static_cast<event_t>(kSlots + slot);
VecTile next_x(1, kWidth);
VecTile next_z(1, kWidth);
TASSIGN(next_x, static_cast<uint64_t>(slot) * kSlotBytes);
TASSIGN(next_z, kZBase + static_cast<uint64_t>(slot) * kSlotBytes);
RowTensor x_row(x + tile * kWidth);
RowTensor z_row(z + tile * kWidth);
// X remains live through STORE_X; Z becomes free after the first vector add.
wait_flag(PIPE_MTE3, PIPE_MTE2, slot_event);
TLOAD(next_x, x_row);
set_flag(PIPE_MTE2, PIPE_V, slot_event);
wait_flag(PIPE_V, PIPE_MTE2, slot_event);
TLOAD(next_z, z_row);
set_flag(PIPE_MTE2, PIPE_V, z_ready_event);
}
static AICORE inline void ComputeAndStore(__gm__ float* y, int64_t tile) {
const int64_t slot = tile % kSlots;
const event_t slot_event = static_cast<event_t>(slot);
const event_t z_ready_event = static_cast<event_t>(kSlots + slot);
VecTile x_s(1, kWidth);
VecTile z_s(1, kWidth);
TASSIGN(x_s, static_cast<uint64_t>(slot) * kSlotBytes);
TASSIGN(z_s, kZBase + static_cast<uint64_t>(slot) * kSlotBytes);
RowTensor y_row(y + tile * kWidth);
wait_flag(PIPE_MTE2, PIPE_V, slot_event);
wait_flag(PIPE_MTE2, PIPE_V, z_ready_event);
// One current-X handle is used for both updates and the final store.
TADD(x_s, x_s, z_s);
set_flag(PIPE_V, PIPE_MTE2, slot_event);
pipe_barrier(PIPE_V);
TADDS(x_s, x_s, 1.0f);
// Keep V->MTE3 IDs disjoint from the V->MTE2 free-Z IDs.
const event_t store_ready_event = static_cast<event_t>(kSlots + slot);
set_flag(PIPE_V, PIPE_MTE3, store_ready_event);
wait_flag(PIPE_V, PIPE_MTE3, store_ready_event);
TSTORE(y_row, x_s);
set_flag(PIPE_MTE3, PIPE_MTE2, slot_event);
}
#endif
} // namespace
AICORE void vec_add_binary_incore_0(__gm__ float* y, __gm__ float* x, __gm__ float* z) {
#if defined(__DAV_VEC__)
set_mask_norm();
set_vector_mask(-1, -1);
// Prime exactly one free token per physical slot. Ready tokens are produced
// only by real loads/computes, so no uninitialized tile can be consumed.
for (int64_t slot = 0; slot < kSlots; ++slot) {
const event_t event = static_cast<event_t>(slot);
set_flag(PIPE_MTE3, PIPE_MTE2, event);
set_flag(PIPE_V, PIPE_MTE2, event);
}
for (int64_t tile = 0; tile < kPrefetch; ++tile) {
LoadPair(x, z, tile);
}
// The same preload / load-next / compute-current / drain schedule as the PTO.
for (int64_t tile = 0; tile < kTrips - kPrefetch; ++tile) {
LoadPair(x, z, tile + kPrefetch);
ComputeAndStore(y, tile);
}
for (int64_t tile = kTrips - kPrefetch; tile < kTrips; ++tile) {
ComputeAndStore(y, tile);
}
// Each load consumes one free token and each completed use replaces it.
// Thus exactly one free X and Z token remains per slot at kernel exit.
// All load-ready and store-ready tokens have already been consumed.
for (int64_t slot = 0; slot < kSlots; ++slot) {
const event_t event = static_cast<event_t>(slot);
wait_flag(PIPE_MTE3, PIPE_MTE2, event);
wait_flag(PIPE_V, PIPE_MTE2, event);
}
pipe_barrier(PIPE_ALL); // Final completion only; no PIPE_ALL in steady state.
#endif
}
// Why the omitted edges are covered, including preload and drain boundaries:
// STORE_X(i-3) -> LOAD_X(i) -> TADD(i) -> TADDS(i) -> STORE_X(i).
// TADD(i-3) -> LOAD_Z(i) -> TADD(i).
// The first use of each slot consumes its seeded free token instead of the
// predecessor edge. Loads 0/1 are issued in the prologue; every later load is
// issued once by the steady loop. The final two computes consume their actual
// ready tokens in the drain. This covers MTE3->V and MTE2->MTE3 transitively;
// the V barrier orders TADD->TADDS without a second MTE2->V event family.
This reference passes exact FP32 equality for all 65,536 elements at N=64, and also for all 66,560 elements at N=65 with kTrips=65 and matching buffer sizes. The latter checks a different slot position at the drain boundary.
For N=64 it takes 18,831 simulator cycles, close to the 19,064-cycle unroll baseline. Its scalar/address-expression code also differs from PTOAS output, so this is evidence that an efficient schedule exists, not an isolated measurement of a sync-only compiler patch or a requirement for exact emitted C++ text.
Actual performance
The following groups remain in software_pipeline.sync.log under After Remove Redundant Sync:
| Pipe pair |
Group indices |
Requested rotating event IDs |
| MTE2 -> V |
7, 8, 11 |
3 + 3 + 3 = 9 |
| V -> MTE2 |
4, 5, 6 |
3 + 3 + 3 = 9 |
| MTE3 -> V |
9, 12 |
3 + 3 = 6 |
| MTE3 -> MTE2 |
3 |
3 |
| MTE2 -> MTE3 |
13 |
3 |
There is also the non-rotating V -> MTE3 store-ready group (idx=2). The table describes PTOAS's retained demand, not the algorithm's minimum requirements. The v0.63 allocator defines an eight-ID event pool; see SyncEventIdAllocation.h.
After event allocation, groups 6 and 11 become full barriers. Relevant lines copied from After EventId Allocation:
[ 6] COMPOUND pto.tload [PIPE_MTE2]
def=[%4(VEC)]
use=[%arg2(GM)]
PRE : pipe_barrier <PIPE_ALL -> PIPE_ALL> idx=6 forEnd=10 eventIdNum=3
[ 8] COMPOUND pto.tadds [PIPE_V]
def=[%3(VEC)]
use=[%3(VEC)]
PRE : pipe_barrier <PIPE_V -> PIPE_V> idx=1
PRE : pipe_barrier <PIPE_ALL -> PIPE_ALL> idx=11 forEnd=10 eventIdNum=3
PRE : wait_flag <PIPE_MTE3 -> PIPE_V> idx=12 forEnd=10 eventIdNum=3 eventIds=[3,4,5]
Their matching set operations are marked useless. The emitted C++ confirms the barrier positions, before the second load and between the two vector operations. These are actual excerpts from inside the loop:
pipe_barrier(PIPE_ALL);
TLOAD(v103, v110);
pipe_barrier(PIPE_V);
pipe_barrier(PIPE_ALL);
event_t v140 = (event_t) v138;
wait_flag(PIPE_MTE3, PIPE_V, v140);
TADDS(v119, v119, v20);
These two barriers execute in each of the 62 steady-state iterations; they are separate from the normal kernel-tail completion barrier.
The requested improvement is to eliminate transitively covered dependencies before event allocation, while preserving slot lifetimes and correct priming/draining. Simply dropping the barrier fallback or collapsing a rotating dependency to one event is not a safe general replacement.
Profiling data (optional)
Measured with the Ascend operator simulator, one active AIV, CANN 9.0.0, Ascend910B1 / dav-c220. These are in-core simulated cycle spans, not hardware wall-clock or end-to-end latency measurements.
| N=64 variant |
UB |
Steady-loop PIPE_ALL barriers per iteration |
Active AIV span, cycles |
Exact correctness |
| PTOAS unroll baseline |
24 KiB |
0 |
19,064 |
PASS |
| PTOAS rotating software pipeline |
24 KiB |
2 |
74,909 |
PASS |
| Manual expected C++ above |
24 KiB |
0 |
18,831 |
PASS |
The current software output is 3.93x the unroll baseline's cycle span. All three traces contain exactly 128 MTE2 loads, 64 VADD, 64 VADDS, and 64 MTE3 stores.
Inputs are generated once with NumPy np.random.seed(19) followed by two sequential np.random.random(65536).astype(np.float32) calls, and reused byte-for-byte. Output is initialized to NaN; the host checks every element using FP32 expected = x + z; expected += 1.0f and exact equality.
X SHA256: b373d6442009c54d11e2f2c83f0c88e81822e35d187c50542e49643f6bd0357f
Z SHA256: aaa012d7e19f43f34bab677b12de7e244cb51c3214ffb2365a316471825677ed
Cycle span is computed from core0.veccore0/trace.json: keep complete (ph="X") instruction events, excluding the process whose process_name is CACHEMISS, and calculate (max(ts + dur) - min(ts)) * 1850. This includes scalar setup, synchronization, and tail instructions. All rows use the same measurement convention.
Generated/reference C++ SHA256 values:
unroll.cpp: 4175c1586abcdd2fd5757c0a12fb166d6c1e82fbe409173b01d2b96ad61b5216
software_pipeline.cpp: 44f7cac9090ec6febc770d36b997dfcedae5dc57c313302c7536c46e5bbc748f
expected.cpp: 197e140b1144102903b54ece2e161d51ecbce3d30989866bf5e136fff4ef1a8b
Git commit
PTOAS v0.63, release tag commit 470eab4; the tested binary reports ptoas 0.63.
PTO-ISA used to compile the simulator cases: 3b4faf67aebb3e0d41be7952c56908b3adba7a8f.
Summary
A real PyPTO-generated, 3-slot in-place vector software pipeline produces redundant synchronization groups in PTOAS v0.63. After
Remove Redundant Sync, bothMTE2 -> VandV -> MTE2still require three groups of three event IDs. Event allocation falls back to twoPIPE_ALLbarriers per steady-state iteration.For
y = (x + z) + 1, the resulting kernel takes 74,909 simulator cycles, versus 19,064 for the existing 3-way unroll implementation. Both use 24 KiB UB and execute the same loads, vector operations, and stores. A manually synchronized C++ reference with the same rotating buffers and preload schedule passes correctness and takes 18,831 cycles.The software input already selects each current slot once and reuses that tile handle for
tadd,tadds, andtstore; repeatedmulti_tile_getcalls are not required to reproduce the problem.Related: #1519 — the earlier preload/slot correctness issue is fixed; the current report concerns correct execution with excessive synchronization.
Related: #233 — another auto-sync performance report, with a different workload and reproduction.
Command line
Save the two inputs below as
unroll.ptoandsoftware_pipeline.pto, then run:These inputs have explicit UB addresses and use level3. They do not depend on the PTOAS memory planner.
Reproduction input
(float* y, float* x, float* z); three distinct contiguous GM buffers.[64, 1024], FP32; each tile is[1, 1024]/ 4096 bytes.x + z, then add1.0f, then store toy.slots = 3,prefetch = slots - 1 = 2; preload tiles 0/1, load tilei + 2and compute/store tilei, then drain tiles 62/63. X and Z each occupy three slots; the result reuses X.Both are actual PTO inputs from the PyPTO experiment. Only source locations were removed and repeated types replaced with MLIR aliases for readability. Recompiling these compact inputs produces byte-identical C++ to the measured originals.
unroll.pto — complete baseline input
software_pipeline.pto — complete failing-performance input
Expected performance
Preserve the required RAW/WAR/WAW dependencies, including slot reuse, preload, and drain, without exhausting event IDs or inserting
PIPE_ALLinside this loop. A final completion barrier is expected.A sufficient dependency graph is:
Consequently, several retained direct edges are transitively covered:
idx=11)idx=5/4)idx=9/12)idx=13)The following complete, hand-written expected C++ reference demonstrates a sufficient event protocol. It is not output from a modified PTOAS compiler. It keeps the software input's 24 KiB UB layout, in-place operations, three slots, and two-tile preload, and uses:
Only free-slot tokens are primed. Ready tokens come from real loads/computes; remaining free-slot tokens are drained at exit.
expected.cpp — complete reference, compiled and simulator-checked
This reference passes exact FP32 equality for all 65,536 elements at N=64, and also for all 66,560 elements at N=65 with
kTrips=65and matching buffer sizes. The latter checks a different slot position at the drain boundary.For N=64 it takes 18,831 simulator cycles, close to the 19,064-cycle unroll baseline. Its scalar/address-expression code also differs from PTOAS output, so this is evidence that an efficient schedule exists, not an isolated measurement of a sync-only compiler patch or a requirement for exact emitted C++ text.
Actual performance
The following groups remain in
software_pipeline.sync.logunderAfter Remove Redundant Sync:There is also the non-rotating V -> MTE3 store-ready group (
idx=2). The table describes PTOAS's retained demand, not the algorithm's minimum requirements. The v0.63 allocator defines an eight-ID event pool; see SyncEventIdAllocation.h.After event allocation, groups 6 and 11 become full barriers. Relevant lines copied from
After EventId Allocation:Their matching set operations are marked
useless. The emitted C++ confirms the barrier positions, before the second load and between the two vector operations. These are actual excerpts from inside the loop:These two barriers execute in each of the 62 steady-state iterations; they are separate from the normal kernel-tail completion barrier.
The requested improvement is to eliminate transitively covered dependencies before event allocation, while preserving slot lifetimes and correct priming/draining. Simply dropping the barrier fallback or collapsing a rotating dependency to one event is not a safe general replacement.
Profiling data (optional)
Measured with the Ascend operator simulator, one active AIV, CANN 9.0.0,
Ascend910B1/dav-c220. These are in-core simulated cycle spans, not hardware wall-clock or end-to-end latency measurements.The current software output is 3.93x the unroll baseline's cycle span. All three traces contain exactly 128 MTE2 loads, 64 VADD, 64 VADDS, and 64 MTE3 stores.
Inputs are generated once with NumPy
np.random.seed(19)followed by two sequentialnp.random.random(65536).astype(np.float32)calls, and reused byte-for-byte. Output is initialized to NaN; the host checks every element using FP32expected = x + z; expected += 1.0fand exact equality.Cycle span is computed from
core0.veccore0/trace.json: keep complete (ph="X") instruction events, excluding the process whoseprocess_nameisCACHEMISS, and calculate(max(ts + dur) - min(ts)) * 1850. This includes scalar setup, synchronization, and tail instructions. All rows use the same measurement convention.Generated/reference C++ SHA256 values:
Git commit
PTOAS v0.63, release tag commit 470eab4; the tested binary reports
ptoas 0.63.PTO-ISA used to compile the simulator cases: 3b4faf67aebb3e0d41be7952c56908b3adba7a8f.