Skip to content

[Performance] Redundant sync exhausts event IDs in a 3-slot in-place Vec pipeline #1569

Description

@Hzfengsy

Summary

A real PyPTO-generated, 3-slot in-place vector software pipeline produces redundant synchronization groups in PTOAS v0.63. After Remove Redundant Sync, both MTE2 -> V and V -> MTE2 still require three groups of three event IDs. Event allocation falls back to two PIPE_ALL barriers per steady-state iteration.

For y = (x + z) + 1, the resulting kernel takes 74,909 simulator cycles, versus 19,064 for the existing 3-way unroll implementation. Both use 24 KiB UB and execute the same loads, vector operations, and stores. A manually synchronized C++ reference with the same rotating buffers and preload schedule passes correctness and takes 18,831 cycles.

The software input already selects each current slot once and reuses that tile handle for tadd, tadds, and tstore; repeated multi_tile_get calls are not required to reproduce the problem.

Related: #1519 — the earlier preload/slot correctness issue is fixed; the current report concerns correct execution with excessive synchronization.

Related: #233 — another auto-sync performance report, with a different workload and reproduction.

Command line

Save the two inputs below as unroll.pto and software_pipeline.pto, then run:

ptoas --version
ptoas --pto-arch=a3 --pto-level=level3 --enable-insert-sync \
  --pto-insert-sync-debug=3 unroll.pto -o unroll.cpp 2> unroll.sync.log
ptoas --pto-arch=a3 --pto-level=level3 --enable-insert-sync \
  --pto-insert-sync-debug=3 software_pipeline.pto \
  -o software_pipeline.cpp 2> software_pipeline.sync.log

These inputs have explicit UB addresses and use level3. They do not depend on the PTOAS memory planner.

Reproduction input

  • Function ABI: (float* y, float* x, float* z); three distinct contiguous GM buffers.
  • Shape: [64, 1024], FP32; each tile is [1, 1024] / 4096 bytes.
  • Computation: x + z, then add 1.0f, then store to y.
  • Unroll: three tile pairs per loop; the result reuses the corresponding X buffer.
  • Software pipeline: slots = 3, prefetch = slots - 1 = 2; preload tiles 0/1, load tile i + 2 and compute/store tile i, then drain tiles 62/63. X and Z each occupy three slots; the result reuses X.

Both are actual PTO inputs from the PyPTO experiment. Only source locations were removed and repeated types replaced with MLIR aliases for readability. Recompiling these compact inputs produces byte-identical C++ to the measured originals.

unroll.pto — complete baseline input
!tile = !pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=1024, v_row=?, v_col=?, blayout=row_major, slayout=none_box, fractal=512, pad=0>
!view = !pto.tensor_view<?x?xf32>
!part = !pto.partition_tensor_view<1x1024xf32>

module attributes {pto.target_arch = "a2a3"} {
  func.func @vec_add_binary_incore_0(%arg0: !pto.ptr<f32>, %arg1: !pto.ptr<f32>, %arg2: !pto.ptr<f32>) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
  %c0_i64 = arith.constant 0 : i64
  %c4096_i64 = arith.constant 4096 : i64
  %c8192_i64 = arith.constant 8192 : i64
  %c12288_i64 = arith.constant 12288 : i64
  %c16384_i64 = arith.constant 16384 : i64
  %c20480_i64 = arith.constant 20480 : i64
  %c64_index = arith.constant 64 : index
  %c1024_index = arith.constant 1024 : index
  %c1_index = arith.constant 1 : index
  %c0_index = arith.constant 0 : index
  %c63_index = arith.constant 63 : index
  %c3_index = arith.constant 3 : index
  %c2_index = arith.constant 2 : index
  %cst_13 = arith.constant 1.00000000000000000e+00 : f32
  %y__ssa_v0_view = pto.make_tensor_view %arg0, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
  %x__ssa_v0_view = pto.make_tensor_view %arg1, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
  %z__ssa_v0_view = pto.make_tensor_view %arg2, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
  scf.for %i__idx_v0 = %c0_index to %c63_index step %c3_index {
    %a__tile = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    %12 = arith.maxsi %i__idx_v0, %c0_index : index
    %x__ssa_v0_pview = pto.partition_view %x__ssa_v0_view, offsets = [%12, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%x__ssa_v0_pview : !part) outs(%a__tile : !tile)
    %b__tile = pto.alloc_tile addr = %c4096_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    %13 = arith.maxsi %i__idx_v0, %c0_index : index
    %z__ssa_v0_pview = pto.partition_view %z__ssa_v0_view, offsets = [%13, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%z__ssa_v0_pview : !part) outs(%b__tile : !tile)
    %0 = pto.alloc_tile addr = %c8192_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    %14 = arith.addi %i__idx_v0, %c1_index : index
    %15 = arith.maxsi %14, %c0_index : index
    %16 = pto.partition_view %x__ssa_v0_view, offsets = [%15, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%16 : !part) outs(%0 : !tile)
    %1 = pto.alloc_tile addr = %c12288_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    %17 = arith.addi %i__idx_v0, %c1_index : index
    %18 = arith.maxsi %17, %c0_index : index
    %19 = pto.partition_view %z__ssa_v0_view, offsets = [%18, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%19 : !part) outs(%1 : !tile)
    %2 = pto.alloc_tile addr = %c16384_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    %20 = arith.addi %i__idx_v0, %c2_index : index
    %21 = arith.maxsi %20, %c0_index : index
    %22 = pto.partition_view %x__ssa_v0_view, offsets = [%21, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%22 : !part) outs(%2 : !tile)
    %3 = pto.alloc_tile addr = %c20480_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    %23 = arith.addi %i__idx_v0, %c2_index : index
    %24 = arith.maxsi %23, %c0_index : index
    %25 = pto.partition_view %z__ssa_v0_view, offsets = [%24, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%25 : !part) outs(%3 : !tile)
    %c__tile = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    pto.tadd ins(%a__tile, %b__tile : !tile, !tile) outs(%c__tile : !tile)
    %t__tile = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    pto.tadds ins(%c__tile, %cst_13 : !tile, f32) outs(%t__tile : !tile)
    %26 = arith.maxsi %i__idx_v0, %c0_index : index
    %y__iter_v1_pview = pto.partition_view %y__ssa_v0_view, offsets = [%26, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tstore ins(%t__tile : !tile) outs(%y__iter_v1_pview : !part)
    %4 = pto.alloc_tile addr = %c8192_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    pto.tadd ins(%0, %1 : !tile, !tile) outs(%4 : !tile)
    %5 = pto.alloc_tile addr = %c8192_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    pto.tadds ins(%4, %cst_13 : !tile, f32) outs(%5 : !tile)
    %28 = arith.addi %i__idx_v0, %c1_index : index
    %29 = arith.maxsi %28, %c0_index : index
    %y__tile_pview = pto.partition_view %y__ssa_v0_view, offsets = [%29, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tstore ins(%5 : !tile) outs(%y__tile_pview : !part)
    %6 = pto.alloc_tile addr = %c16384_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    pto.tadd ins(%2, %3 : !tile, !tile) outs(%6 : !tile)
    %7 = pto.alloc_tile addr = %c16384_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
    pto.tadds ins(%6, %cst_13 : !tile, f32) outs(%7 : !tile)
    %31 = arith.addi %i__idx_v0, %c2_index : index
    %32 = arith.maxsi %31, %c0_index : index
    %33 = pto.partition_view %y__ssa_v0_view, offsets = [%32, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tstore ins(%7 : !tile) outs(%33 : !part)
  }
  %8 = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
  %34 = pto.partition_view %x__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tload ins(%34 : !part) outs(%8 : !tile)
  %9 = pto.alloc_tile addr = %c4096_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
  %35 = pto.partition_view %z__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tload ins(%35 : !part) outs(%9 : !tile)
  %10 = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
  pto.tadd ins(%8, %9 : !tile, !tile) outs(%10 : !tile)
  %11 = pto.alloc_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !tile
  pto.tadds ins(%10, %cst_13 : !tile, f32) outs(%11 : !tile)
  %y__rv_v2_main_pview = pto.partition_view %y__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tstore ins(%11 : !tile) outs(%y__rv_v2_main_pview : !part)
  return
  }
}
software_pipeline.pto — complete failing-performance input
!tile = !pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=1024, v_row=?, v_col=?, blayout=row_major, slayout=none_box, fractal=512, pad=0>
!view = !pto.tensor_view<?x?xf32>
!part = !pto.partition_tensor_view<1x1024xf32>
!ring = !pto.multi_tile_buf<!pto.tile_buf<loc=vec, dtype=f32, rows=1, cols=1024, v_row=?, v_col=?, blayout=row_major, slayout=none_box, fractal=512, pad=0>, count=3>

module attributes {pto.target_arch = "a2a3"} {
  func.func @vec_add_binary_incore_0(%arg0: !pto.ptr<f32>, %arg1: !pto.ptr<f32>, %arg2: !pto.ptr<f32>) attributes {pto.kernel_kind = #pto.kernel_kind<vector>} {
  %c0_i64 = arith.constant 0 : i64
  %c12288_i64 = arith.constant 12288 : i64
  %c4096_i64 = arith.constant 4096 : i64
  %c16384_i64 = arith.constant 16384 : i64
  %c8192_i64 = arith.constant 8192 : i64
  %c20480_i64 = arith.constant 20480 : i64
  %c1_index = arith.constant 1 : index
  %c1024_index = arith.constant 1024 : index
  %c64_index = arith.constant 64 : index
  %c0_index = arith.constant 0 : index
  %c62_index = arith.constant 62 : index
  %c2_index = arith.constant 2 : index
  %c3_index = arith.constant 3 : index
  %cst_13 = arith.constant 1.00000000000000000e+00 : f32
  %c63_index = arith.constant 63 : index
  %y__ssa_v0_view = pto.make_tensor_view %arg0, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
  %x__ssa_v0_view = pto.make_tensor_view %arg1, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
  %z__ssa_v0_view = pto.make_tensor_view %arg2, shape = [%c64_index, %c1024_index], strides = [%c1024_index, %c1_index] {layout = #pto.layout<nd>} : !view
  %pipe_a__tile_0_mb = pto.alloc_multi_tile addr = %c0_i64 valid_row = %c1_index valid_col = %c1024_index : !ring
  %pipe_b__tile_1_mb = pto.alloc_multi_tile addr = %c12288_i64 valid_row = %c1_index valid_col = %c1024_index : !ring
  %a__tile = pto.multi_tile_get %pipe_a__tile_0_mb[%c0_index] : !ring -> !tile
  %x__ssa_v0_pview = pto.partition_view %x__ssa_v0_view, offsets = [%c0_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tload ins(%x__ssa_v0_pview : !part) outs(%a__tile : !tile)
  %b__tile = pto.multi_tile_get %pipe_b__tile_1_mb[%c0_index] : !ring -> !tile
  %z__ssa_v0_pview = pto.partition_view %z__ssa_v0_view, offsets = [%c0_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tload ins(%z__ssa_v0_pview : !part) outs(%b__tile : !tile)
  %0 = pto.multi_tile_get %pipe_a__tile_0_mb[%c1_index] : !ring -> !tile
  %14 = pto.partition_view %x__ssa_v0_view, offsets = [%c1_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tload ins(%14 : !part) outs(%0 : !tile)
  %1 = pto.multi_tile_get %pipe_b__tile_1_mb[%c1_index] : !ring -> !tile
  %15 = pto.partition_view %z__ssa_v0_view, offsets = [%c1_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tload ins(%15 : !part) outs(%1 : !tile)
  scf.for %i__idx_v0_pipe = %c0_index to %c62_index step %c1_index {
    %16 = arith.addi %i__idx_v0_pipe, %c2_index : index
    %17 = arith.remui %16, %c3_index : index
    %2 = pto.multi_tile_get %pipe_a__tile_0_mb[%17] : !ring -> !tile
    %18 = arith.addi %i__idx_v0_pipe, %c2_index : index
    %19 = arith.maxsi %18, %c0_index : index
    %20 = pto.partition_view %x__ssa_v0_view, offsets = [%19, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%20 : !part) outs(%2 : !tile)
    %21 = arith.addi %i__idx_v0_pipe, %c2_index : index
    %22 = arith.remui %21, %c3_index : index
    %3 = pto.multi_tile_get %pipe_b__tile_1_mb[%22] : !ring -> !tile
    %23 = arith.addi %i__idx_v0_pipe, %c2_index : index
    %24 = arith.maxsi %23, %c0_index : index
    %25 = pto.partition_view %z__ssa_v0_view, offsets = [%24, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tload ins(%25 : !part) outs(%3 : !tile)
    %26 = arith.remui %i__idx_v0_pipe, %c3_index : index
    %4 = pto.multi_tile_get %pipe_a__tile_0_mb[%26] : !ring -> !tile
    %27 = arith.remui %i__idx_v0_pipe, %c3_index : index
    %5 = pto.multi_tile_get %pipe_b__tile_1_mb[%27] : !ring -> !tile
    pto.tadd ins(%4, %5 : !tile, !tile) outs(%4 : !tile)
    pto.tadds ins(%4, %cst_13 : !tile, f32) outs(%4 : !tile)
    %28 = arith.maxsi %i__idx_v0_pipe, %c0_index : index
    %y__iter_v1_pview = pto.partition_view %y__ssa_v0_view, offsets = [%28, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
    pto.tstore ins(%4 : !tile) outs(%y__iter_v1_pview : !part)
  }
  %6 = pto.multi_tile_get %pipe_a__tile_0_mb[%c2_index] : !ring -> !tile
  %7 = pto.multi_tile_get %pipe_b__tile_1_mb[%c2_index] : !ring -> !tile
  pto.tadd ins(%6, %7 : !tile, !tile) outs(%6 : !tile)
  pto.tadds ins(%6, %cst_13 : !tile, f32) outs(%6 : !tile)
  %y__rv_v2_steady_pview = pto.partition_view %y__ssa_v0_view, offsets = [%c62_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tstore ins(%6 : !tile) outs(%y__rv_v2_steady_pview : !part)
  %10 = pto.multi_tile_get %pipe_a__tile_0_mb[%c0_index] : !ring -> !tile
  %11 = pto.multi_tile_get %pipe_b__tile_1_mb[%c0_index] : !ring -> !tile
  pto.tadd ins(%10, %11 : !tile, !tile) outs(%10 : !tile)
  pto.tadds ins(%10, %cst_13 : !tile, f32) outs(%10 : !tile)
  %y__tile_pview = pto.partition_view %y__ssa_v0_view, offsets = [%c63_index, %c0_index], sizes = [%c1_index, %c1024_index] : !view -> !part
  pto.tstore ins(%10 : !tile) outs(%y__tile_pview : !part)
  return
  }
}

Expected performance

Preserve the required RAW/WAR/WAW dependencies, including slot reuse, preload, and drain, without exhausting event IDs or inserting PIPE_ALL inside this loop. A final completion barrier is expected.

A sufficient dependency graph is:

LOAD_X(i) --------------------> ADD(i)
LOAD_Z(i) --------------------> ADD(i)
ADD(i) -> ADDS(i) -> STORE_X(i) -> LOAD_X(i + 3)
ADD(i) ------------------------> LOAD_Z(i + 3)

Consequently, several retained direct edges are transitively covered:

Retained group Covering path for this input
LOAD_X -> ADDS (idx=11) LOAD_X -> ADD -> ADDS
ADD/ADDS -> next LOAD_X (idx=5/4) ADD -> ADDS -> STORE_X -> next LOAD_X
Previous STORE_X -> ADD/ADDS (idx=9/12) Previous STORE_X -> LOAD_X -> ADD -> ADDS
LOAD_X -> STORE_X (idx=13) LOAD_X -> ADD -> ADDS -> STORE_X

The following complete, hand-written expected C++ reference demonstrates a sufficient event protocol. It is not output from a modified PTOAS compiler. It keeps the software input's 24 KiB UB layout, in-place operations, three slots, and two-tile preload, and uses:

Dependency Pipe pair Event IDs
X load ready MTE2 -> V 0–2
Z load ready MTE2 -> V 3–5
Z slot free after ADD V -> MTE2 0–2
Result ready for store V -> MTE3 3–5
X slot free after STORE MTE3 -> MTE2 0–2

Only free-slot tokens are primed. Ready tokens come from real loads/computes; remaining free-slot tokens are drained at exit.

expected.cpp — complete reference, compiled and simulator-checked
// MANUAL EXPECTED REFERENCE: this file was not emitted by PTOAS.
// Workload: y = (x + z) + 1, 64 rows of 1024 FP32 values, slots=3, prefetch=2.
// Change kTrips, and the sibling .pto shape metadata, together to test another N.
// UB layout matches the actual software-pipeline case exactly: X=[0,12288),
// Z=[12288,24576). No extra output or intermediate buffer is introduced.
#include "pto/pto-inst.hpp"

using namespace pto;

namespace {
constexpr int64_t kTrips = 64;
constexpr int64_t kWidth = 1024;
constexpr int64_t kSlots = 3;
constexpr int64_t kPrefetch = kSlots - 1;
constexpr uint64_t kSlotBytes = kWidth * sizeof(float);
constexpr uint64_t kZBase = kSlots * kSlotBytes;
static_assert(kTrips >= kPrefetch, "This reference requires at least two tiles");

#ifdef __DAV_VEC__
// Dynamic valid extents match the actual emitted C++ tile type; their runtime
// values remain the same constants, 1 and 1024.
using VecTile = Tile<TileType::Vec, float, 1, kWidth, BLayout::RowMajor, -1, -1, SLayout::NoneBox, 512,
                     PadValue::Null, CompactMode::Null>;
using RowTensor =
    GlobalTensor<float, Shape<1, 1, 1, 1, kWidth>, Stride<kWidth, kWidth, kWidth, kWidth, 1>, Layout::ND>;

static AICORE inline void LoadPair(__gm__ float* x, __gm__ float* z, int64_t tile) {
  const int64_t slot = tile % kSlots;
  const event_t slot_event = static_cast<event_t>(slot);
  const event_t z_ready_event = static_cast<event_t>(kSlots + slot);
  VecTile next_x(1, kWidth);
  VecTile next_z(1, kWidth);
  TASSIGN(next_x, static_cast<uint64_t>(slot) * kSlotBytes);
  TASSIGN(next_z, kZBase + static_cast<uint64_t>(slot) * kSlotBytes);
  RowTensor x_row(x + tile * kWidth);
  RowTensor z_row(z + tile * kWidth);

  // X remains live through STORE_X; Z becomes free after the first vector add.
  wait_flag(PIPE_MTE3, PIPE_MTE2, slot_event);
  TLOAD(next_x, x_row);
  set_flag(PIPE_MTE2, PIPE_V, slot_event);
  wait_flag(PIPE_V, PIPE_MTE2, slot_event);
  TLOAD(next_z, z_row);
  set_flag(PIPE_MTE2, PIPE_V, z_ready_event);
}

static AICORE inline void ComputeAndStore(__gm__ float* y, int64_t tile) {
  const int64_t slot = tile % kSlots;
  const event_t slot_event = static_cast<event_t>(slot);
  const event_t z_ready_event = static_cast<event_t>(kSlots + slot);
  VecTile x_s(1, kWidth);
  VecTile z_s(1, kWidth);
  TASSIGN(x_s, static_cast<uint64_t>(slot) * kSlotBytes);
  TASSIGN(z_s, kZBase + static_cast<uint64_t>(slot) * kSlotBytes);
  RowTensor y_row(y + tile * kWidth);

  wait_flag(PIPE_MTE2, PIPE_V, slot_event);
  wait_flag(PIPE_MTE2, PIPE_V, z_ready_event);
  // One current-X handle is used for both updates and the final store.
  TADD(x_s, x_s, z_s);
  set_flag(PIPE_V, PIPE_MTE2, slot_event);
  pipe_barrier(PIPE_V);
  TADDS(x_s, x_s, 1.0f);
  // Keep V->MTE3 IDs disjoint from the V->MTE2 free-Z IDs.
  const event_t store_ready_event = static_cast<event_t>(kSlots + slot);
  set_flag(PIPE_V, PIPE_MTE3, store_ready_event);
  wait_flag(PIPE_V, PIPE_MTE3, store_ready_event);
  TSTORE(y_row, x_s);
  set_flag(PIPE_MTE3, PIPE_MTE2, slot_event);
}
#endif
}  // namespace

AICORE void vec_add_binary_incore_0(__gm__ float* y, __gm__ float* x, __gm__ float* z) {
#if defined(__DAV_VEC__)
  set_mask_norm();
  set_vector_mask(-1, -1);

  // Prime exactly one free token per physical slot. Ready tokens are produced
  // only by real loads/computes, so no uninitialized tile can be consumed.
  for (int64_t slot = 0; slot < kSlots; ++slot) {
    const event_t event = static_cast<event_t>(slot);
    set_flag(PIPE_MTE3, PIPE_MTE2, event);
    set_flag(PIPE_V, PIPE_MTE2, event);
  }
  for (int64_t tile = 0; tile < kPrefetch; ++tile) {
    LoadPair(x, z, tile);
  }

  // The same preload / load-next / compute-current / drain schedule as the PTO.
  for (int64_t tile = 0; tile < kTrips - kPrefetch; ++tile) {
    LoadPair(x, z, tile + kPrefetch);
    ComputeAndStore(y, tile);
  }
  for (int64_t tile = kTrips - kPrefetch; tile < kTrips; ++tile) {
    ComputeAndStore(y, tile);
  }

  // Each load consumes one free token and each completed use replaces it.
  // Thus exactly one free X and Z token remains per slot at kernel exit.
  // All load-ready and store-ready tokens have already been consumed.
  for (int64_t slot = 0; slot < kSlots; ++slot) {
    const event_t event = static_cast<event_t>(slot);
    wait_flag(PIPE_MTE3, PIPE_MTE2, event);
    wait_flag(PIPE_V, PIPE_MTE2, event);
  }
  pipe_barrier(PIPE_ALL);  // Final completion only; no PIPE_ALL in steady state.
#endif
}

// Why the omitted edges are covered, including preload and drain boundaries:
//   STORE_X(i-3) -> LOAD_X(i) -> TADD(i) -> TADDS(i) -> STORE_X(i).
//   TADD(i-3) -> LOAD_Z(i) -> TADD(i).
// The first use of each slot consumes its seeded free token instead of the
// predecessor edge. Loads 0/1 are issued in the prologue; every later load is
// issued once by the steady loop. The final two computes consume their actual
// ready tokens in the drain. This covers MTE3->V and MTE2->MTE3 transitively;
// the V barrier orders TADD->TADDS without a second MTE2->V event family.

This reference passes exact FP32 equality for all 65,536 elements at N=64, and also for all 66,560 elements at N=65 with kTrips=65 and matching buffer sizes. The latter checks a different slot position at the drain boundary.

For N=64 it takes 18,831 simulator cycles, close to the 19,064-cycle unroll baseline. Its scalar/address-expression code also differs from PTOAS output, so this is evidence that an efficient schedule exists, not an isolated measurement of a sync-only compiler patch or a requirement for exact emitted C++ text.

Actual performance

The following groups remain in software_pipeline.sync.log under After Remove Redundant Sync:

Pipe pair Group indices Requested rotating event IDs
MTE2 -> V 7, 8, 11 3 + 3 + 3 = 9
V -> MTE2 4, 5, 6 3 + 3 + 3 = 9
MTE3 -> V 9, 12 3 + 3 = 6
MTE3 -> MTE2 3 3
MTE2 -> MTE3 13 3

There is also the non-rotating V -> MTE3 store-ready group (idx=2). The table describes PTOAS's retained demand, not the algorithm's minimum requirements. The v0.63 allocator defines an eight-ID event pool; see SyncEventIdAllocation.h.

After event allocation, groups 6 and 11 become full barriers. Relevant lines copied from After EventId Allocation:

  [   6] COMPOUND pto.tload [PIPE_MTE2]
    def=[%4(VEC)]
    use=[%arg2(GM)]
    PRE : pipe_barrier <PIPE_ALL -> PIPE_ALL> idx=6 forEnd=10 eventIdNum=3
  [   8] COMPOUND pto.tadds [PIPE_V]
    def=[%3(VEC)]
    use=[%3(VEC)]
    PRE : pipe_barrier <PIPE_V -> PIPE_V> idx=1
    PRE : pipe_barrier <PIPE_ALL -> PIPE_ALL> idx=11 forEnd=10 eventIdNum=3
    PRE : wait_flag <PIPE_MTE3 -> PIPE_V> idx=12 forEnd=10 eventIdNum=3 eventIds=[3,4,5]

Their matching set operations are marked useless. The emitted C++ confirms the barrier positions, before the second load and between the two vector operations. These are actual excerpts from inside the loop:

    pipe_barrier(PIPE_ALL);
    TLOAD(v103, v110);
    pipe_barrier(PIPE_V);
    pipe_barrier(PIPE_ALL);
    event_t v140 = (event_t) v138;
    wait_flag(PIPE_MTE3, PIPE_V, v140);
    TADDS(v119, v119, v20);

These two barriers execute in each of the 62 steady-state iterations; they are separate from the normal kernel-tail completion barrier.

The requested improvement is to eliminate transitively covered dependencies before event allocation, while preserving slot lifetimes and correct priming/draining. Simply dropping the barrier fallback or collapsing a rotating dependency to one event is not a safe general replacement.

Profiling data (optional)

Measured with the Ascend operator simulator, one active AIV, CANN 9.0.0, Ascend910B1 / dav-c220. These are in-core simulated cycle spans, not hardware wall-clock or end-to-end latency measurements.

N=64 variant UB Steady-loop PIPE_ALL barriers per iteration Active AIV span, cycles Exact correctness
PTOAS unroll baseline 24 KiB 0 19,064 PASS
PTOAS rotating software pipeline 24 KiB 2 74,909 PASS
Manual expected C++ above 24 KiB 0 18,831 PASS

The current software output is 3.93x the unroll baseline's cycle span. All three traces contain exactly 128 MTE2 loads, 64 VADD, 64 VADDS, and 64 MTE3 stores.

Inputs are generated once with NumPy np.random.seed(19) followed by two sequential np.random.random(65536).astype(np.float32) calls, and reused byte-for-byte. Output is initialized to NaN; the host checks every element using FP32 expected = x + z; expected += 1.0f and exact equality.

X SHA256: b373d6442009c54d11e2f2c83f0c88e81822e35d187c50542e49643f6bd0357f
Z SHA256: aaa012d7e19f43f34bab677b12de7e244cb51c3214ffb2365a316471825677ed

Cycle span is computed from core0.veccore0/trace.json: keep complete (ph="X") instruction events, excluding the process whose process_name is CACHEMISS, and calculate (max(ts + dur) - min(ts)) * 1850. This includes scalar setup, synchronization, and tail instructions. All rows use the same measurement convention.

Generated/reference C++ SHA256 values:

unroll.cpp:            4175c1586abcdd2fd5757c0a12fb166d6c1e82fbe409173b01d2b96ad61b5216
software_pipeline.cpp: 44f7cac9090ec6febc770d36b997dfcedae5dc57c313302c7536c46e5bbc748f
expected.cpp:          197e140b1144102903b54ece2e161d51ecbce3d30989866bf5e136fff4ef1a8b

Git commit

PTOAS v0.63, release tag commit 470eab4; the tested binary reports ptoas 0.63.

PTO-ISA used to compile the simulator cases: 3b4faf67aebb3e0d41be7952c56908b3adba7a8f.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions