Skip to content

[Bug] PTODSL-vmi,vinterpret_cast 跨位宽小 VL 触发 VMI-LAYOUT-CONTRACT #1380

Description

@chenben09

Component

PTO Dialect / ODS (include/PTO/IR)

Description

在 A5 平台测试 pto.vmi.vinterpret_cast 的跨位宽(lane-count-changing)场景时,ISA 规格(docs/isa/vmi-isa/06-convert.md)声明:"Any T_src, T_dst ... with L · bitwidth(T_src) == L · bitwidth(T_dst)",且 00-architecture-overview.md 明确 compact/partial 向量(L·bw < 2048,即 sub-physical 小 VL)合法(单物理寄存器承载、低 L 槽有效)。因此小 VL 跨位宽 reinterpret 在规格层应受支持。

Reproduction (minimal)

@jit_kernel(name, params)
def kernel(
    src_ptr: pto.ptr(src_ptype, "gm"),
    out_ptr: pto.ptr(dst_ptype, "gm"),
):
    src_ub = pto.castptr(pto.i64(0), pto.ptr(src_ptype, "ub"))
    out_ub = pto.castptr(pto.i64(src_nbytes), pto.ptr(dst_ptype, "ub"))
    pto.mte_gm_ub(src_ptr, src_ub, 0, src_nbytes, nburst=(1, 0, 0))
    pto.set_flag("MTE2", "V", event_id=0)
    pto.wait_flag("MTE2", "V", event_id=0)
    source = pto.vmi.vload(src_ub, pto.const(0, dtype=pto.index), size=src_vl)
    result = pto.vmi.vinterpret_cast(source, to_dtype)
    pto.vmi.vstore(result, out_ub, pto.const(0, dtype=pto.index))
    pto.set_flag("V", "MTE3", event_id=0)
    pto.wait_flag("V", "MTE3", event_id=0)


触发参数示例:`src=f16, dst=f32, vl=4(dst lanes), src_vl=8`(16→32 lane-halving);或 `src=f32, dst=f16, vl=4, src_vl=2`(32→16 lane-doubling)。只要 src/dst 位宽不同且 vl 不填满物理寄存器即触发。

Expected behavior

编译通过 + 运行通过

Actual behavior / error logs

loc("3"("kernel.mlir":18:12)): error: VMI-LAYOUT-CONTRACT: pto.vmi.bitcast operand #0 has type
  !pto.vmi.vreg<8xf16, #pto.vmi.layout<num_groups = 8, slots = 8>>
  but requires !pto.vmi.vreg<8xf16, #pto.vmi.layout<contiguous>>;
  pto.vmi.ensure_layout has no registered materialization support:
  source/result layouts do not match a supported ensure_layout table row
Error: VPTO emission pipeline failed.


报错场景:
跨位宽小 VL(vl2/4/8,非 bf16)共 46/46 FAIL;跨位宽大 VL(vl64,contiguous)全 PASS;同位宽(任意 VL)全 PASS。compile(ir_check)全 PASS,失败只发生在 PTOAS lowering(ptoas MLIR→device .o)阶段。

前端 MLIR(compile 产物,layout-less,合法):

%2 = pto.vmi.vload %0[%c0] : !pto.ptr<f16, ub> -> !pto.vmi.vreg<8xf16>
%3 = pto.vmi.vinterpret_cast %2 : !pto.vmi.vreg<8xf16> -> !pto.vmi.vreg<4xf32>

VPTO emission 的 layout propagation pass 给小 VL vload 结果分配了 `layout<num_groups=N, slots=8>`(sub-physical 分组布局),而 bitcast 要求 `contiguous`。

部分失败用例:
f32_f16_vl4_u_a4_n_compact   # 32→16 lane-doubling
f16_f32_vl4_u_a4_n_compact   # 16→32 lane-halving
i32_i8_vl4_u_a4_n_compact    # 32→8
i8_i32_vl4_u_a4_n_compact    # 8→32
f16_i8_vl4_u_a4_n_compact    # 16→8
i8_f16_vl4_u_a4_n_compact    # 8→16

Git commit

10883a8

Host platform

Linux (aarch64)

Target Ascend arch (if relevant)

a5

PTOAS build level (if relevant)

None

Metadata

Metadata

Labels

QAIssues from the QA teambugSomething isn't workingvmi

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions