@jit_kernel(name, params)
def kernel(
src_ptr: pto.ptr(src_ptype, "gm"),
out_ptr: pto.ptr(dst_ptype, "gm"),
):
src_ub = pto.castptr(pto.i64(0), pto.ptr(src_ptype, "ub"))
out_ub = pto.castptr(pto.i64(src_nbytes), pto.ptr(dst_ptype, "ub"))
pto.mte_gm_ub(src_ptr, src_ub, 0, src_nbytes, nburst=(1, 0, 0))
pto.set_flag("MTE2", "V", event_id=0)
pto.wait_flag("MTE2", "V", event_id=0)
source = pto.vmi.vload(src_ub, pto.const(0, dtype=pto.index), size=src_vl)
result = pto.vmi.vinterpret_cast(source, to_dtype)
pto.vmi.vstore(result, out_ub, pto.const(0, dtype=pto.index))
pto.set_flag("V", "MTE3", event_id=0)
pto.wait_flag("V", "MTE3", event_id=0)
触发参数示例:`src=f16, dst=f32, vl=4(dst lanes), src_vl=8`(16→32 lane-halving);或 `src=f32, dst=f16, vl=4, src_vl=2`(32→16 lane-doubling)。只要 src/dst 位宽不同且 vl 不填满物理寄存器即触发。
loc("3"("kernel.mlir":18:12)): error: VMI-LAYOUT-CONTRACT: pto.vmi.bitcast operand #0 has type
!pto.vmi.vreg<8xf16, #pto.vmi.layout<num_groups = 8, slots = 8>>
but requires !pto.vmi.vreg<8xf16, #pto.vmi.layout<contiguous>>;
pto.vmi.ensure_layout has no registered materialization support:
source/result layouts do not match a supported ensure_layout table row
Error: VPTO emission pipeline failed.
报错场景:
跨位宽小 VL(vl2/4/8,非 bf16)共 46/46 FAIL;跨位宽大 VL(vl64,contiguous)全 PASS;同位宽(任意 VL)全 PASS。compile(ir_check)全 PASS,失败只发生在 PTOAS lowering(ptoas MLIR→device .o)阶段。
前端 MLIR(compile 产物,layout-less,合法):
%2 = pto.vmi.vload %0[%c0] : !pto.ptr<f16, ub> -> !pto.vmi.vreg<8xf16>
%3 = pto.vmi.vinterpret_cast %2 : !pto.vmi.vreg<8xf16> -> !pto.vmi.vreg<4xf32>
VPTO emission 的 layout propagation pass 给小 VL vload 结果分配了 `layout<num_groups=N, slots=8>`(sub-physical 分组布局),而 bitcast 要求 `contiguous`。
部分失败用例:
f32_f16_vl4_u_a4_n_compact # 32→16 lane-doubling
f16_f32_vl4_u_a4_n_compact # 16→32 lane-halving
i32_i8_vl4_u_a4_n_compact # 32→8
i8_i32_vl4_u_a4_n_compact # 8→32
f16_i8_vl4_u_a4_n_compact # 16→8
i8_f16_vl4_u_a4_n_compact # 8→16
Component
PTO Dialect / ODS (include/PTO/IR)
Description
在 A5 平台测试
pto.vmi.vinterpret_cast的跨位宽(lane-count-changing)场景时,ISA 规格(docs/isa/vmi-isa/06-convert.md)声明:"Any T_src, T_dst ... with L · bitwidth(T_src) == L · bitwidth(T_dst)",且00-architecture-overview.md明确 compact/partial 向量(L·bw < 2048,即 sub-physical 小 VL)合法(单物理寄存器承载、低 L 槽有效)。因此小 VL 跨位宽 reinterpret 在规格层应受支持。Reproduction (minimal)
@jit_kernel(name, params) def kernel( src_ptr: pto.ptr(src_ptype, "gm"), out_ptr: pto.ptr(dst_ptype, "gm"), ): src_ub = pto.castptr(pto.i64(0), pto.ptr(src_ptype, "ub")) out_ub = pto.castptr(pto.i64(src_nbytes), pto.ptr(dst_ptype, "ub")) pto.mte_gm_ub(src_ptr, src_ub, 0, src_nbytes, nburst=(1, 0, 0)) pto.set_flag("MTE2", "V", event_id=0) pto.wait_flag("MTE2", "V", event_id=0) source = pto.vmi.vload(src_ub, pto.const(0, dtype=pto.index), size=src_vl) result = pto.vmi.vinterpret_cast(source, to_dtype) pto.vmi.vstore(result, out_ub, pto.const(0, dtype=pto.index)) pto.set_flag("V", "MTE3", event_id=0) pto.wait_flag("V", "MTE3", event_id=0) 触发参数示例:`src=f16, dst=f32, vl=4(dst lanes), src_vl=8`(16→32 lane-halving);或 `src=f32, dst=f16, vl=4, src_vl=2`(32→16 lane-doubling)。只要 src/dst 位宽不同且 vl 不填满物理寄存器即触发。Expected behavior
编译通过 + 运行通过
Actual behavior / error logs
Git commit
10883a8
Host platform
Linux (aarch64)
Target Ascend arch (if relevant)
a5
PTOAS build level (if relevant)
None