Summary
在 packed UE8M0 scale-factor packing 中,vcmax(group=4) 的结果是 grouped 4xui16,随后需要做 width-changing vinterpret_cast:
目的是把两个 8-bit exponent 拼成一个 16-bit packed word,再 vcvt 回 4xui16 写入 SF UB。
当前 PTOAS 报 layout contract 失败:width-changing bitcast 需要 source 为 contiguous,但 source 从 group reduce 继承的是:
num_groups = 4, slots = 1
ensure_layout 没有 grouped -> contiguous 的 materialization 规则,因此编译失败。
Minimal PTODSL shape
mask128 = pto.vmi.create_mask(128, size=128)
mask4 = pto.vmi.create_mask(1, size=4, group=4)
scale4 = pto.vmi.vcmax(amax128, mask128, group=4)
# same-width reinterpret: group layout is preserved
bits = pto.vmi.vinterpret_cast(scale4, to_dtype=pto.ui32)
# group layout is carried through the elementwise/integer conversion
exp16 = pto.vmi.vcvt(
pto.vmi.vshrs(bits, 23, mask4),
to_dtype=pto.ui16,
saturate="NOSAT",
)
# width-changing reinterpret: 4xui16 -> 2xui32
packed32 = pto.vmi.vshrs(
pto.vmi.vinterpret_cast(exp16, to_dtype=pto.ui32),
8,
pto.vmi.create_mask(2, size=2),
)
packed16 = pto.vmi.vcvt(packed32, to_dtype=pto.ui16, saturate="NOSAT")
The failing operation is the second vinterpret_cast:
%packed32 = pto.vmi.vinterpret_cast %exp16
: !pto.vmi.vreg<4xui16, #pto.vmi.layout<num_groups = 4, slots = 1>>
-> !pto.vmi.vreg<2xui32>
Actual error
VMI-LAYOUT-CONTRACT: pto.vmi.bitcast operand #0 has type
!pto.vmi.vreg<4xui16, #pto.vmi.layout<num_groups = 4, slots = 1>>
but requires
!pto.vmi.vreg<4xui16, #pto.vmi.layout<contiguous>>;
pto.vmi.ensure_layout has no registered materialization support:
source/result layouts do not match a supported ensure_layout table row
Environment
- PTOAS
0.66 (v0.66-8-gb729c955c)
- TileKernels
per_block_cast VMI path
- Reduce/bitcast shape:
128xf32 -> vcmax(group=4) -> 4xf32 -> ... -> 4xui16 -> 2xui32
Expected behavior
Allow a grouped 4xui16 value to feed a width-changing vinterpret_cast to 2xui32, either by:
- adding a legal
num_groups=4, slots=1 -> contiguous materialization path, or
- recognizing this specific width-changing bitcast as layout-preserving where physically valid.
The former is the direct fix; the latter may be a useful optimization if the grouped physical layout already matches the required byte layout.
Workaround
Avoiding the width-changing vinterpret_cast and packing the two exponents entirely in the original 16-bit domain. This is not implemented in the kernel because it changes the VMI dataflow and may affect the desired ASC-aligned form.
Related issues
Summary
在 packed UE8M0 scale-factor packing 中,
vcmax(group=4)的结果是 grouped4xui16,随后需要做 width-changingvinterpret_cast:目的是把两个 8-bit exponent 拼成一个 16-bit packed word,再
vcvt回4xui16写入 SF UB。当前 PTOAS 报 layout contract 失败:width-changing bitcast 需要 source 为
contiguous,但 source 从 group reduce 继承的是:ensure_layout没有grouped -> contiguous的 materialization 规则,因此编译失败。Minimal PTODSL shape
The failing operation is the second
vinterpret_cast:Actual error
Environment
0.66(v0.66-8-gb729c955c)per_block_castVMI path128xf32 -> vcmax(group=4) -> 4xf32 -> ... -> 4xui16 -> 2xui32Expected behavior
Allow a grouped
4xui16value to feed a width-changingvinterpret_castto2xui32, either by:num_groups=4, slots=1 -> contiguousmaterialization path, orThe former is the direct fix; the latter may be a useful optimization if the grouped physical layout already matches the required byte layout.
Workaround
Avoiding the width-changing
vinterpret_castand packing the two exponents entirely in the original 16-bit domain. This is not implemented in the kernel because it changes the VMI dataflow and may affect the desired ASC-aligned form.Related issues
ensure_layoutmaterialization.group_reduce_maxf(group=4); would avoid producing the grouped source in the first place.group_reduce_maxf(group=4).