Repository navigation
undefined reference to `__truncsfbf2' when compiling ir with clang #97896
Description
Activity
- addedclangClang issues not falling into any other categoryClang issues not falling into any other category
on Jul 6, 2024 You stated no CPU version or a request for SSE2 in godbolt.
You stated no CPU version or a request for SSE2 in godbolt.
In the configuration below, I have specified the CPU and enabled SSE, but why am I still getting an error? I would appreciate it if you could help me.
--target=x86_64-unknown-linux-gnu -mcpu=x86-64 -march=native -msse -msse2 -mfpmath=sseThis error is likely due to the absence of the
compiler-rtruntime library in your build of the LLVM project. The issue can be resolved by building the LLVM project withcompiler-rtenabled. Below are the generic steps to build the LLVM project with compiler-rt and compile your program without encountering this error:git clone <llvm-project-link> cd llvm-project mkdir build-compiler-rt cd build-compiler-rt cmake ../llvm \ -DLLVM_ENABLE_PROJECTS="clang;compiler-rt" \ -DCMAKE_BUILD_TYPE=Release \ -DCMAKE_INSTALL_PREFIX=/path/to/llvm-install make -j<N> make installCompile the program using the built Clang and specify compiler-rt as the runtime library:
/path/to/build-compiler-rt/bin/clang -o <output> -rtlib=compiler-rt --target=x86_64-unknown-linux-gnu -march=native <program(.c/.ll/etc..)>
These steps ensure that thecompiler-rtruntime is available, which supports the_bf16type and resolves the error.This is actually an LLVM bug. LLVM inserting a floating point format conversion here (via the call to
__truncsfbf2) where there is none in the IR will quieten signalling NaNs, changing the bfloat's bit pattern and causing a miscompilation compilation.Reacted by Michael R. CrusoeSimpler repro:
define internal i16 @foo(bfloat %x) unnamed_addr { %r = bitcast bfloat %x to i16 ret i16 %r }
foo: push rax pslld xmm0, 16 call __truncsfbf2@PLT pextrw eax, xmm0, 0 pop rcx ret
This is a miscompilation similar to #97981, bitcasts need to be libcall-free. Checking a few targets at https://llvm.godbolt.org/z/r574feeT4, it seems like this may be limited to x86 (at least for targets where bfloat doesn't crash completely).
- addedfloating-pointFloating-point mathFloating-point mathand removedclangClang issues not falling into any other categoryClang issues not falling into any other category
on Sep 19, 2026 llvmorg-github-actions commented
on Sep 19, 2026 More actions@llvm/issue-subscribers-backend-x86
Author: DigOrDog
Details
# Description When I was compiling this IR with clang, I encountered this error. It seems to be right from a syntactic perspective, which is quite strange. From [half-precision-floating-point](https://clang.llvm.org/docs/LanguageExtensions.html#half-precision-floating-point) document, _bf16 is supported on X86 (when SSE2 is available)(For X86, SSE2 is available on 64-bit and all recent 32-bit processors.)define i64 @<!-- -->f(bfloat %CastFPTrunc, ptr %x) { BB: store bfloat %CastFPTrunc, ptr %x, align 2 ret i64 0 } define i32 @<!-- -->main() { entry: ret i32 0 }
Seems to be a fast-isel issue? https://llvm.godbolt.org/z/c5G68dWaT
bb.0 (%ir-block.0): liveins: $xmm0 %0:fr16 = COPY $xmm0 %2:vr128 = COPY %0:fr16 %3:vr128 = PSLLDri %2:vr128(tied-def 0), 16 %4:fr32 = COPY killed %3:vr128 %1:fr32 = COPY %4:fr32 ADJCALLSTACKDOWN64 0, 0, 0, implicit-def dead $rsp, implicit-def dead $eflags, implicit-def dead $ssp, implicit $rsp, implicit $ssp $xmm0 = COPY %1:fr32 CALL64pcrel32 target-flags(x86-plt) &__truncsfbf2, <regmask $bh $bl $bp $bph $bpl $bx $ebp $ebx $hbp $hbx $rbp $rbx $r12 $r13 $r14 $r15 $r12b $r13b $r14b $r15b $r12bh $r13bh $r14bh $r15bh $r12d $r13d $r14d $r15d $r12w $r13w $r14w $r15w $r12wh and 3 more...>, implicit $rsp, implicit $ssp, implicit $xmm0, implicit-def $rsp, implicit-def $ssp, implicit-def $xmm0 ADJCALLSTACKUP64 0, 0, implicit-def dead $rsp, implicit-def dead $eflags, implicit-def dead $ssp, implicit $rsp, implicit $ssp %6:fr16 = COPY $xmm0 %7:vr128 = COPY %6:fr16 %8:gr32 = PEXTRWrri killed %7:vr128, 0 %5:gr16 = COPY %8.sub_16bit:gr32 $ax = COPY %5:gr16 RET64 implicit $ax
Looks like there's an extra fregs step that probably hits
, whereas the f16 version goes directly to GPRs:llvm-project/llvm/lib/CodeGen/TargetLoweringBase.cpp
Lines 309 to 310 in 534934c
if (OpVT == MVT::f32) return FPROUND_F32_BF16; bb.0 (%ir-block.0): liveins: $xmm0 %0:fr16 = COPY $xmm0 %2:vr128 = COPY %0:fr16 %3:gr32 = PEXTRWrri killed %2:vr128, 0 %1:gr16 = COPY %3.sub_16bit:gr32 $ax = COPY %1:gr16 RET64 implicit $ax
it seems like this may be limited to x86
Looks like (some) riscv is also affected https://llvm.godbolt.org/z/znjz4f7z3
And loongarch https://llvm.godbolt.org/z/5zGM9z5oe
Description
When I was compiling this IR with clang, I encountered this error. It seems to be right from a syntactic perspective, which is quite strange. From half-precision-floating-point document, _bf16 is supported on X86 (when SSE2 is available)(For X86, SSE2 is available on 64-bit and all recent 32-bit processors.)
https://godbolt.org/z/MbcMvcWe6