Repository navigation
RISC-V: "Illegal instruction" crash when vector instructions detected by V8 #64538
Description
Activity
- addedv8 engineIssues and PRs related to the V8 dependency.Issues and PRs related to the V8 dependency.riscv64Issues and PRs related to the 64-bit RISC-V architecture.Issues and PRs related to the 64-bit RISC-V architecture.
on Jul 16, 2026 That is likely not the only problem as even with that bypass of the vector instructions in that file enabled there are still
Illegal instructionfailures in some test cases e.g.=== release test-http-set-global-proxy-from-env-fetch === Path: client-proxy/test-http-set-global-proxy-from-env-fetch [CLOSE] null SIGILL node:internal/modules/run_main:107 triggerUncaughtException( ^Bypassing all of the RVV stuff in V8 (sxa@88497e2) allows all of the tests to complete without illegal instruction errors
I might be well off here since I've only had a few hours on a K3 (remote access back in March, benchmarking llama.cpp), but there's one thing from that session that might be worth ruling out.
The K3 is heterogeneous in a way that matters for RVV: the X100 cores are VLEN 256 and the A100 cores are VLEN 1024, on the same chip. I measured both, and RVV code ran fine on each of them, so it isn't that the vector unit is missing or broken.
What I never tested is what happens when a long-running process migrates between the two clusters. If V8 probes VLEN once at startup, or if vector state doesn't survive migration between cores with different VLEN, then a later vector instruction could trap. That would fit
vmv.s.x v24, x19being plain baseline RVV 1.0 rather than anything exotic, and it would also fit the standalone Highway suite passing, since a short test is far less likely to migrate than a full npm run.If it's worth five minutes, pinning with taskset to X100-only cores and then A100-only should settle it. If the SIGILL survives pinning, I'm wrong and it really is codegen.
You've spent a lot more time inside this than I have, so there's a good chance you've already ruled it out.
If it's worth five minutes, pinning with taskset to X100-only cores and then A100-only should settle it. If the SIGILL survives pinning, I'm wrong and it really is codegen.
Yeah that was one of my first thoughts too, but pinning to A100 didnt' seem to change the behaviour. I'll give it another try tomorrow just to confirm though!
Confirmed again that it does still crash with
Illegal Instructionon the A100 cores, albeit with slightly different output as it includes a bit of stack trace - here's the output since I hadn't logged it previously:$ ./npm Illegal instruction $ systemd-run --scope -p AllowedCPUs=8-15 ./npm ==== AUTHENTICATING FOR org.freedesktop.systemd1.manage-units ==== [... ==== AUTHENTICATION COMPLETE ==== Running as unit: run-p844737-i844738.scope; invocation ID: ecab0467508f437a9dc23e002c27f250 # # Fatal error in , line 0 # unimplemented code # # # #FailureMessage Object: 0x3fe36f2e98 ----- Native stack trace ----- 1: 0x2ac04adb36 [npm] 2: 0x2ac1ec2816 V8_Fatal(char const*, ...) [npm] 3: 0x2ac12359c0 [npm] 4: 0x2ac0caa598 [npm] 5: 0x2ac0ca928c [npm] 6: 0x2ac0ca899a [npm] 7: 0x2ac0cdd770 [npm] 8: 0x2ac0cdda24 [npm] 9: 0x2ac0cdb45c [npm] 10: 0x2ac0cdb318 [npm] 11: 0x2ac0cfe3ee [npm] 12: 0x2ac139ab04 [npm] Illegal instructionConfirmed again that it does still crash with
Illegal Instructionon the A100 cores, albeit with slightly different output as it includes a bit of stack trace - here's the output since I hadn't logged it previously:$ ./npm Illegal instruction $ systemd-run --scope -p AllowedCPUs=8-15 ./npm ==== AUTHENTICATING FOR org.freedesktop.systemd1.manage-units ==== [... ==== AUTHENTICATION COMPLETE ==== Running as unit: run-p844737-i844738.scope; invocation ID: ecab0467508f437a9dc23e002c27f250 # # Fatal error in , line 0 # unimplemented code # # # #FailureMessage Object: 0x3fe36f2e98 ----- Native stack trace ----- 1: 0x2ac04adb36 [npm] 2: 0x2ac1ec2816 V8_Fatal(char const*, ...) [npm] 3: 0x2ac12359c0 [npm] 4: 0x2ac0caa598 [npm] 5: 0x2ac0ca928c [npm] 6: 0x2ac0ca899a [npm] 7: 0x2ac0cdd770 [npm] 8: 0x2ac0cdda24 [npm] 9: 0x2ac0cdb45c [npm] 10: 0x2ac0cdb318 [npm] 11: 0x2ac0cfe3ee [npm] 12: 0x2ac139ab04 [npm] Illegal instructionThis method is incorrect. The vendor kernel has hacked the vector context-related parts to enable different vlen sizes, so running RVV on an A100 may not work.
Another issue is that scheduling may fail on CPUs 8-15, where the behavior is controlled by the HMP[4] module.1: spacemit-com/linux-6.18@ce988b8
2: spacemit-com/linux-6.18@33024f2
3: spacemit-com/linux-6.18@2cb5950
4: spacemit-com/linux-6.18@6968f3eThat stack trace helps, and I think it points somewhere slightly different from "instructions incompatible with the K3".
Frame 2 is
V8_Fatal, and the message isunimplemented code. That string is V8's own:kUnimplementedCodeMessageindeps/v8/src/base/logging.h:66, with#define UNIMPLEMENTED() FATAL(::v8::base::kUnimplementedCodeMessage)right below it at line 70. So V8 is deliberately aborting, and theIllegal instructionyou see is just howV8_Fatalends the process on riscv64 rather than the CPU rejecting anything.There is one
UNIMPLEMENTED()in the riscv64 backend that RVV codegen reaches. Indeps/v8/src/codegen/riscv/assembler-riscv.h,VU::SetSimd128(line 625),SetSimd128Half(661) andSetSimd128x2(690) each switch onCpuFeatures::vlen()with cases for 128, 256 and 512, thenstatic_assert(kMaxRvvVLEN <= 512)andUNIMPLEMENTED()in the default.kMaxRvvVLENis 512, indeps/v8/src/codegen/riscv/base-constants-riscv.h:343. The value is not a build-time assumption either:deps/v8/src/base/cpu.ccprobes the hardware at startup with__riscv_vlenb()(line 64, used at line 1091).When I had remote access to a K3 back in March, the A100 cores measured VLEN 1024. If the probe returns 1024, no case matches and the first SIMD path taken aborts. That would fit everything in this thread: the K1 at 256 is fine, your commit disabling the RVV paths fixes it, and it has nothing to do with which cluster the process ends up on.
It would also explain the qemu result. As far as I know qemu defaults to VLEN 128 unless you pass
vlen=explicitly, so it would take the first case and never reach the abort. Not reproducing under emulation is what this theory predicts rather than evidence against it.@RevySR, thank you for the kernel pointers, and that also means the test I suggested earlier was a bad one: if HMP controls placement then pinning proves nothing either way, and sxa's A100 run not changing anything was never evidence against VLEN.
Something that does not depend on scheduling: build a tiny program that prints
__riscv_vlenb() * 8and run it a number of times on that machine. If 1024 ever comes back, V8 has no case for it and the abort follows wherever the code lands. If it is always 256, I am wrong and this is ordinary codegen after all.I do not have any K3 hardware, the March session was remote access on someone else's board and mine is a few weeks out, so I cannot check this myself. You have both been inside this a lot longer than I have, so happy to be told I am chasing the wrong
UNIMPLEMENTED().Something that does not depend on scheduling: build a tiny program that prints __riscv_vlenb() * 8 and run it a number of times on that machine. If 1024 ever comes back, V8 has no case for it and the abort follows wherever the code lands. If it is always 256, I am wrong and this is ordinary codegen after all.
It consistently returns 256 on the (default) X100 cores and when run under
systemd-run --scope -p AllowedCPUs=8-15consistently returns 1024.happy to be told I am chasing the wrong UNIMPLEMENTED().
I wouldn't claim to be any more of a V8 expert at the moment than you, but that analysis sounds feasible ...
I ran a build with some crude debug sxa@6ad0dbd to clarify what V8 was identifying on the X100 cores in case it was incorrect:
23:14:12 RISC-V CPU Vector length detected as 256 23:14:12 supports_wasm_simd_128=1, vlen=256 23:14:12 RISC-V Extension zba=0,zbb=0,zbs=0,ZICOND=0(Probably not useful for anything else but if anyone wants the build with that patch in it to run elsewhere ping me on Slack - it's in the CI but that's not publicly visible)
@luyahan Do you have any advice/guides on setting up a V8 build environment on RISC-V (ideally natively - I'm struggling to get a suitable build environemnt for it at the moment)? Also are you aware of any problems like this on K3 hardware? Noting that Node.js is currently using V8 14.6
I ran a build overtnight with sxa@4ad3692 which hard codes the vlen to be overridden to 128 and that doesn't seem to crash.
$ npm i bson RISC-V CPU Vector length detected as 128 supports_wasm_simd_128=1, vlen=128 RISC-V Extension zba=0,zbb=0,zbs=0,ZICOND=0 npm warn cli npm v11.17.0 does not support Node.js v26.5.1-pre. This version of npm supports the following node versions: `^20.17.0 || >=22.9.0`. You can find the latest version at https://nodejs.org/. added 1 package in 1sNoting that I'm on vacation for a few days so I won't be doing anything more on this myself in the immediate future.
I got hold of a K1 and tried your reproducer on it, and I think the K1 is not actually immune, which changes the shape of this a bit.
Stock
node-v26.5.0-linux-riscv64-pointer-compressionfrom unofficial-builds, on a BananaPi F3 (SpaceMiT K1, VLEN 256).node -p "process.version"is fine.npm --versionandnpm i bsonboth die withIllegal instruction, exit 132.The kernel trap names the instruction for us:
cause: 2,badaddr: 000000004209ec57. Assembling that back,0x4209ec57isvmv.s.x v24, s3, and s3 is x19. Samevmv.s.x v24, x19you flagged at the top of this issue.statushas VS=Dirty, so the vector unit was on and in use.On the why, I think there is a second bug in
VU::SetSimd128, distinct from theUNIMPLEMENTED()that catches the A100. The VLEN 256 branch,deps/v8/src/codegen/riscv/assembler-riscv.h:639:case 256: lmul = (sew + 1) > kRvvELEN ? m1 : mf2;
kRvvELENis 64, a bit count (base-constants-riscv.h:338), butsewis aVSewenum index, so E8=0, E16=1, E32=2, E64=3. The comment right above the switch does the algebra in exponent form,elen >= sew + n, where elen for ELEN=64 is 3. So the guard wants 3, not 64. As written(sew + 1) > 64is never true, and mf2 gets picked for every SEW, E64 included. RVV 1.0 wantsSEW <= LMUL * ELEN, which at LMUL=1/2 caps SEW at 32, soe64, mf2is reserved: vill goes up and the next vector instruction traps.Measured on the K1, vsetvli then reading vtype back to back:
e8,mf2 vill=0 vl=16 e16,mf2 vill=0 vl=8 e32,mf2 vill=0 vl=4 e64,mf2 vill=1 vl=0 <-- the only one that goes vill e8,m1 vill=0 vl=32 e64,m1 vill=0 vl=4And the faulting instruction itself, under the two configurations:
e32,m1 + vmv.s.x v24,x19 : OK e64,mf2 + vmv.s.x v24,x19 : SIGILLSo
vmv.s.xis implemented perfectly well. It is the vtype it inherits that kills it.The same shape is in
SetSimd128Half(661) andSetSimd128x2(690). All six guards compare an enum index against 64. At VLEN 256 the Half variant lands on mf4, which breaks E32 as well as E64. It fits your vlen=128 override too, since the 128 case inSetSimd128is a barelmul = m1with no guard at all, so nothing reserved gets emitted. That would make your override a fix for two separate bugs at once.Which leaves why your K1 looked clean, and I would not assume your board and mine are in the same state. Your F3 logs over in the citgm thread show
Linux 6.1.15-legacy-k1, while mine are 6.6.99 and 6.18.33, both with V in HWCAP. If the older kernel does not expose V to userspace, V8 never turns RVV on and never emits the reserved config, and then the discriminator is "did V8 enable RVV" rather than which board you are sitting in front of. That is a guess on my part, but HWCAP on that F3, or the detected vlen line from your instrumented build, would settle it in a minute.Happy to share the small vsetvli probe if it is worth running on the K3, and to run anything else you want tested on a K1. You have been inside this far longer than I have, so if I have the wrong end of it, tell me and I will drop it.
Your theory on the kernel makes sense and I've just verified that my "verbose" build on the BPi doesn't print out the VLEN detection messages which indicates that it's not going down those paths. That's good to know. Great diagnosis on the other analysis too.
13 remaining items
Noting also that in the upstream V8 code the check on
__GLIBC_MINOR__as well as the fallback were removed in March along with the refactoring ofcpu.ccinto architecture specific files in v8/v8@89ad153 by @luyahan which was included as of V8 15.0 from what I can see.I'll also note that #65161 has been started to trial Node with V8 15.2
We have rvv issue on Debian builds, too, see https://bugs.debian.org/1144759.
Depending on the riscv processor, builds fail or pass.@kapouer Thanks for the extra data point, and given that you're also seeing that it's ok on a system with
vlen=128it sounds like it's probably the same problem.
The bug report references a build log but from what I can see the failures mentioned in the issue aren't in the log. Was that log from a non-RVV system? Thetest-inspector-network-content-typetest seems to pass ok in the log:ok 2120 parallel/test-inspector-network-content-type --- duration_ms: 2227.27900 ...Build logs: https://buildd.debian.org/status/logs.php?pkg=nodejs&arch=riscv64
we see failures on rv-manda-04, which is Milk-V Jupiter (SpacemiT M1)FYI I'm testing an AI assisted patch which looks like it might have resolved most of the crashes: sxa@4a7a7d2
Still trying to get a standalone V8 build environment working properly before I can look at testing and submitting anything upstream, but if anyone wants to take a look at the patch there it is :-)
Reacted by Bruno Verachten and Jérémy LalNoting that I've just successfully (I think) done a successful build using the V8 15.2 branch from #65161 and it seems to be running without any crashes so far in the test suite. Also seems to run the gouthar's test case and
npm install bsonwithout problems. Good news! Means we now need to consider what to do with 26 and below. Some of the files (cpu.ccin particular) have been refactored in the latest V8 so the recent changes cannot be trivially backported in general but since it has been at least partially fixed somehow (I'm assuming that it's running with RVV extensions enabled - that's something I'll need to confirm) we probably have a good path forward.EDIT: It does still crash on the A100 cores (with gounthar's example and npm) so some additional patching will be required regardless:
# # Fatal error with no security impact: # unimplemented code # ----- Native stack trace ----- 1: 0x2ae0391dc2 [./node] 2: 0x2ae1e6cbfe v8::base::FatalNoSecurityImpact(char const*, ...) [./node] 3: 0x2ae11b79da [./node] 4: 0x2ae10ae074 [./node] 5: 0x2ae10672d6 [./node] 6: 0x2ae1065db2 [./node] 7: 0x2ae10fae14 [./node] 8: 0x2ae10fdbde [./node] 9: 0x2ae1052c28 [./node] 10: 0x2ae1311314 [./node] Illegal instruction ./node gounthar.jsReacted by Bruno VerachtenPatch for the v8-152 branch that avoids the crash on A100 cores: sxa@1f85667
For those with CI access there's a build with this patch as the
node.bin.xzartifact on https://ci.nodejs.org/job/node-test-riscv64-experimental/217/label=zz-test-sxa-bianbu-riscv64-1/
The previous build (216) is thev8-152branch without the additional patch.Reporting a related-but-distinct riscv64 wasm SIGILL from a Lichee Pi 3A (SpacemiT K1, VLEN 256, kernel 6.1.15 → no
riscv_hwprobe) that may be worth cross-referencing:- On the non-pointer-compression unofficial-builds v24.20.0/v26.0.0, V8 here refuses SIMD modules at compile time (
CompileError: Wasm SIMD unsupported— soVU::SetSimd128*can't be reached), yet a plain non-SIMD Rust→wasm module (Next.jsswc-wasm-nodejsfallback) deterministically SIGILLs ~5–8 s after "Ready" with no traffic. - Kernel shows
badaddr 0x0and gdbra = Builtins_JSToWasmWrapperAsm+184, PC in a zero-filled code page — i.e. a wasm call jumping into an unpopulated lazy-compilation entry, not an unsupported RVV instruction word. --no-wasm-lazy-compilationeliminates it 100% (SIMD support untouched) — orthogonal to the SIMD/vlen fixes discussed here.
Full writeup: #65724.
Caveat / open item relevant to both threads: the pointer-compression flavor also SIGILLs on this board (so #64538's vlen/VLEN-1024 branch is not the only pc-flavor failure mode on K1 either), but I haven't yet captured its
badaddr— if it turns out to be a real RVV encoding it belongs here; I'll post the signature once I've re-flashed and rerun. The board is ssh-reachable and I'm glad to run experiments for either thread.- On the non-pointer-compression unofficial-builds v24.20.0/v26.0.0, V8 here refuses SIMD modules at compile time (
I've uploaded some additional patched builds for those who don't have access to the CI and want to experiment. They're in https://unofficial-builds.nodejs.org/download/test/riscv64/ (May not stay there, no guarantees/warranty etc. etc.)
node-rve216-v8152unpatched.xz- A version with V8 15.2 but no extra patchesnode-rve217-v8152patched.xz- Includes v8 15.2 but has the patch from RISC-V: "Illegal instruction" crash when vector instructions detected by V8 #64538 (comment)node-rve211-v8146patched.xz- Default version of V8 but with the patch from RISC-V: "Illegal instruction" crash when vector instructions detected by V8 #64538 (comment) (I think that's the right build I've uploaded!)
These are compressed versions of the node binary only, so overwrite
bin/nodeof one of the unofficial-builds tarballs if you want to test it in a normal environment.I now have a V8 15.3 standalone (unpatched by me other than to disable snapshots for reasons I won't go into here) build and it passes gouthar's test on the X100 cores but not the vlen=1024 A100 cores. gzipped binary (Because github won't take .xz ones) is here in case anyone wants to experiment with it : d8153-nosnapshot.gz. Like the node binaries in my last comment above this is just the
d8binary compressed.Confirmed that when applying the patches from the earlier comment designed for the v8-152 branch onto a standalone V8 15.3 gouthar's code no longer crashes on the A100 cores 👍🏻 Here's the binary: d8153-rvvpatches.gz
- added a commit that references this issue
on Sep 20, 2026 Ref the last comment the same is true for bleeding edge V8 so the fixes for A100 cores are still required there
Simplified patch without the separate helper function: rvv-vlenfix-v8-152.patch / https://github.com/sxa/node/commit/18f8188f5d8b487661186c9ccbdb0966cd7be741.patch
Debug version of the same patch: rvv-vlenfix-v8-152-with-debug.patch / https://github.com/sxa/node/commit/f8335c5e717cb568884859efafe3c83daaefbb6f.patch
Version
v26.5.0 (but any v26 and above is affected)
Platform
Subsystem
V8
What steps will reproduce the bug?
run
npmor anynode-gypon a SpacemiT-K3 board. Also fails in the build-addons step withmake run-cias that executesnode-gyp. Other boards such as the SpacemiT-K1 do not show the crash (Both should have vector instructions available)How often does it reproduce? Is there a required condition?
100% of the time
What is the expected behavior? Why is that the expected behavior?
There should be no crash when running npm
What do you see instead?
Illegal InstructionAdditional information
Based on the fact that sxa@1ecdd38 fixes it seems likely that this was introduced as part of the vector instructions introduced in v8/v8@de361fd - I haven't yet tracked it down to any specific operation in that function. At the moment I've also only determined that it fixes the initial crash and hopefully that doesn' uncover more issues.
FYI @luhenry @luyahan @nodejs/platform-riscv64
Split out from nodejs/build#4099 (comment)