Repository navigation
JVM SIGSEGV in vframe::java_sender via allocation-profiler JVMTI GetStackTrace on a virtual thread (JDK 25, Alpine/musl) — regression in 1.62.0 (1.60.1 OK) #11518
Description
Activity
Hi @soidisant ,
Thanks for the report! Yes, hs_err*.log would be useful.
FTR, we had some issues in 1.62.0 where concurrent access interleaved with arriving signals could have lead to use-after-free, hence memory corruption that could lead to spurious crashes like you are seeing.
Otherwise, if not caused by the UaF this would point to a JVM bug because we are using standard JVMTI API from a standard context, so it should not be crashing like that.
HI @jbachorik
Thanks for your response
I added the log to the original message (redacted package names and some private naming)Please, upgrade to 1.63.0 - we have a number of fixes in profiler that are targeting the concurrency issues and memory corruption. Internal testing and dogfooding is showing that the problems are gone in 1.63.0
Reacted by Benoît HavretConfirmed no more problem with 1.63.0 (yet :) )
Reopening — 1.63.0 did not actually fix this for us.
I closed this on 2026-06-04 ("confirmed no more problem with 1.63.0"), but that was
a false positive: I had only verified on our DEV environment, which runs on real
hardware (AWS EC2, Intel Xeon) and never reproduced the crash. The bug only
manifests on our INTEG environment, which runs on an emulated QEMU CPU — and there,
1.63.0 still hard-crashes (SIGSEGV ~5 min after a deploy, on 2026-06-09).That environment difference turns out to be the trigger itself (A/B below).
New crash signature (1.63.0) — different victim frame, consistent with the
residual memory corruption mentioned above:SIGSEGV (SI_KERNEL, si_addr=0x0)insideld-musl-x86_64.so.1(musl mallocng
internals), onC1 CompilerThread0, with repeated "error occurred during error
reporting" → native-heap corruption.- ddprof native engine loaded (
libjavaProfiler-dd-*.so),Native Stacks=true,
perf_events_paranoid=3(perf_event_open blocked by the default Docker seccomp
profile → ddprof on the software unwind path). - Happens during JVM warmup right after a (re)deploy (elapsed ~293s), not under
sustained load — per-restart, not traffic-driven.
Consistent with @TheDevOps in #11373 ("drastically improved but not at 0" on
1.63.0, crashing on a different frame).Isolating factor — the emulated QEMU CPU. Clean A/B with a byte-identical
Docker image (same image SHA; container not privileged, no added caps, default
seccomp). Only the host CPU differs:Env host CPU arch_perfmon result DEV Intel Xeon Platinum 8488C (real, EC2) present stable INTEG QEMU Virtual CPU version 2.5+ absent SIGSEGV at warmup ddprof is enabled on both agent versions (verified:
libjavaProfiler-dd-*.so
is loaded under 1.60.1 too, and runs fine on the same QEMU host) — so this is a
native-engine regression between 1.60.1 and 1.63.0, not an activation change:
1.60.1's ddprof handles the emulated CPU, 1.63.0's mis-walks the stack and
corrupts the heap.Env: Temurin JDK 25.0.3, Alpine 3.23 (musl), Spring Boot 4, container
-Xmx6g,
G1. Freshhs_err_pid*.log+ crash.jfravailable (redacted) — happy to attach.Workarounds confirmed: pin 1.60.1, or
DD_PROFILING_DDPROF_ENABLED=falseon 1.63.0.Reacted by ayemillerWe have also seen a quite few crashes from crash tracker reports, it looks like upstream bug: https://bugs.openjdk.org/browse/JDK-8363973
hs_err_pid1.redacted.log
Tracer Version(s)
1.62.0 (crashes). Reverting to 1.60.1 resolves it — so this is a regression introduced between 1.60.1 and 1.62.0.
Java Version(s)
OpenJDK Runtime Environment Temurin-25.0.1+8 (build 25.0.1+8-LTS), linux-amd64
JVM Vendor
Eclipse Adoptium / Temurin
OS / Environment
Alpine Linux v3.22 (musl libc), QEMU 4 cores / 11G, container
-Xmx6g, G1 GC. Spring Boot 4.0.6 app, virtual threads in use. Continuous profiler enabled (nativeddprof, allocation profiling on).Bug Report
After upgrading the agent to 1.62.0, the JVM hard-crashes (
SIGSEGV,is_crash) intermittently — ~10 min after startup under normal traffic. The crash is in the Datadog allocation profiler: on a sampled allocation it calls JVMTIGetStackTrace, and walking a virtual thread's frames segfaults insidevframe::java_sender().This appears to be a sibling of #9830 but on a different stackwalk path: #9830 was the async/ASGCT path (
vframeStreamForte::forte_next), mitigated by-Ddd.profiling.ddprof.cstack=vm(default since 1.55.0). This one is the JVMTIGetStackTracepath used by the allocation sampler on virtual threads, whichcstack=vmdoes not appear to cover (we are on 1.62.0, well past that default, and still crashing).si_code: 128 (SI_KERNEL), si_addr: 0x0; the faulting thread is a virtual-thread carrier (ForkJoinPool-1-worker),_thread_in_vm. TheArrayList.grow/QueryExecutorImpl.processResultsJava frames are just the innocent allocation being sampled (no large allocation — verified, the underlying collection is small).Workaround: downgrade to 1.60.1 (clean). We expect
DD_PROFILING_ALLOCATION_ENABLED=false(orDD_PROFILING_DDPROF_ENABLED=false, given musl) would also avoid it.Crash stack (top frames from
hs_err_pid1.log):I can attach the full
hs_err_pid1.log(and the.jfr) if useful.