Skip to content

JVM SIGSEGV in vframe::java_sender via allocation-profiler JVMTI GetStackTrace on a virtual thread (JDK 25, Alpine/musl) — regression in 1.62.0 (1.60.1 OK) #11518

Description

@soidisant

hs_err_pid1.redacted.log

Tracer Version(s)

1.62.0 (crashes). Reverting to 1.60.1 resolves it — so this is a regression introduced between 1.60.1 and 1.62.0.

Java Version(s)

OpenJDK Runtime Environment Temurin-25.0.1+8 (build 25.0.1+8-LTS), linux-amd64

JVM Vendor

Eclipse Adoptium / Temurin

OS / Environment

Alpine Linux v3.22 (musl libc), QEMU 4 cores / 11G, container -Xmx6g, G1 GC. Spring Boot 4.0.6 app, virtual threads in use. Continuous profiler enabled (native ddprof, allocation profiling on).

Bug Report

After upgrading the agent to 1.62.0, the JVM hard-crashes (SIGSEGV, is_crash) intermittently — ~10 min after startup under normal traffic. The crash is in the Datadog allocation profiler: on a sampled allocation it calls JVMTI GetStackTrace, and walking a virtual thread's frames segfaults inside vframe::java_sender().

This appears to be a sibling of #9830 but on a different stackwalk path: #9830 was the async/ASGCT path (vframeStreamForte::forte_next), mitigated by -Ddd.profiling.ddprof.cstack=vm (default since 1.55.0). This one is the JVMTI GetStackTrace path used by the allocation sampler on virtual threads, which cstack=vm does not appear to cover (we are on 1.62.0, well past that default, and still crashing).

si_code: 128 (SI_KERNEL), si_addr: 0x0; the faulting thread is a virtual-thread carrier (ForkJoinPool-1-worker), _thread_in_vm. The ArrayList.grow / QueryExecutorImpl.processResults Java frames are just the innocent allocation being sampled (no large allocation — verified, the underlying collection is small).

Workaround: downgrade to 1.60.1 (clean). We expect DD_PROFILING_ALLOCATION_ENABLED=false (or DD_PROFILING_DDPROF_ENABLED=false, given musl) would also avoid it.

Crash stack (top frames from hs_err_pid1.log):

#  SIGSEGV (0xb) at pc=..., pid=1, tid=197
# Java VM: OpenJDK 64-Bit Server VM Temurin-25.0.1+8 (25.0.1+8-LTS, mixed mode, sharing, tiered, compressed oops, compressed class ptrs, g1 gc, linux-amd64)
# Problematic frame:
# V  [libjvm.so+0x1190128]  vframe::java_sender() const+0x38

Current thread: JavaThread "ForkJoinPool-1-worker-18" daemon [_thread_in_vm, id=197]

Native frames: (J=compiled Java code, j=interpreted, Vv=VM code, C=native code)
V  [libjvm.so]  vframe::java_sender() const+0x38
V  [libjvm.so]  JvmtiEnvBase::get_stack_trace(javaVFrame*, int, int, jvmtiFrameInfo*, int*)
V  [libjvm.so]  GetStackTraceClosure::do_vthread(Handle)
V  [libjvm.so]  JvmtiHandshake::execute(...)
V  [libjvm.so]  JvmtiEnv::GetStackTrace(...)
V  [libjvm.so]  jvmti_GetStackTrace
C  [libjavaProfiler-dd-*.so]  Profiler::recordJVMTISample(...)
C  [libjavaProfiler-dd-*.so]  ObjectSampler::recordAllocation(...)
C  [libjavaProfiler-dd-*.so]  ObjectSampler::SampledObjectAlloc(...)
V  [libjvm.so]  JvmtiExport::post_sampled_object_alloc(JavaThread*, oopDesc*)
V  [libjvm.so]  MemAllocator::allocate() / InstanceKlass::allocate_objArray(...)
V  [libjvm.so]  OptoRuntime::new_array_C(...)

Java frames: (J=compiled Java code, j=interpreted, Vv=VM code)
J  java.util.ArrayList.grow()  java.base@25.0.1
J  org.postgresql.core.v3.QueryExecutorImpl.processResults(...)
J  jdk.internal.vm.Continuation.run()  java.base@25.0.1
J  java.lang.VirtualThread.runContinuation()  java.base@25.0.1
J  java.util.concurrent.ForkJoinPool.runWorker(...)  java.base@25.0.1
j  java.util.concurrent.ForkJoinWorkerThread.run()  java.base@25.0.1

siginfo: si_signo: 11 (SIGSEGV), si_code: 128 (SI_KERNEL), si_addr: 0x0

I can attach the full hs_err_pid1.log (and the .jfr) if useful.

Activity

  1. jbachorik commented on Jun 1, 2026

    @jbachorik
    Contributor

    Hi @soidisant ,

    Thanks for the report! Yes, hs_err*.log would be useful.

    FTR, we had some issues in 1.62.0 where concurrent access interleaved with arriving signals could have lead to use-after-free, hence memory corruption that could lead to spurious crashes like you are seeing.

    Otherwise, if not caused by the UaF this would point to a JVM bug because we are using standard JVMTI API from a standard context, so it should not be crashing like that.

  2. soidisant commented on Jun 1, 2026

    @soidisant
    Author

    HI @jbachorik

    Thanks for your response
    I added the log to the original message (redacted package names and some private naming)

  3. jbachorik commented on Jun 3, 2026

    @jbachorik
    Contributor

    Please, upgrade to 1.63.0 - we have a number of fixes in profiler that are targeting the concurrency issues and memory corruption. Internal testing and dogfooding is showing that the problems are gone in 1.63.0

  4. soidisant commented on Jun 4, 2026

    @soidisant
    Author

    Confirmed no more problem with 1.63.0 (yet :) )

  5. soidisant commented on Jun 9, 2026

    @soidisant
    Author

    Reopening — 1.63.0 did not actually fix this for us.

    I closed this on 2026-06-04 ("confirmed no more problem with 1.63.0"), but that was
    a false positive: I had only verified on our DEV environment, which runs on real
    hardware (AWS EC2, Intel Xeon) and never reproduced the crash. The bug only
    manifests on our INTEG environment, which runs on an emulated QEMU CPU — and there,
    1.63.0 still hard-crashes (SIGSEGV ~5 min after a deploy, on 2026-06-09).

    That environment difference turns out to be the trigger itself (A/B below).

    New crash signature (1.63.0) — different victim frame, consistent with the
    residual memory corruption mentioned above:

    • SIGSEGV (SI_KERNEL, si_addr=0x0) inside ld-musl-x86_64.so.1 (musl mallocng
      internals), on C1 CompilerThread0, with repeated "error occurred during error
      reporting" → native-heap corruption.
    • ddprof native engine loaded (libjavaProfiler-dd-*.so), Native Stacks=true,
      perf_events_paranoid=3 (perf_event_open blocked by the default Docker seccomp
      profile → ddprof on the software unwind path).
    • Happens during JVM warmup right after a (re)deploy (elapsed ~293s), not under
      sustained load — per-restart, not traffic-driven.

    Consistent with @TheDevOps in #11373 ("drastically improved but not at 0" on
    1.63.0, crashing on a different frame).

    Isolating factor — the emulated QEMU CPU. Clean A/B with a byte-identical
    Docker image (same image SHA; container not privileged, no added caps, default
    seccomp). Only the host CPU differs:

    Env host CPU arch_perfmon result
    DEV Intel Xeon Platinum 8488C (real, EC2) present stable
    INTEG QEMU Virtual CPU version 2.5+ absent SIGSEGV at warmup

    ddprof is enabled on both agent versions (verified: libjavaProfiler-dd-*.so
    is loaded under 1.60.1 too, and runs fine on the same QEMU host) — so this is a
    native-engine regression between 1.60.1 and 1.63.0, not an activation change:
    1.60.1's ddprof handles the emulated CPU, 1.63.0's mis-walks the stack and
    corrupts the heap.

    Env: Temurin JDK 25.0.3, Alpine 3.23 (musl), Spring Boot 4, container -Xmx6g,
    G1. Fresh hs_err_pid*.log + crash .jfr available (redacted) — happy to attach.

    Workarounds confirmed: pin 1.60.1, or DD_PROFILING_DDPROF_ENABLED=false on 1.63.0.

  6. zhengyu123 commented on Jun 10, 2026

    @zhengyu123
    Contributor

    We have also seen a quite few crashes from crash tracker reports, it looks like upstream bug: https://bugs.openjdk.org/browse/JDK-8363973

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions