Skip to content

Worker SIGSEGV / heap corruption on records that fail to decode — native decode path is not fail-safe ([customer-cluster], 5.1.3/5.1.4) #1370

Description

@maurice-harper

Summary

On the [customer-cluster] production cluster, http worker threads SIGSEGV repeatedly with heap-corruption signatures, on the same nodes that are continuously logging large volumes of "Error decoding record: Data read, but end of buffer not reached" (the decode-failure / silent-degradation issue tracked in #1348, and the audit-store corruption tracked in #1132).

The decode path's JS try/catch (resources/RecordEncoder.ts:425-450) already makes a thrown decode error fail-safe (logs Error decoding record, returns null — this is the #1348 flood). A native out-of-bounds access in the decode pipeline cannot be caught that way — it corrupts the heap and crashes the process. That is the gap this issue is about: decode is fail-safe against thrown errors but not against native memory faults, so the same corrupt records that degrade silently in #1348 can instead take down a worker.

Scope note: #1348 already owns the decode-failure / silent-degradation analysis and the structon structure-desync hypothesis; #1132 owns the createAuditEntry ENTRY_HEADER aliasing race that corrupts the txnlog. This issue is the crash escalation of those, filed separately because the impact (worker SIGSEGV, not silent degradation) and the fix surface (native fail-safety) are different. Please de-dup/merge as the team sees fit.

Harper version: harper-pro 5.1.4 (us-west-1) / observed since 5.1.3 (matches #1348). Core resources: resources/RecordEncoder.ts, resources/blob.ts, resources/auditStore.ts.

Evidence — two nodes, same cluster, two stores

Native worker crashes do not appear in hdb.log/system.log — only the OS kernel log (dmesg/journalctl -k) captures them. All faults are error 4 (user-mode read of a not-present page) at small fault offsets (0x10/0x65/0xe0/0x108) — the signature of a near-null pointer dereference in native code.

Node A — [customer-node-A] (5.1.4)

dmesg (36-day uptime, clean kernel log):

Jun 17 22:46:08  http[317288]  segfault at 10   error 4  in node          ← triggered the restart→5.1.4 upgrade
Jun 16 13:08:15  http[58715]   segfault at 108  error 4  in libc.so.6
Jun 15 20:42:09  http[55970]   segfault at 108  error 4  in libc.so.6
Jun 15 20:40:34  http[23554]   segfault at 108  error 4  in libc.so.6
Jun 15 16:35:55  http[3843395] segfault at 108  error 4  in libc.so.6
Jun 11 18:08:14  http[3605464] segfault at e0   error 4  in libc.so.6
May 20 02:04:49  http[904808]  general protection fault   in node

Decode failures here are on the application/RocksDB read path (Hints.get → rocksdb-js Store.decodeValue → RecordEncoder.decode → msgpackr checkedRead) — this is exactly #1348. ~1400 caught decode errors in system.log.

Node B (leader) — [customer-node-B] (5.1.4)

journalctl -k (persistent; dmesg here is unusable — only ~1h deep and flooded by the Datadog agent container's apparmor ptrace DENIED audit spam every ~10s, which evicts kernel entries):

Jun 17 22:46:48  http[152961]        segfault at 65  error 4  in node   ← coincident w/ the 22:46:50 upgrade restart
Jun 10 23:54:29  libuv-worker[3529417]  general protection fault

Captured in-process SIGSEGV — log/crash.log (May 20), a glibc abort (heap corruption detected) surfacing in OpenSSL TLS record read:

PID 1 received SIGSEGV for address: 0x0
  libc abort  ← glibc detected heap corruption at malloc/free
  CRYPTO_malloc → tls_get_more_records → tls_read_record → ssl3_read_bytes → SSL_read
  node::crypto::TLSWrap::ClearOut → OnStreamRead → LibuvStreamWrap::OnUvRead → uv_run → Worker::Run

TLS is just where the next allocation touched the already-corrupted heap — not the cause. Decode failures on this leader are in the audit store (auditStore.getValue → blob.decodeFromDatabase → RecordEncoder.decode → msgpackr checkedRead), logged by [main/0] during transaction-log replay after the improper shutdown ("Harper was not properly shutdown, replaying transaction logs"). This is the read side of the corruption #1132 describes.

Why "heap corruption from native decode" is the leading hypothesis

I have not proven the corrupt-record → SIGSEGV link (the caught decodes return null and leave no crash; the crashing ones leave no JS trace, only a kernel entry). This issue asserts the correlation + a plausible mechanism and asks engineering to isolate it.

Alternative heap-corruption sources to rule in/out: (a) #1132's ENTRY_HEADER shared-buffer aliasing race (already a confirmed native-buffer correctness bug); (b) @harperfast/rocksdb-js native decode; (c) msgpackr native extractor. Not the cause on 5.1.4: the May 14 us-west-1 crash.log was a @datadog/pprof WallProfiler::CleanupHook crash — the pprof prebuild is absent in 5.1.4, so that specific profiler crash is resolved and is a separate red herring.

Repro material (raw bytes from logs)

Each Error decoding record line ends with data: <hex> (the undecodable value). Note the shared 42 79 ed … 0e0000 40000000 … framing — same as #1348.

Node A (Hints/RocksDB store):

[customer-data-redacted]
[customer-data-redacted]
[customer-data-redacted]

Node B (audit store):

68656173742d312e65682d70726f642e67656e642e6861727065726661627269632e636f6daa4561   ("…heast-1.[customer-cluster].harperfabric.com" + trailing bytes)

Suggested next steps

  1. Isolate the crash: build/run with AddressSanitizer (or a debug @harperfast/rocksdb-js + msgpackr) and feed a captured data: payload through the decode path to see whether the native extractor reads OOB — confirms or kills the hypothesis directly.
  2. Test the native-path-off control: force the pure-JS msgpackr decode path (no native extractor) on an affected node and see whether the decode failures stay caught-and-logged but the SIGSEGVs stop. If they stop, the native extractor is implicated.
  3. Make native decode fail-safe: ensure a malformed buffer cannot drive a native OOB access — validate length/structure bounds before the native call, or guarantee the bounds-checked path for stored values. The JS try/catch in RecordEncoder.decode is necessary but insufficient because it can't catch a native memory fault.
  4. Land createAuditEntry shared ENTRY_HEADER view races on no-encodedRecord path, corrupts txnlog #1132 (writer defensive copy + reader RangeError guard) to remove one confirmed corruption source, and progress Replicated records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent degradation #1348 (structure-desync) to stop generating new corrupt records.
  5. Operational: the Datadog agent apparmor ptrace DENIED audit spam is blinding dmesg on these hosts — worth fixing so native crashes remain diagnosable from the kernel ring buffer.

Related


Filed from a live production investigation (fabric-investigation) of [customer-cluster] on 2026-06-18. Compiled by Claude (Opus 4.8). Crash figures are from dmesg/journalctl -k/crash.log; decode figures from system.log at time of writing. The corrupt-record → SIGSEGV link is a hypothesis with supporting correlation, not an isolated root cause.

Activity

  1. heskew commented on Jun 18, 2026

    @heskew
    Contributor

    Second production cluster with this signature — plus a customer-visible cascade (crash → parked index → blocked upload)

    Independent corroboration from a different prod cluster (v4-origin, recently in-place upgraded 5.0.x→5.1.3), from a read-only fabric-investigation. Same heap-corruption-abort signature as Node B above:

    • Kernel log (persistent, 3-week window): traps: http[…] general protection fault ip:…50f error:0 in libc.so.6 — that …50f ip is abort+0x170 (glibc detected heap corruption and called abort()).
    • crash.log: PID 1 received SIGSEGV for address: 0x0, stack abort → malloc / operator new (_Znwm) → segfault-handler — the malloc-time heap-corruption signature, not a logic fault.
    • Recurrence: one node took a storm of ~15 in a single day; another took one at its in-place-upgrade boot; both stable since. Recurs over time, consistent with the pattern above.

    New angle — downstream blast radius. Here the upgrade-boot crash didn't just restart a worker: crash-recovery left a secondary index rebuilding/parked, queries on that attribute returned 503 "not indexed yet", and an app that treats 503 as a per-row business "skip" dropped a full data upload (100% skipped), repeatedly, until the index recovered hours later. So a single crash here can escalate into a hard, silent data-load blocker — the trigger upstream of the indexing-access mitigations in #1354 (#1363 typed-retryable 503, #1371 don't-park-on-transient), which make the fallout graceful but don't stop the crash.

    Caveats (correlation, same un-isolated status as this issue):

    • Corruption source not isolated here either — native fault, no JS trace; only the kernel log + crash.log captured it.
    • The decode misreads at the crash boot were the harper-pro#352 handled / fail-safe kind (hdb_nodes … did not decode … authorizing by cert … self-heals — that fix working), so I can't pin this crash on the Fix getThisNodeName() reference #352 path; it's a temporal correlation at the in-place-upgrade boot.
    • The crash → parked-index → skip step is inferred from the timeline (node logs at level: warn, so the runIndexing/indexingFailed lifecycle isn't persisted) + the observed 503 behavior.

    Filed by Claude (Opus 4.8) from a read-only fabric-investigation.

  2. changed the title [-]Worker SIGSEGV / heap corruption on records that fail to decode — native decode path is not fail-safe (eh-prod.gend, 5.1.3/5.1.4)[/-] [+]Worker SIGSEGV / heap corruption on records that fail to decode — native decode path is not fail-safe ([customer-cluster], 5.1.3/5.1.4)[/+] on Jun 20, 2026
  3. added this to the v5.2 milestone on Jul 7, 2026
  4. modified the milestones: v5.2, v5.3 on Jul 23, 2026
  5. added theissue type on Aug 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:storageStorage engine, LMDB/RocksDB, compactionbugSomething isn't working

    Type

    Fields

    Priority

    P2

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions