You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Worker SIGSEGV / heap corruption on records that fail to decode — native decode path is not fail-safe ([customer-cluster], 5.1.3/5.1.4) #1370
On the [customer-cluster] production cluster, http worker threads SIGSEGV repeatedly with heap-corruption signatures, on the same nodes that are continuously logging large volumes of "Error decoding record: Data read, but end of buffer not reached" (the decode-failure / silent-degradation issue tracked in #1348, and the audit-store corruption tracked in #1132).
The decode path's JS try/catch (resources/RecordEncoder.ts:425-450) already makes a thrown decode error fail-safe (logs Error decoding record, returns null — this is the #1348 flood). A native out-of-bounds access in the decode pipeline cannot be caught that way — it corrupts the heap and crashes the process. That is the gap this issue is about: decode is fail-safe against thrown errors but not against native memory faults, so the same corrupt records that degrade silently in #1348 can instead take down a worker.
Scope note: #1348 already owns the decode-failure / silent-degradation analysis and the structon structure-desync hypothesis; #1132 owns the createAuditEntryENTRY_HEADER aliasing race that corrupts the txnlog. This issue is the crash escalation of those, filed separately because the impact (worker SIGSEGV, not silent degradation) and the fix surface (native fail-safety) are different. Please de-dup/merge as the team sees fit.
Native worker crashes do not appear in hdb.log/system.log — only the OS kernel log (dmesg/journalctl -k) captures them. All faults are error 4 (user-mode read of a not-present page) at small fault offsets (0x10/0x65/0xe0/0x108) — the signature of a near-null pointer dereference in native code.
Node A — [customer-node-A] (5.1.4)
dmesg (36-day uptime, clean kernel log):
Jun 17 22:46:08 http[317288] segfault at 10 error 4 in node ← triggered the restart→5.1.4 upgrade
Jun 16 13:08:15 http[58715] segfault at 108 error 4 in libc.so.6
Jun 15 20:42:09 http[55970] segfault at 108 error 4 in libc.so.6
Jun 15 20:40:34 http[23554] segfault at 108 error 4 in libc.so.6
Jun 15 16:35:55 http[3843395] segfault at 108 error 4 in libc.so.6
Jun 11 18:08:14 http[3605464] segfault at e0 error 4 in libc.so.6
May 20 02:04:49 http[904808] general protection fault in node
Decode failures here are on the application/RocksDB read path (Hints.get → rocksdb-js Store.decodeValue → RecordEncoder.decode → msgpackr checkedRead) — this is exactly #1348. ~1400 caught decode errors in system.log.
Node B (leader) — [customer-node-B] (5.1.4)
journalctl -k (persistent; dmesg here is unusable — only ~1h deep and flooded by the Datadog agent container's apparmor ptrace DENIED audit spam every ~10s, which evicts kernel entries):
Jun 17 22:46:48 http[152961] segfault at 65 error 4 in node ← coincident w/ the 22:46:50 upgrade restart
Jun 10 23:54:29 libuv-worker[3529417] general protection fault
Captured in-process SIGSEGV — log/crash.log (May 20), a glibc abort (heap corruption detected) surfacing in OpenSSL TLS record read:
TLS is just where the next allocation touched the already-corrupted heap — not the cause. Decode failures on this leader are in the audit store (auditStore.getValue → blob.decodeFromDatabase → RecordEncoder.decode → msgpackr checkedRead), logged by [main/0] during transaction-log replay after the improper shutdown ("Harper was not properly shutdown, replaying transaction logs"). This is the read side of the corruption #1132 describes.
Why "heap corruption from native decode" is the leading hypothesis
Confirmed: workers SIGSEGV repeatedly with classic heap-corruption signatures (glibc abort at free/malloc; error 4 near-null reads in node/libc).
Hypothesis (mechanism, not yet isolated): the native decode/encode pipeline — msgpackr's native extractor and/or @harperfast/rocksdb-js native code — reads/writes out of bounds when handed a corrupt buffer (e.g. a bad length prefix), corrupting the heap. The corruption then surfaces as a fault in whatever native code next allocates (OpenSSL TLS, libuv, V8), which is why the crash sites look unrelated to decoding.
I have not proven the corrupt-record → SIGSEGV link (the caught decodes return null and leave no crash; the crashing ones leave no JS trace, only a kernel entry). This issue asserts the correlation + a plausible mechanism and asks engineering to isolate it.
Alternative heap-corruption sources to rule in/out: (a) #1132's ENTRY_HEADER shared-buffer aliasing race (already a confirmed native-buffer correctness bug); (b) @harperfast/rocksdb-js native decode; (c) msgpackr native extractor. Not the cause on 5.1.4: the May 14 us-west-1 crash.log was a @datadog/pprofWallProfiler::CleanupHook crash — the pprof prebuild is absent in 5.1.4, so that specific profiler crash is resolved and is a separate red herring.
Repro material (raw bytes from logs)
Each Error decoding record line ends with data: <hex> (the undecodable value). Note the shared 42 79 ed … 0e0000 40000000 … framing — same as #1348.
Isolate the crash: build/run with AddressSanitizer (or a debug @harperfast/rocksdb-js + msgpackr) and feed a captured data: payload through the decode path to see whether the native extractor reads OOB — confirms or kills the hypothesis directly.
Test the native-path-off control: force the pure-JS msgpackr decode path (no native extractor) on an affected node and see whether the decode failures stay caught-and-logged but the SIGSEGVs stop. If they stop, the native extractor is implicated.
Make native decode fail-safe: ensure a malformed buffer cannot drive a native OOB access — validate length/structure bounds before the native call, or guarantee the bounds-checked path for stored values. The JS try/catch in RecordEncoder.decode is necessary but insufficient because it can't catch a native memory fault.
Operational: the Datadog agent apparmor ptrace DENIED audit spam is blinding dmesg on these hosts — worth fixing so native crashes remain diagnosable from the kernel ring buffer.
Filed from a live production investigation (fabric-investigation) of [customer-cluster] on 2026-06-18. Compiled by Claude (Opus 4.8). Crash figures are from dmesg/journalctl -k/crash.log; decode figures from system.log at time of writing. The corrupt-record → SIGSEGV link is a hypothesis with supporting correlation, not an isolated root cause.
Second production cluster with this signature — plus a customer-visible cascade (crash → parked index → blocked upload)
Independent corroboration from a different prod cluster (v4-origin, recently in-place upgraded 5.0.x→5.1.3), from a read-only fabric-investigation. Same heap-corruption-abort signature as Node B above:
Kernel log (persistent, 3-week window): traps: http[…] general protection fault ip:…50f error:0 in libc.so.6 — that …50f ip is abort+0x170 (glibc detected heap corruption and called abort()).
crash.log: PID 1 received SIGSEGV for address: 0x0, stack abort → malloc / operator new (_Znwm) → segfault-handler — the malloc-time heap-corruption signature, not a logic fault.
Recurrence: one node took a storm of ~15 in a single day; another took one at its in-place-upgrade boot; both stable since. Recurs over time, consistent with the pattern above.
New angle — downstream blast radius. Here the upgrade-boot crash didn't just restart a worker: crash-recovery left a secondary index rebuilding/parked, queries on that attribute returned 503 "not indexed yet", and an app that treats 503 as a per-row business "skip" dropped a full data upload (100% skipped), repeatedly, until the index recovered hours later. So a single crash here can escalate into a hard, silent data-load blocker — the trigger upstream of the indexing-access mitigations in #1354 (#1363 typed-retryable 503, #1371 don't-park-on-transient), which make the fallout graceful but don't stop the crash.
Caveats (correlation, same un-isolated status as this issue):
Corruption source not isolated here either — native fault, no JS trace; only the kernel log + crash.log captured it.
The decode misreads at the crash boot were the harper-pro#352handled / fail-safe kind (hdb_nodes … did not decode … authorizing by cert … self-heals — that fix working), so I can't pin this crash on the Fix getThisNodeName() reference #352 path; it's a temporal correlation at the in-place-upgrade boot.
The crash → parked-index → skip step is inferred from the timeline (node logs at level: warn, so the runIndexing/indexingFailed lifecycle isn't persisted) + the observed 503 behavior.
Filed by Claude (Opus 4.8) from a read-only fabric-investigation.
changed the title [-]Worker SIGSEGV / heap corruption on records that fail to decode — native decode path is not fail-safe (eh-prod.gend, 5.1.3/5.1.4)[/-][+]Worker SIGSEGV / heap corruption on records that fail to decode — native decode path is not fail-safe ([customer-cluster], 5.1.3/5.1.4)[/+]on Jun 20, 2026
Summary
On the
[customer-cluster]production cluster,httpworker threads SIGSEGV repeatedly with heap-corruption signatures, on the same nodes that are continuously logging large volumes of "Error decoding record: Data read, but end of buffer not reached" (the decode-failure / silent-degradation issue tracked in #1348, and the audit-store corruption tracked in #1132).The decode path's JS
try/catch(resources/RecordEncoder.ts:425-450) already makes a thrown decode error fail-safe (logsError decoding record, returnsnull— this is the #1348 flood). A native out-of-bounds access in the decode pipeline cannot be caught that way — it corrupts the heap and crashes the process. That is the gap this issue is about: decode is fail-safe against thrown errors but not against native memory faults, so the same corrupt records that degrade silently in #1348 can instead take down a worker.Harper version: harper-pro 5.1.4 (us-west-1) / observed since 5.1.3 (matches #1348). Core resources:
resources/RecordEncoder.ts,resources/blob.ts,resources/auditStore.ts.Evidence — two nodes, same cluster, two stores
Native worker crashes do not appear in
hdb.log/system.log— only the OS kernel log (dmesg/journalctl -k) captures them. All faults areerror 4(user-mode read of a not-present page) at small fault offsets (0x10/0x65/0xe0/0x108) — the signature of a near-null pointer dereference in native code.Node A —
[customer-node-A](5.1.4)dmesg(36-day uptime, clean kernel log):Decode failures here are on the application/RocksDB read path (
Hints.get→rocksdb-js Store.decodeValue→RecordEncoder.decode→ msgpackrcheckedRead) — this is exactly #1348. ~1400 caught decode errors insystem.log.Node B (leader) —
[customer-node-B](5.1.4)journalctl -k(persistent;dmesghere is unusable — only ~1h deep and flooded by the Datadogagentcontainer'sapparmor ptrace DENIEDaudit spam every ~10s, which evicts kernel entries):Captured in-process SIGSEGV —
log/crash.log(May 20), a glibcabort(heap corruption detected) surfacing in OpenSSL TLS record read:TLS is just where the next allocation touched the already-corrupted heap — not the cause. Decode failures on this leader are in the audit store (
auditStore.getValue→blob.decodeFromDatabase→RecordEncoder.decode→ msgpackrcheckedRead), logged by[main/0]during transaction-log replay after the improper shutdown ("Harper was not properly shutdown, replaying transaction logs"). This is the read side of the corruption #1132 describes.Why "heap corruption from native decode" is the leading hypothesis
abortatfree/malloc;error 4near-null reads innode/libc).@harperfast/rocksdb-jsnative code — reads/writes out of bounds when handed a corrupt buffer (e.g. a bad length prefix), corrupting the heap. The corruption then surfaces as a fault in whatever native code next allocates (OpenSSL TLS, libuv, V8), which is why the crash sites look unrelated to decoding.I have not proven the corrupt-record → SIGSEGV link (the caught decodes return null and leave no crash; the crashing ones leave no JS trace, only a kernel entry). This issue asserts the correlation + a plausible mechanism and asks engineering to isolate it.
Alternative heap-corruption sources to rule in/out: (a) #1132's
ENTRY_HEADERshared-buffer aliasing race (already a confirmed native-buffer correctness bug); (b)@harperfast/rocksdb-jsnative decode; (c) msgpackr native extractor. Not the cause on 5.1.4: the May 14 us-west-1crash.logwas a@datadog/pprofWallProfiler::CleanupHookcrash — the pprof prebuild is absent in 5.1.4, so that specific profiler crash is resolved and is a separate red herring.Repro material (raw bytes from logs)
Each
Error decoding recordline ends withdata: <hex>(the undecodable value). Note the shared42 79 ed … 0e0000 40000000 …framing — same as #1348.Node A (Hints/RocksDB store):
Node B (audit store):
Suggested next steps
@harperfast/rocksdb-js+ msgpackr) and feed a captureddata:payload through the decode path to see whether the native extractor reads OOB — confirms or kills the hypothesis directly.try/catchinRecordEncoder.decodeis necessary but insufficient because it can't catch a native memory fault.RangeErrorguard) to remove one confirmed corruption source, and progress Replicated records fail to decode (structon 'end of buffer not reached') — cluster-wide, silent degradation #1348 (structure-desync) to stop generating new corrupt records.apparmor ptrace DENIEDaudit spam is blindingdmesgon these hosts — worth fixing so native crashes remain diagnosable from the kernel ring buffer.Related
createAuditEntryENTRY_HEADERrace corrupts txnlog (a confirmed corruption source + reader hardening)Filed from a live production investigation (fabric-investigation) of
[customer-cluster]on 2026-06-18. Compiled by Claude (Opus 4.8). Crash figures are fromdmesg/journalctl -k/crash.log; decode figures fromsystem.logat time of writing. The corrupt-record → SIGSEGV link is a hypothesis with supporting correlation, not an isolated root cause.