Repository navigation
Stale shared-structures: nodes silently return empty/partial query results without error #1163
Description
Activity
- addedbugSomething isn't workingSomething isn't workingarea:storageStorage engine, LMDB/RocksDB, compactionStorage engine, LMDB/RocksDB, compaction
on Jun 8, 2026 Field corroboration from a production cluster, harper-pro 5.0.28 — 2026-06-08
Active right now on a production cluster, and it's the cause of customer-reported "missing results" (some
Table.search()calls silently return empty for records that exist — matches this issue's title exactly).Error (live, ongoing):
[http/N] [error]: Error decoding record Error: Could not find typed structure 1On the worst node (region
jp-osa-1) these run continuously — 66 occurrences, earliest 20:34 GMT, latest 21:45 GMT (checked at 21:46). A record references shared structure #1 that the node's structures buffer doesn't contain, so it decodes to nothing andsearch()returns empty for it with no surfaced error to the caller.Cluster distribution (count of
Could not find typed structurein current container logs):node (region) count notes jp-osa-1 66 ongoing, latest 21:45 br-gru-1 3 us-east-1 2 ap-northeast-1 2 us-central-1 1 The affected nodes are exactly the ones returning partial/empty query results; a restart clears it temporarily (reloads structures), then it recurs.
Possible linkage to harper-pro#289 (worth checking): this cluster's replication mesh is currently wedged (most nodes 0/30 peers connected after a synchronized restart). The 5.0.18 "force structures reload on first audit record per subscription" recovery can't fire when no audit records are flowing — so with replication down, a node holding a stale/incomplete shared-structures buffer has no path back to a good state and keeps emitting decode failures. If that linkage holds, fixing the replication wedge may also stop the steady-state recurrence here. (Hypothesis — not verified at the code level.)
Customer impact: a production customer cluster (identity redacted) returned partial/empty early-hints query results to end users on affected nodes.
— Investigated via
/fabric-investigation(Claude)Field report: severe system-table manifestation (and 5.0.28/5.0.30 did not fully resolve it)
Observed on a production Fabric cluster (16 nodes, multi-region) during routine operations and a rolling upgrade. Adding this as confirmation data, because the issue currently states the 5.0.28 shared-structures fix was expected to resolve this — on this cluster it did not: the cluster was running 5.0.28 and still hit the decode failures, and the worst incident occurred on the 5.0.30 restart.
What happened
On 5 of the 16 nodes, immediately after the 5.0.30 process restart:
- Startup logged the familiar signature —
Could not find typed structure N/incomplete: true(core/resources/blob.ts:969,core/resources/auditStore.ts:558). cluster_statusthen reported the node'shdb_nodestable as containing only itself — i.e. cluster membership looked wiped. Healthy nodes that decoded cleanly showed the full 16-row membership and never exhibited this.
Why this is worse than "empty query results"
The decode failure here hit the
system.hdb_nodeschange stream, not just an application query. The replication layer subscribes tohdb_nodesupdates; when a change event's value fails to decode it surfaces asundefined, and the replication code interpreted that nullish value as a node deletion, emitting a burst ofNode was deleted, unsubscribing from node …for every peer at once. Net effect: the node tore down all of its outbound replication subscriptions — an operational outage, triggered purely by a read/decode failure rather than any real data change.Evidence it was a decode failure and not a genuine deletion:
- A healthy node's
read_audit_logonsystem.hdb_nodesshowed onlypatchoperations, zerodeletes, across the whole window. A genuine replicated delete would be audited cluster-wide. - The
Node was deletedburst recurred on subsequent restarts — consistent with persistent stale shared-structures re-failing to decode, not a one-time event. - Only nodes that hit the decode errors showed the symptom; nodes that decoded cleanly retained full membership.
Takeaways
- The 5.0.28 shared-structures fix is necessary-but-insufficient here; stale structures still occur post-5.0.28 and post-5.0.30 on this cluster.
- When the decode failure lands on a system table that drives control-plane behavior (here, replication membership), the blast radius is much larger than silent empty query results.
- Recovery required a re-clone of the affected nodes from a healthy leader.
I'm filing a separate replication-robustness issue for the downstream behavior (an undecodable change-stream value being treated as a deletion / subscription teardown), and putting up a guard so a decode failure can no longer masquerade as a node removal. That hardening is defense-in-depth and does not replace fixing the stale-structures root cause tracked here.
- Startup logged the familiar signature —
Opened #1237 (draft) addressing the silent aspect of this issue:
RecordEncoder.decodenow detects the terminal missing-shared-structure errors and routes them to a distinct, non-fatal path — an analytics counter (decode-missing-structure) plus a dedicated warning — instead of logging at error and silently returningnull. Affected nodes become observable/alertable rather than silently serving partial results.This is intentionally non-fatal (still returns
null): a hard throw regressed DB initialization, since internal meta/__dbis__scans legitimately tolerate an undecodable record. It does not recover the missing structure — a node genuinely missing the structure (e.g. a replica that received the record but not the structure-buffer update) still returnsnull, just loudly. Full recovery remains gated on the structure-delivery/replication path.Generated by Claude (Opus 4.7).
Metadata
Metadata
Assignees
Labels
Type
Fields
Priority
Summary
Nodes affected by stale msgpackr shared-structures silently return empty or incomplete query results — data present cluster-wide comes back as empty on the affected node — without surfacing any error to the caller.
Symptoms
Could not find typed structure 1(msgpackr decode failures)Distinction from related issues
This is distinct from the Unknown constant decode error (PR #1154) where records become permanently unreadable. Here records are intact but queries silently return partial/empty results due to structure decode failures — the failure mode is invisible to callers.
Expected fix
The msgpackr shared-structures fix shipping in 5.0.28 is expected to resolve this. Filing to track confirmation post-rollout.
Reporter
Joshua Johnson