Skip to content

Stale shared-structures: nodes silently return empty/partial query results without error #1163

Description

@kriszyp

Summary

Nodes affected by stale msgpackr shared-structures silently return empty or incomplete query results — data present cluster-wide comes back as empty on the affected node — without surfacing any error to the caller.

Symptoms

  • Queries that should return records return empty results on affected nodes
  • Log shows Could not find typed structure 1 (msgpackr decode failures)
  • Symptom clears on process restart
  • Other nodes in the cluster return the same queries correctly

Distinction from related issues

This is distinct from the Unknown constant decode error (PR #1154) where records become permanently unreadable. Here records are intact but queries silently return partial/empty results due to structure decode failures — the failure mode is invisible to callers.

Expected fix

The msgpackr shared-structures fix shipping in 5.0.28 is expected to resolve this. Filing to track confirmation post-rollout.

Reporter

Joshua Johnson

Activity

  1. added this to the v5.0 milestone on Jun 8, 2026
  2. added
    bugSomething isn't working
    area:storageStorage engine, LMDB/RocksDB, compaction
    on Jun 8, 2026
  3. kriszyp commented on Jun 8, 2026

    @kriszyp
    MemberAuthor

    Field corroboration from a production cluster, harper-pro 5.0.28 — 2026-06-08

    Active right now on a production cluster, and it's the cause of customer-reported "missing results" (some Table.search() calls silently return empty for records that exist — matches this issue's title exactly).

    Error (live, ongoing):

    [http/N] [error]: Error decoding record Error: Could not find typed structure 1
    

    On the worst node (region jp-osa-1) these run continuously — 66 occurrences, earliest 20:34 GMT, latest 21:45 GMT (checked at 21:46). A record references shared structure #1 that the node's structures buffer doesn't contain, so it decodes to nothing and search() returns empty for it with no surfaced error to the caller.

    Cluster distribution (count of Could not find typed structure in current container logs):

    node (region) count notes
    jp-osa-1 66 ongoing, latest 21:45
    br-gru-1 3
    us-east-1 2
    ap-northeast-1 2
    us-central-1 1

    The affected nodes are exactly the ones returning partial/empty query results; a restart clears it temporarily (reloads structures), then it recurs.

    Possible linkage to harper-pro#289 (worth checking): this cluster's replication mesh is currently wedged (most nodes 0/30 peers connected after a synchronized restart). The 5.0.18 "force structures reload on first audit record per subscription" recovery can't fire when no audit records are flowing — so with replication down, a node holding a stale/incomplete shared-structures buffer has no path back to a good state and keeps emitting decode failures. If that linkage holds, fixing the replication wedge may also stop the steady-state recurrence here. (Hypothesis — not verified at the code level.)

    Customer impact: a production customer cluster (identity redacted) returned partial/empty early-hints query results to end users on affected nodes.

    — Investigated via /fabric-investigation (Claude)

  4. kriszyp commented on Jun 10, 2026

    @kriszyp
    MemberAuthor

    Field report: severe system-table manifestation (and 5.0.28/5.0.30 did not fully resolve it)

    Observed on a production Fabric cluster (16 nodes, multi-region) during routine operations and a rolling upgrade. Adding this as confirmation data, because the issue currently states the 5.0.28 shared-structures fix was expected to resolve this — on this cluster it did not: the cluster was running 5.0.28 and still hit the decode failures, and the worst incident occurred on the 5.0.30 restart.

    What happened

    On 5 of the 16 nodes, immediately after the 5.0.30 process restart:

    1. Startup logged the familiar signature — Could not find typed structure N / incomplete: true (core/resources/blob.ts:969, core/resources/auditStore.ts:558).
    2. cluster_status then reported the node's hdb_nodes table as containing only itself — i.e. cluster membership looked wiped. Healthy nodes that decoded cleanly showed the full 16-row membership and never exhibited this.

    Why this is worse than "empty query results"

    The decode failure here hit the system.hdb_nodes change stream, not just an application query. The replication layer subscribes to hdb_nodes updates; when a change event's value fails to decode it surfaces as undefined, and the replication code interpreted that nullish value as a node deletion, emitting a burst of Node was deleted, unsubscribing from node … for every peer at once. Net effect: the node tore down all of its outbound replication subscriptions — an operational outage, triggered purely by a read/decode failure rather than any real data change.

    Evidence it was a decode failure and not a genuine deletion:

    • A healthy node's read_audit_log on system.hdb_nodes showed only patch operations, zero deletes, across the whole window. A genuine replicated delete would be audited cluster-wide.
    • The Node was deleted burst recurred on subsequent restarts — consistent with persistent stale shared-structures re-failing to decode, not a one-time event.
    • Only nodes that hit the decode errors showed the symptom; nodes that decoded cleanly retained full membership.

    Takeaways

    • The 5.0.28 shared-structures fix is necessary-but-insufficient here; stale structures still occur post-5.0.28 and post-5.0.30 on this cluster.
    • When the decode failure lands on a system table that drives control-plane behavior (here, replication membership), the blast radius is much larger than silent empty query results.
    • Recovery required a re-clone of the affected nodes from a healthy leader.

    I'm filing a separate replication-robustness issue for the downstream behavior (an undecodable change-stream value being treated as a deletion / subscription teardown), and putting up a guard so a decode failure can no longer masquerade as a node removal. That hardening is defense-in-depth and does not replace fixing the stale-structures root cause tracked here.

  5. self-assigned this
    on Jun 10, 2026
  6. kriszyp commented on Jun 10, 2026

    @kriszyp
    MemberAuthor

    Opened #1237 (draft) addressing the silent aspect of this issue: RecordEncoder.decode now detects the terminal missing-shared-structure errors and routes them to a distinct, non-fatal path — an analytics counter (decode-missing-structure) plus a dedicated warning — instead of logging at error and silently returning null. Affected nodes become observable/alertable rather than silently serving partial results.

    This is intentionally non-fatal (still returns null): a hard throw regressed DB initialization, since internal meta/__dbis__ scans legitimately tolerate an undecodable record. It does not recover the missing structure — a node genuinely missing the structure (e.g. a replica that received the record but not the structure-buffer update) still returns null, just loudly. Full recovery remains gated on the structure-delivery/replication path.

    Generated by Claude (Opus 4.7).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:storageStorage engine, LMDB/RocksDB, compactionbugSomething isn't working

Type

No type

Fields

Priority

None yet

Projects

No projects

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions