Skip to content

Prune table drop markers instead of keeping one per dropped name forever #1018

Description

@kriszyp

Problem

A table drop leaves a durable drop marker (harper#2962): a /dropped/<table> row in the database's catalog, read by getTableDrops() in resources/databases.ts. harper-pro#956 sends these markers to peers in DB_SCHEMA frames, so a peer that was offline for a drop retires its stale copy on rejoin instead of resurrecting the table.

Markers are never pruned. That causes three problems:

  1. The list grows without bound. There is one marker per table name ever dropped in a database. Apps that create and drop tables dynamically (per tenant, per time bucket) accumulate markers indefinitely.
  2. Every schema frame carries all of them. sendDBSchema reads the full marker list from the catalog (uncached, by design) and sends it on every DB_SCHEMA frame. Cost scales with the number of drops ever made.
  3. Past 10,000 markers, the oldest drops stop protecting peers. The receiver processes at most 10,000 markers per database (replication/tableLifecycle.ts), and the sender orders newest first. A peer that stays offline across more than 10,000 later drops of distinct names can keep a stale table whose marker falls past the cutoff. Nodes that hold the marker keep refusing that table's definition and records, so it does not spread through them, but the stale copy survives on that peer.

The 10,000 bound was accepted for harper-pro#956 (see the review thread on replication/tableLifecycle.ts). This issue tracks the underlying retention problem.

Proposed direction

Prune a marker once it can no longer matter, instead of paginating an ever-growing list:

  • Peer-acknowledged pruning: a marker matters only to a node that has not yet received it. Track, per known peer, the last time a DB_SCHEMA frame carrying the database's markers was delivered. Prune a marker once every known peer has received a frame sent after the marker's droppedTime. This depends on a reliable cluster membership list, so a peer that was removed does not hold pruning forever.
  • Or an age horizon: prune markers older than the longest outage after which a node can still resume incrementally (it needs a full copy past that point anyway). This is simpler, but only correct if a node that is offline longer than that horizon is guaranteed to re-copy instead of resuming. That needs verifying.

The alternative raised in review, paginating DB_SCHEMA[4] into bounded frames and applying every batch before the definitions, removes the cutoff. But it keeps the unbounded growth and per-frame cost, and it changes the wire format.

Acceptance

  • Marker count per database stays bounded under continuous create/drop churn.
  • A peer offline across any number of drops still retires every stale table it holds on rejoin, or is forced into a full copy.
  • dropTableOfflinePeer.test.mjs scenarios still pass, plus a test covering pruning, including a peer that rejoins after its marker was pruned.

Related: harper-pro#956, harper#2962, harper#3119, harper#1212.

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    Fields

    Priority

    P3

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions