Skip to content

feat(plugin): the queue keeper is the render queue — nextRenderTime index, claim floor and sweep removed (v0.93.0) - #219

Merged
harper-joseph merged 4 commits into
mainfrom
feat/queue-keeper
Sep 28, 2026
Merged

harper-joseph merged 4 commits into
mainfrom
feat/queue-keeper

Conversation

@harper-joseph

@harper-joseph harper-joseph commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Why

#215: the render queue was the nextRenderTime index, walked three ways — the claim scan, the ready-set sweep (every 5 minutes, about 27 s on a 300k-row due set in production) and the backlog snapshot (every 15 minutes, capped at management.scanCap). Every walk degraded with the store, truncated when the queue was deepest, and was minutes stale. The claim floor, its guard band, periodic reset and unpin hatch existed only to keep the claim scan's seek point off the dead index entries every reschedule leaves behind.

This PR makes the queue an in-memory structure — the queue keeper — with RenderSchedule as its durable side, and deletes the index and everything built for it, with no transitional fallback. The benchmarks in #216 and #217 cleared the approach first: the subscription is exact under replication, at a cost within noise.

What changed

The keeper — worker 0 holds every schedule row this node owns by residency.

  • util/queueKeeper.js (new, pure): rows grouped by class (route × cadence × sitemap flag), each class ordered by due minute. The best K is a k-way merge over class heads using scoreOf at minute resolution; a test pins it to brute-force scoring and sorting. It computes every queue-state count.
  • util/queueKeeperService.js (new): the lifecycle. Once it has published, nothing stops it serving claims.
    • Load, serving from the first chunk. Subscribe, then walk the table by primary key in local chunks, publishing the ready set from what is held so far. Events apply as they arrive; a walked row is applied only if no event touched its key since the walk began (an event is at least as new as anything the walk read). A partial queue is safe to serve because every grant is checked against its row.
    • Never reloaded once live. A route or default-interval change reclassifies every held row in memory. A membership change reclassifies (dropping rows no longer owned) and runs the verification walk to add the rows gained. A closed or failed subscription is reopened (with backoff) and walked.
    • Publish every queue.keeper.publishInterval (1 s), leased rows skipped. Then point-reads the top 64 published rows (each key at most once a minute) and repairs any it holds wrongly.
    • Verify every queue.keeper.verifyInterval (1 h): a full walk that applies the table's value wherever the keeper differs (skipping keys an event touched meanwhile) and point-reads held keys the walk did not see. Counted as queue_health keeper_repaired, expected 0. A walk a resync needs is queued behind one already running, and retried if it fails.
    • Unreadable rows. If the ascending walk cannot get past a key that did not decode, the rest is read from the top down; only if that stops short too is the queue partial (exact: false, logged).
    • Peers. A node whose system.hdb_nodes names another node waits (up to 2 minutes) to see one before loading, since ownership is not knowable before the node list is. A single node loads at once.

Claims (claimSchedules, RenderQueue.claim):

  • Take entries from the ready set in order, point-read each one's row locally, skip any no longer due (queue_health claim_stale), and lease the rest. The renderer gets the row's live fromSitemap.
  • The lease grant is exclusive (util/renderLease.js): a key holding a live lease is refused, never renewed, and after winning a fresh slot a grant rescans the key's probe window and gives the slot back if the key is live elsewhere. So claims need no mutex, and the duplicate-grant path #218 measured is closed on the claim path. Two tests pin it; each fails with its guard removed.
  • A stalled worker 0 degrades instead of stopping claims. A heartbeat older than ten publish intervals (never under 30 s) is reported (stale), not enforced: the last set is still granted from until it drains; only then does the node count as not serving.
  • A render that never reports is bounded. A key whose last render.failureRetry.fastRetries + 2 leases all expired with no result — a renderer crashing on the URL, or results not reaching the node — is held back in the lease table instead of granted again: 2 leases, then 4, 8, … capped at its cadence. Nothing durable is written, no strike is counted, and the first result for the key clears it; so a node-wide delivery outage delays rows by a few leases, not a cadence. Named in a warning, counted as queue_health claim_wedged. (The removed unpin hatch filed a row a cadence forward once an hour.)
  • The ready set packs keys end to end, so a key of any length fits. A publish flips the slot before resetting the cursor, and a take reads the slot after its increment, so a take straddling a publish can hand an entry out twice (the exclusive grant refuses the second) but never skips the new head.

Status: a fourth queue status, unready, while the keeper is not serving or is loading with nothing due found yet. The keeper reports each change itself, so becoming ready is broadcast and the fleet is woken at once (HostHealth.applyMqttStatus), instead of finding out on its next idle poll (15 s) or the 60 s status sync. A deployed render fleet reads an unknown status as empty (idle backoff), which is right for a node that can grant nothing; the console shows it as a warn pill. The work-arrived hints (render-now, revalidate) use QueueState.noteWork(), which only moves empty → queued.

Removed, with the index:

Removed What it was for
nextRenderTime @indexed (schema) ordering claims, the sweep and the backlog walk; every reschedule paid an index delete and insert
The claim floor: its header words in the lease buffer, lowerFloorFor on every write, the guard band, periodic reset, reset-claim-floor action, unpin hatch keeping the claim scan's seek point off dead index entries
claimFromIndex, runClaimPass, the claim-scan warnings the index claim scan
sweepReadySet / startReadySweep, createTopK ranking by walking the whole due set
scanUpcoming and the cap override on POST /prerender_admin/backlog the capped backlog walk
Config: queue.claimFloor.*, queue.claimScanCap, queue.ready.enabled, queue.ready.sweepInterval, queue.ready.sweepCap, queue.keeper.enabled tuning the above
queue_health series: claim_scan_ms, ready_sweep_ms, ready_published, ready_cadence, below_floor, below_floor_age_ms, floor_pin_age_ms observing the above

New queue_health series: keeper_load_ms, keeper_publish_ms, keeper_verify_ms, keeper_repaired, keeper_live, claim_stale, claim_wedged.

Queue state: GET /prerender_admin/queue-state (new, node-local) — now / coming / lateness / flow / trust, exact, from a shared buffer any worker reads; 503 with the reason until the load finishes or whenever the keeper cannot vouch for its numbers. The README has the field table.

Admin API: overview.claimFloor is replaced by overview.leases (occupancy, oldest lease); explain and schedule no longer carry claimFloor / belowClaimFloor. The backlog snapshot's queue half comes from the keeper (source: 'keeper'; queueUnavailable while it is not live).

Also: walkUrlRange takes searchOptions and tags its cannot-advance throw (WALK_CANNOT_ADVANCE); lease buffer key render_queue_v2 (layout changed: no floor words, a miss count per slot); README (one "render queue" section replaces four) and METRICS.md updated; a pre-existing customer domain in test/readyQueue.test.js fixtures replaced with example.com.

Deploying

In place, with no data migration: RenderSchedule rows are unchanged and the keeper builds from them. Do not start from an empty store; losing the schedule rows re-files every target at a jittered time.

  • Plugin release only. The console is unchanged in this PR. Against this plugin, console 0.17.0 behaves as follows (verified from its code, not run), and nothing reads bad:
    • its claim-floor, below-floor and ready-sweep tiles read —;
    • "In flight" falls back to the snapshot;
    • "Prioritised claims" reads 100%, because every claim_granted is tagged ready;
    • Deep recompute still works, because its cap is ignored.
  • A mixed-version cluster is safe, so one node can go first. The index setting does not cross nodes. Verified in harper-pro 5.2.14: the replicated schema carries only each attribute's name, type and primary-key flag, and a schema-defined table keeps its local definition. So 0.92.0 nodes keep their index and claim floor while 0.93.0 nodes run the keeper. No writer's due time changed, and the peer explain read tolerates the missing claimFloor.
  • A restarted node grants claims within about a publish interval of its keeper starting, because it serves while loading. A clustered node first waits for server.nodes, which fills from the local hdb_nodes scan within seconds. The peer count excludes this node the same way harper-pro's replicator does (verified from source; not yet run on harper-pro).
  • Worker 0 does less work. A sweep of about 27 s every 5 minutes becomes sub-millisecond publishes each second plus an hourly verification walk. The keeper adds about 48 MB of heap at 250,000 rows.
  • Render fleet: a deployed renderer reads unready as empty (idle backoff) until queued arrives (verified in browser 1.38.0 HostHealth).
  • A config.yaml or stored override that still sets a removed key logs "Unknown configuration key" and is otherwise ignored; delete them.
  • Upgrade (observed, single node, 0.92.0 → 0.93.0 in place): the index drop needed no manual step. Harper leaves the old index's entries on disk: 5.2.14's removal looks the index up by the table's name rather than the attribute's, so nothing is queued to drop. That is dead weight only, and a rollback clears them before rebuilding.
  • A rollback takes two restarts per node.
    • Going back to 0.92.0 re-adds the index, and Harper rebuilds it from the records. Observed on a single node after 1,200 changes made without the index: the index matched the table exactly at +2, +17 and +62 s. The rebuild took 0.1 s at 2,000 rows and will take longer at production size; until it finishes, 0.92.0's claims fail with "nextRenderTime" is not indexed yet.
    • On a multi-thread node, harper-pro#681 makes that failure permanent on some threads. The issue is open, and 5.2.14 still has no receiver for indexing-finished, so any worker thread that opened the table mid-rebuild keeps failing every claim with that error against a complete index. It has been seen in production on a new index, on 12 of 64 threads.
    • So once the rebuild finishes, restart the node once more (or clear isIndexing on the stuck threads).
  • Rows claimed just before a restart are claimed again after it (the lease buffer is new), as at any restart.

Verification

  • Unit: plugin node --test 1385/1385, console 371/371, eslint and prettier --check clean. Four test files that tested only removed machinery are deleted (renderQueueFloor, renderQueuePin, readySweep, backlogScanCap); their two write-contract tests moved to scheduleWrite.test.js. Claim-path tests (renderQueueVariants, …Redirect, …Targetless, suppressionStatus) claim through a small keeper stand-in (test/support/keeperStandIn.js). New tests cover serving while loading, the in-memory resync on membership and route changes, a walk queued behind a running one, a failed resubscribe retried with backoff, config changes during a load, the unreadable-row tail walk (and the partial case), the walk never applying a row older than an event, the exclusive-grant races, the crash-loop hold growing and resetting, and the stale-but-served claim path.
  • Real Harper (harperfast/harper:5.2.13, packed plugin, 2 worker threads), 2,001 rows seeded under 0.92.0, then the new build installed over it and the container restarted; run five times across the revisions, the last on the final build:
    • Under 0.92.0 describe_table showed nextRenderTime (indexed); after the swap, no index and the same 2,001 records.
    • The first claim, 2 s after the restart, was granted (the homepage first, 2 cadences late); the replicated QueueStatus row already read queued. There was no peer wait: open-source Harper has no hdb_nodes, so the node counts as single. (An earlier revision, before serving while loading and the peer check, answered [] and unready for the 2-minute grace, then flipped to queued in the second it went live.)
    • The load read 2,001 rows in 33 ms. queue-state: due 1,507 of 2,001 (exactly the seeded due count), 1,005 sitemap / 502 discovered, exact: true. A publish took 3.6 ms; the head check read 64 rows and repaired 0.
    • Claims came in keeper order; 10 granted, 10 distinct.
    • Verification walk: 2,001 scanned and owned; 0 missing, mismatched, deleted or unowned; 25 ms.
    • Backlog snapshot: source: keeper, overdue 1,507, not truncated. overview.leases present, claimFloor absent. 0 error lines in the log.
  • Cost at production scale, pure keeper, 250,000 rows: topK(5000) p50 0.14–0.27 ms, state() about 0.5 ms, state document about 2.2 KB, heap about 48 MB.
  • Review: degraded — the Codex CLI's login had expired and Gemini had no key in this environment, so fresh-context Claude reviewers did the adversarial passes, four in all. Earlier passes changed the design (it was first built beside the sweep, then as the index with a fallback). The last two found, and this PR fixes with tests: an unreadable row that failed every load forever; a grant race onto two fresh slots; the take/publish head skip; the heartbeat held behind a slow publish; a resync walk dropped when a verification was running; a failed resubscribe never retried; the crash-loop bound pushing every row a cadence forward during a node-wide delivery failure (now a lease-table hold); a dead worker 0 reporting queued forever; empty reported mid-load; config changes during a load dropped; and a set of stale comments.

Not covered

  • A churned production-size store: the load is expected to take 20–30× longer per row than on a fresh one (bench/queue-index). It now decides only how long a restart serves a partial order, not how long it grants nothing.
  • More than two nodes, and a live membership change in a real cluster: the in-memory resync is unit-tested only. countConfiguredPeers is untested against harper-pro (open-source Harper has no hdb_nodes, so the single-node path is what the Docker runs exercised).
  • The console catch-up (drop the floor and sweep panels, read overview.leases, a queue-state view) and a cluster queue-state fan-out for autoscaling.
  • #218's mutex fix itself: setPause and the status sync still use the store mutex.

🤖 Generated with Claude Code

…e index and everything built for it are gone; v0.93.0

Worker 0 holds every schedule row this node owns, grouped by class (route x
cadence x sitemap flag) into minute buckets, loaded by a local primary-key
walk and kept current by a RenderSchedule subscription. RenderSchedule is its
durable side, read by primary key only; nextRenderTime is no longer indexed.

- claims take from the ready set it publishes each second, point-read each
  entry's durable row (stale ones skipped, queue_health claim_stale) and lease
  it; the grant is exclusive (a live lease is refused, never renewed, and a
  race onto two fresh slots gives one back), so claims need no mutex (#218)
- the ready set packs keys end to end, so any key length fits
- each publish re-reads its head and repairs; a verification walk every
  queue.keeper.verifyInterval repairs anything missed (keeper_repaired); a
  repair never reverts a newer event
- waits for a cluster peer before loading, rebuilds on membership, route or
  default-interval change; a load that cannot get past an unreadable row goes
  live on what it read, marked not exact
- until it is live the node grants no claims and reports queued
  (keeper_live gauge)

GET /prerender_admin/queue-state (new) serves its exact counts; the backlog
snapshot reads the same document.

Removed: the claim floor and its unpin hatch, the ready-set sweep, the index
claim scan, the capped backlog walk, the reset-claim-floor action, config
queue.claimFloor.*, queue.claimScanCap, queue.ready.enabled/sweepInterval/
sweepCap and queue.keeper.enabled, and the queue_health series that watched
them. overview.claimFloor becomes overview.leases.

Refs #215, #216, #217, #218.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request upgrades the prerender plugin to version 0.93.0, introducing a major architectural redesign of the render queue. The previous database-index-based claim scan, claim floor, and ready-set sweep mechanisms have been replaced by an in-memory 'Queue Keeper' service on worker 0. The keeper maintains the queue state in memory, kept current via a table subscription, and publishes the prioritized ready set to shared memory for concurrent claims. Consequently, the secondary index on nextRenderTime in the RenderSchedule table has been removed, and several obsolete configuration options and endpoints have been cleaned up. A new GET /prerender_admin/queue-state endpoint and updated metrics have been introduced to monitor the keeper's health and queue state. I have no feedback to provide as there are no review comments.

harper-joseph and others added 3 commits September 27, 2026 17:44
…rves claims

A node whose keeper is not serving (waiting for peers, loading, retrying a
failed load, or gone quiet) reported 'queued', so becoming ready changed
nothing and was never broadcast: the fleet found out on its next idle poll,
up to backoff.idleMs later, or at the next status sync. It now reports
'unready', and the keeper reports 'queued'/'empty' itself the moment it goes
live (and 'unready' when it withdraws), so the fleet is woken by the change.

A deployed render fleet treats an unknown status as 'empty' (HostHealth:
back off to the idle interval), which is right for a node that can grant
nothing. The work-arrived hints (render-now, revalidate) use a new
QueueState.noteWork() that only moves empty -> queued, so they no longer
lift 'unready' (or 'paused').

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…enders that never report

- serves while loading: the ready set is published from the first chunk of
  the load; events apply live and a walked row is applied only if no event
  touched its key since the walk began
- never reloads once live: a route or default-interval change reclassifies
  every held row in memory; a membership change reclassifies (dropping rows
  no longer owned) and runs the verification walk (adding rows gained); a
  closed subscription is reopened, then walked. Claims are served throughout
- a stale heartbeat is reported, not enforced: a stalled worker 0 degrades
  to serving its last set (each grant checked against its row) until it
  drains, instead of stopping every claim on the node
- no 2-minute peer wait on a node whose system.hdb_nodes names no other node
- an unreadable row that stops the ascending walk: the rest of the table is
  read from the top down; only if that stops short too is the load partial
- the verification walk applies the walked row directly (skipping keys an
  event touched meanwhile) instead of a point read per suspect, so a
  membership change costs one walk, not a read per gained row
- the head check reads each key at most once a minute
- a row whose last fastRetries+2 leases expired with no result (a renderer
  crashing on it) is filed one cadence forward instead of granted again,
  with no strike (queue_health claim_wedged); the lease slot counts misses

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- the crash-loop bound is now a hold in the lease table (2 leases, then 4,
  8, ... capped at the cadence), not a schedule write a cadence forward: a
  node-wide failure to get results back delays rows by a few leases, and
  the first result for a key clears it. Nothing durable is written
- a walk a resync needs is queued behind a verification already running,
  and retried after a backoff if it fails
- a resubscribe that fails is tried again after a backoff; the node list is
  recorded only once the keeper holds rows by it
- a stale keeper stops counting as serving once its last set has drained,
  so a worker 0 that never came back reports unready, not queued forever
- claims report unready, not empty, while the load has found nothing due
- a config change or failed event during the load is applied when it ends;
  interval changes re-arm the timers while loading
- each walk owns its touched set

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant