Repository navigation
feat(plugin): the queue keeper is the render queue — nextRenderTime index, claim floor and sweep removed (v0.93.0) - #219
Conversation
…e index and everything built for it are gone; v0.93.0 Worker 0 holds every schedule row this node owns, grouped by class (route x cadence x sitemap flag) into minute buckets, loaded by a local primary-key walk and kept current by a RenderSchedule subscription. RenderSchedule is its durable side, read by primary key only; nextRenderTime is no longer indexed. - claims take from the ready set it publishes each second, point-read each entry's durable row (stale ones skipped, queue_health claim_stale) and lease it; the grant is exclusive (a live lease is refused, never renewed, and a race onto two fresh slots gives one back), so claims need no mutex (#218) - the ready set packs keys end to end, so any key length fits - each publish re-reads its head and repairs; a verification walk every queue.keeper.verifyInterval repairs anything missed (keeper_repaired); a repair never reverts a newer event - waits for a cluster peer before loading, rebuilds on membership, route or default-interval change; a load that cannot get past an unreadable row goes live on what it read, marked not exact - until it is live the node grants no claims and reports queued (keeper_live gauge) GET /prerender_admin/queue-state (new) serves its exact counts; the backlog snapshot reads the same document. Removed: the claim floor and its unpin hatch, the ready-set sweep, the index claim scan, the capped backlog walk, the reset-claim-floor action, config queue.claimFloor.*, queue.claimScanCap, queue.ready.enabled/sweepInterval/ sweepCap and queue.keeper.enabled, and the queue_health series that watched them. overview.claimFloor becomes overview.leases. Refs #215, #216, #217, #218. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request upgrades the prerender plugin to version 0.93.0, introducing a major architectural redesign of the render queue. The previous database-index-based claim scan, claim floor, and ready-set sweep mechanisms have been replaced by an in-memory 'Queue Keeper' service on worker 0. The keeper maintains the queue state in memory, kept current via a table subscription, and publishes the prioritized ready set to shared memory for concurrent claims. Consequently, the secondary index on nextRenderTime in the RenderSchedule table has been removed, and several obsolete configuration options and endpoints have been cleaned up. A new GET /prerender_admin/queue-state endpoint and updated metrics have been introduced to monitor the keeper's health and queue state. I have no feedback to provide as there are no review comments.
…rves claims A node whose keeper is not serving (waiting for peers, loading, retrying a failed load, or gone quiet) reported 'queued', so becoming ready changed nothing and was never broadcast: the fleet found out on its next idle poll, up to backoff.idleMs later, or at the next status sync. It now reports 'unready', and the keeper reports 'queued'/'empty' itself the moment it goes live (and 'unready' when it withdraws), so the fleet is woken by the change. A deployed render fleet treats an unknown status as 'empty' (HostHealth: back off to the idle interval), which is right for a node that can grant nothing. The work-arrived hints (render-now, revalidate) use a new QueueState.noteWork() that only moves empty -> queued, so they no longer lift 'unready' (or 'paused'). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…enders that never report - serves while loading: the ready set is published from the first chunk of the load; events apply live and a walked row is applied only if no event touched its key since the walk began - never reloads once live: a route or default-interval change reclassifies every held row in memory; a membership change reclassifies (dropping rows no longer owned) and runs the verification walk (adding rows gained); a closed subscription is reopened, then walked. Claims are served throughout - a stale heartbeat is reported, not enforced: a stalled worker 0 degrades to serving its last set (each grant checked against its row) until it drains, instead of stopping every claim on the node - no 2-minute peer wait on a node whose system.hdb_nodes names no other node - an unreadable row that stops the ascending walk: the rest of the table is read from the top down; only if that stops short too is the load partial - the verification walk applies the walked row directly (skipping keys an event touched meanwhile) instead of a point read per suspect, so a membership change costs one walk, not a read per gained row - the head check reads each key at most once a minute - a row whose last fastRetries+2 leases expired with no result (a renderer crashing on it) is filed one cadence forward instead of granted again, with no strike (queue_health claim_wedged); the lease slot counts misses Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
- the crash-loop bound is now a hold in the lease table (2 leases, then 4, 8, ... capped at the cadence), not a schedule write a cadence forward: a node-wide failure to get results back delays rows by a few leases, and the first result for a key clears it. Nothing durable is written - a walk a resync needs is queued behind a verification already running, and retried after a backoff if it fails - a resubscribe that fails is tried again after a backoff; the node list is recorded only once the keeper holds rows by it - a stale keeper stops counting as serving once its last set has drained, so a worker 0 that never came back reports unready, not queued forever - claims report unready, not empty, while the load has found nothing due - a config change or failed event during the load is applied when it ends; interval changes re-arm the timers while loading - each walk owns its touched set Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Why
#215: the render queue was the
nextRenderTimeindex, walked three ways — the claim scan, the ready-set sweep (every 5 minutes, about 27 s on a 300k-row due set in production) and the backlog snapshot (every 15 minutes, capped atmanagement.scanCap). Every walk degraded with the store, truncated when the queue was deepest, and was minutes stale. The claim floor, its guard band, periodic reset and unpin hatch existed only to keep the claim scan's seek point off the dead index entries every reschedule leaves behind.This PR makes the queue an in-memory structure — the queue keeper — with
RenderScheduleas its durable side, and deletes the index and everything built for it, with no transitional fallback. The benchmarks in #216 and #217 cleared the approach first: the subscription is exact under replication, at a cost within noise.What changed
The keeper — worker 0 holds every schedule row this node owns by residency.
util/queueKeeper.js(new, pure): rows grouped by class (route × cadence × sitemap flag), each class ordered by due minute. The best K is a k-way merge over class heads usingscoreOfat minute resolution; a test pins it to brute-force scoring and sorting. It computes every queue-state count.util/queueKeeperService.js(new): the lifecycle. Once it has published, nothing stops it serving claims.queue.keeper.publishInterval(1 s), leased rows skipped. Then point-reads the top 64 published rows (each key at most once a minute) and repairs any it holds wrongly.queue.keeper.verifyInterval(1 h): a full walk that applies the table's value wherever the keeper differs (skipping keys an event touched meanwhile) and point-reads held keys the walk did not see. Counted asqueue_healthkeeper_repaired, expected 0. A walk a resync needs is queued behind one already running, and retried if it fails.exact: false, logged).system.hdb_nodesnames another node waits (up to 2 minutes) to see one before loading, since ownership is not knowable before the node list is. A single node loads at once.Claims (
claimSchedules,RenderQueue.claim):queue_healthclaim_stale), and lease the rest. The renderer gets the row's livefromSitemap.util/renderLease.js): a key holding a live lease is refused, never renewed, and after winning a fresh slot a grant rescans the key's probe window and gives the slot back if the key is live elsewhere. So claims need no mutex, and the duplicate-grant path #218 measured is closed on the claim path. Two tests pin it; each fails with its guard removed.stale), not enforced: the last set is still granted from until it drains; only then does the node count as not serving.render.failureRetry.fastRetries + 2leases all expired with no result — a renderer crashing on the URL, or results not reaching the node — is held back in the lease table instead of granted again: 2 leases, then 4, 8, … capped at its cadence. Nothing durable is written, no strike is counted, and the first result for the key clears it; so a node-wide delivery outage delays rows by a few leases, not a cadence. Named in a warning, counted asqueue_healthclaim_wedged. (The removed unpin hatch filed a row a cadence forward once an hour.)Status: a fourth queue status,
unready, while the keeper is not serving or is loading with nothing due found yet. The keeper reports each change itself, so becoming ready is broadcast and the fleet is woken at once (HostHealth.applyMqttStatus), instead of finding out on its next idle poll (15 s) or the 60 s status sync. A deployed render fleet reads an unknown status asempty(idle backoff), which is right for a node that can grant nothing; the console shows it as awarnpill. The work-arrived hints (render-now, revalidate) useQueueState.noteWork(), which only movesempty→queued.Removed, with the index:
nextRenderTime @indexed(schema)lowerFloorForon every write, the guard band, periodic reset,reset-claim-flooraction, unpin hatchclaimFromIndex,runClaimPass, the claim-scan warningssweepReadySet/startReadySweep,createTopKscanUpcomingand thecapoverride onPOST /prerender_admin/backlogqueue.claimFloor.*,queue.claimScanCap,queue.ready.enabled,queue.ready.sweepInterval,queue.ready.sweepCap,queue.keeper.enabledqueue_healthseries:claim_scan_ms,ready_sweep_ms,ready_published,ready_cadence,below_floor,below_floor_age_ms,floor_pin_age_msNew
queue_healthseries:keeper_load_ms,keeper_publish_ms,keeper_verify_ms,keeper_repaired,keeper_live,claim_stale,claim_wedged.Queue state:
GET /prerender_admin/queue-state(new, node-local) —now/coming/lateness/flow/trust, exact, from a shared buffer any worker reads; 503 with the reason until the load finishes or whenever the keeper cannot vouch for its numbers. The README has the field table.Admin API:
overview.claimFlooris replaced byoverview.leases(occupancy, oldest lease);explainandscheduleno longer carryclaimFloor/belowClaimFloor. The backlog snapshot's queue half comes from the keeper (source: 'keeper';queueUnavailablewhile it is not live).Also:
walkUrlRangetakessearchOptionsand tags its cannot-advance throw (WALK_CANNOT_ADVANCE); lease buffer keyrender_queue_v2(layout changed: no floor words, a miss count per slot); README (one "render queue" section replaces four) and METRICS.md updated; a pre-existing customer domain intest/readyQueue.test.jsfixtures replaced withexample.com.Deploying
In place, with no data migration:
RenderSchedulerows are unchanged and the keeper builds from them. Do not start from an empty store; losing the schedule rows re-files every target at a jittered time.—;claim_grantedis taggedready;capis ignored.explainread tolerates the missingclaimFloor.server.nodes, which fills from the localhdb_nodesscan within seconds. The peer count excludes this node the same way harper-pro's replicator does (verified from source; not yet run on harper-pro).unreadyasempty(idle backoff) untilqueuedarrives (verified in browser 1.38.0HostHealth).config.yamlor stored override that still sets a removed key logs "Unknown configuration key" and is otherwise ignored; delete them."nextRenderTime" is not indexed yet.indexing-finished, so any worker thread that opened the table mid-rebuild keeps failing every claim with that error against a complete index. It has been seen in production on a new index, on 12 of 64 threads.isIndexingon the stuck threads).Verification
node --test1385/1385, console 371/371,eslintandprettier --checkclean. Four test files that tested only removed machinery are deleted (renderQueueFloor,renderQueuePin,readySweep,backlogScanCap); their two write-contract tests moved toscheduleWrite.test.js. Claim-path tests (renderQueueVariants,…Redirect,…Targetless,suppressionStatus) claim through a small keeper stand-in (test/support/keeperStandIn.js). New tests cover serving while loading, the in-memory resync on membership and route changes, a walk queued behind a running one, a failed resubscribe retried with backoff, config changes during a load, the unreadable-row tail walk (and the partial case), the walk never applying a row older than an event, the exclusive-grant races, the crash-loop hold growing and resetting, and the stale-but-served claim path.harperfast/harper:5.2.13, packed plugin, 2 worker threads), 2,001 rows seeded under 0.92.0, then the new build installed over it and the container restarted; run five times across the revisions, the last on the final build:describe_tableshowednextRenderTime (indexed); after the swap, no index and the same 2,001 records.QueueStatusrow already readqueued. There was no peer wait: open-source Harper has nohdb_nodes, so the node counts as single. (An earlier revision, before serving while loading and the peer check, answered[]andunreadyfor the 2-minute grace, then flipped toqueuedin the second it went live.)queue-state: due 1,507 of 2,001 (exactly the seeded due count), 1,005 sitemap / 502 discovered,exact: true. A publish took 3.6 ms; the head check read 64 rows and repaired 0.source: keeper, overdue 1,507, not truncated.overview.leasespresent,claimFloorabsent. 0 error lines in the log.topK(5000)p50 0.14–0.27 ms,state()about 0.5 ms, state document about 2.2 KB, heap about 48 MB.queuedforever;emptyreported mid-load; config changes during a load dropped; and a set of stale comments.Not covered
bench/queue-index). It now decides only how long a restart serves a partial order, not how long it grants nothing.countConfiguredPeersis untested against harper-pro (open-source Harper has nohdb_nodes, so the single-node path is what the Docker runs exercised).overview.leases, a queue-state view) and a cluster queue-state fan-out for autoscaling.setPauseand the status sync still use the store mutex.🤖 Generated with Claude Code