Repository navigation
feat(console): read the queue keeper — per-node queue state, keeper health, leases; v0.18.0 - #221
Merged
Merged
Conversation
…ealth, leases; v0.18.0 Catches the console up with plugin 0.93.0, where the queue keeper replaced the nextRenderTime index, the claim floor and the ready-set sweep. - queue-state is proxied and merged per node. A 503 is an answer: the keeper's reason, stats and live `now` fields are kept per node, and the cluster total is withheld (never floored) until every node's keeper can vouch for it. - Queue: a Queue keeper card (per-node phase and verdict, due, in flight, next hour, oldest due, behind, load/publish/verify) with range tiles for repairs, wedged keys, stale skips, publish p95, verify and loads; lateness bins and the most-behind classes; queue flow for the last hour. Due now and the next 24h come from the keepers when all are live, else the labelled snapshot. In flight is overview.leases' exact slot walk, with slot pressure. - Health: a Queue group (backlog, queue keeper worst node, keeper repairs, publish p95, wedged renders, stale claims, lease slots). A failed queue-state read is a bad input check; `unready` is its own warning. - Removed: claim-floor tiles and lag, below-floor alarms, the prioritisation panel, the claim-scan check and chart, reset-claim-floor guidance, the Deep recompute cap. - Guard: every queue_health series the console reads must be in the plugin catalog and every catalog series read or waived (QUEUE_HEALTH in queue.js). - Fixtures: real plugin 0.93.0 payloads (overview, analytics, queue-state live and both 503 phases) captured from harperfast/harper:5.2.13. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… a watch, not an outage During a canary rollout one node runs 0.93.0 and the rest 0.92.x. The old nodes answer queue-state 404, and they still claim from their own index, so calling them "no answer" (bad) and raising the grants-no-claims alarm was wrong. They now read "no keeper" (watch); the cluster total stays withheld. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nothing silently short
- Rows a membership change hands a node are counted by the plugin's
verification walk as repairs and turn `exact` false. Told apart from missed
writes by `trust.keeper.lastResync` (a node-list resync whose walk repaired
the same rows): Keeper repairs reads "N rows gained after a membership
change" at watch, the keeper stays ok, and the backlog is not marked "+".
- Repairs come from one shared reading (range sum and each node's last walk)
on Health and Queue; the sub says how many nodes have walked.
- Behind is judged only past 100 late rows.
- In flight never reads short in silence: a lease sum missing a node yields
to the keepers' complete sum, else is marked "+" with the node named.
- A restart's not-live snapshot (at most one per load) is explained, not a
range-long watch.
- A node on a plugin before v0.93.0 ("Unknown route") is "no keeper", a watch,
at cluster and node scope and for a whole 0.92 cluster; "Management API is
disabled" stays no answer.
- The Queue page keeps the keeper card when the overview fails, and says so.
- A snapshot with no count because a keeper is not live is n/a on Backlog,
pointing at the Queue keeper check.
- Waiting rows are summed per node (Σ max(0, due − leases)).
- Holds, not "wedged keys"; publish wording; verifyInterval read from config;
the scan group is "Scan budgets"; the class list names its node by host;
unread byRoute/pausedOn dropped from the merge; 0.93.1's listsTruncated shown
when present.
- Tests: fan-out through PrerenderConsole.clusterGet (fails without
errorBody), flow arithmetic, lateness bin values, each finding.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
There was a problem hiding this comment.
Code Review
This pull request introduces significant updates to the prerender management console, primarily focused on integrating the new queue keeper (plugin v0.93.0). Key changes include replacing the 'claim floor' mechanism with a 'queue keeper' that provides exact, node-local queue state, updating the UI to reflect these changes, and refining how cluster-wide metrics are aggregated. The documentation and tests have been updated to support these architectural shifts. There are no review comments to address.
5 of 15 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
Plugin 0.93.0 (#219) replaced the
nextRenderTimeindex, the claim floor and the ready-set sweep with the queue keeper. Console 0.17.0 still read the removed machinery: its claim-floor, below-floor and prioritisation panels read—or 100%. It also could not see the keeper at all. It could not show whether each node's keeper is serving, could not readqueue-state, and could not see the seven newqueue_healthseries.Deploy order: deploy plugin 0.93.x to every node before console 0.18.0. During a mixed canary the console withholds the cluster queue total and shows the old nodes as "no keeper".
What changed
queue-stateis proxied and merged per node (util/aggregate.jsmergeQueueState).nowfields are kept per node.fanOutnow keeps a non-200 body aserrorBody, apart frombody.Queue view
now.statusreadsemptyduringstartingon 0.93.0;overview.leases' exact slot walk, with slot pressure (> 75% watch, > 95% bad).unreadygets its own warning pill.Health view: new Queue group with seven checks:
A failed
queue-stateread is a bad input check, so Health never reads "All clear" without it.Removed:
reset-claim-floorguidancecapGuard: a new test requires every
queue_healthseries the console reads (QUEUE_HEALTHinviews/queue.js) to be in the plugin catalog, and every catalog series to be read or waived with a reason. I checked that it fails in both directions when mutated.Verification
packages/console:node --test420/420 (371 before).npm run lintandnpm run format:checkare clean. Rebased on 0.93.1 (04c0d4b). The plugin package is untouched.harperfast/harper:5.2.13, packed 0.93.0, 2 threads, 300k seeded rows):overview,analytics, andqueue-stateboth live and in both 503 phases (starting,loading), with renders that never report, soclaim_wedgedandclaim_stalewere really emitted.fanOut's 503 plumbing is tested throughPrerenderConsole.clusterGetagainst local HTTP nodes. The test fails witherrorBodyremoved.Review
One fresh-context adversarial review ran on the diff. Every finding was fixed and tested; none was dismissed.
Likely bugs
membershipGainmatches the last walk tolastResync(node-listwhy, same repairs, walk inside the resync). Now one watch, "N rows gained after a membership change".leasesrepairsReading.queueUnavailablebacklog)Nits
byRouteandpausedOndropped from the merge.queue.keeper.verifyInterval.scan.*titled "Claim scan"edgesDivergeworstNodenamingTest gaps: the
fanOuttest, flow In/Out/Net arithmetic, lateness bin values, and findings 1, 3 and 5 are all covered.Not covered
unreadyon a fresh node, in flight is read fromoverview.leaseseither way, andlistsTruncatedis shown when present (tested).failed,waiting-for-peersand a stale-state 503, and a realkeeper_repairedor membership resync: synthetic tests only.lateness.listsTruncatedis fixed in 0.93.1 (fix(plugin): queue-state counts in-flight exactly, a fresh node reads unready; v0.93.1 #220), and the console reads it when present. Follow-ups:keeper_repaired/exactdo not separate rows gained in a membership change from missed writes; the console infers it fromlastResync.carriedflag.explaindoes not show a key's hold state.claim_wedgedcounts holds.keeper_liverides the backlog snapshot's cadence and carries no reason.🤖 Generated with Claude Code