Skip to content

feat(console): read the queue keeper — per-node queue state, keeper health, leases; v0.18.0 - #221

Merged
harper-joseph merged 3 commits into
mainfrom
feat/console-keeper
Sep 28, 2026
Merged

harper-joseph merged 3 commits into
mainfrom
feat/console-keeper

Conversation

@harper-joseph

Copy link
Copy Markdown
Contributor

Why

Plugin 0.93.0 (#219) replaced the nextRenderTime index, the claim floor and the ready-set sweep with the queue keeper. Console 0.17.0 still read the removed machinery: its claim-floor, below-floor and prioritisation panels read — or 100%. It also could not see the keeper at all. It could not show whether each node's keeper is serving, could not read queue-state, and could not see the seven new queue_health series.

Deploy order: deploy plugin 0.93.x to every node before console 0.18.0. During a mixed canary the console withholds the cluster queue total and shows the old nodes as "no keeper".

What changed

queue-state is proxied and merged per node (util/aggregate.js mergeQueueState).

  • A 503 is an answer. The keeper's reason, its stats and the live now fields are kept per node. fanOut now keeps a non-200 body as errorBody, apart from body.
  • A node on a plugin before 0.93.0 (404 "Unknown route") is a node without a keeper, not a missing answer.
  • The cluster total is withheld, not floored, until every node's keeper can vouch for its numbers. The view then shows each node's own counts.

Queue view

  • New Queue keeper card with a per-node row:
    • a verdict judged by phase first, because now.status reads empty during starting on 0.93.0;
    • due, in flight, next hour, oldest due and behind;
    • load time, publish time and last verification.
  • Range tiles on that card:
    • repaired: missed writes are bad; rows gained in a membership change are a watch;
    • holds: any is a watch, over 5% of grants is bad;
    • stale skips, publish p95, verify walk and loads.
  • New cards: lateness bins with the most-behind classes, and queue flow for the last hour.
  • Due now and the next-24h histogram come from the keepers when all are live, else from the snapshot, labelled as such.
  • In flight is overview.leases' exact slot walk, with slot pressure (> 75% watch, > 95% bad).
  • A node whose keeper is not serving is the alarm at the top. unready gets its own warning pill.

Health view: new Queue group with seven checks:

  • backlog time to clear
  • queue keeper (worst node)
  • keeper repairs
  • keeper publish p95
  • render holds
  • stale claims
  • lease slots

A failed queue-state read is a bad input check, so Health never reads "All clear" without it.

Removed:

  • claim-floor tiles and lag
  • below-floor alarms (Queue, Health, Inspect)
  • the prioritisation panel
  • the claim-scan check and chart
  • reset-claim-floor guidance
  • the Deep-recompute cap
  • help text naming removed config keys

Guard: a new test requires every queue_health series the console reads (QUEUE_HEALTH in views/queue.js) to be in the plugin catalog, and every catalog series to be read or waived with a reason. I checked that it fails in both directions when mutated.

Verification

  • packages/console: node --test 420/420 (371 before). npm run lint and npm run format:check are clean. Rebased on 0.93.1 (04c0d4b). The plugin package is untouched.
  • Real plugin output (harperfast/harper:5.2.13, packed 0.93.0, 2 threads, 300k seeded rows):
    • Captured overview, analytics, and queue-state both live and in both 503 phases (starting, loading), with renders that never report, so claim_wedged and claim_stale were really emitted.
    • They are committed as test fixtures, and the Queue and Health tests render them.
  • Through the real console proxy (console 0.18.0 packed and installed in the same container, table grown to 1.2M rows so the ~7 s load window was catchable):
    • Configured with two origins of the one instance plus one dead origin: the loading 503s arrived with their reasons and counted as answered, the dead origin showed as not answered, and the total was withheld.
    • With two origins: the cluster summed to 1,207,502 due (2 × 603,751).
    • The Queue and Health views were rendered from exactly what the console served, at cluster and node scope, again after the review fixes.
  • fanOut's 503 plumbing is tested through PrerenderConsole.clusterGet against local HTTP nodes. The test fails with errorBody removed.

Review

One fresh-context adversarial review ran on the diff. Every finding was fixed and tested; none was dismissed.

Likely bugs

# Finding Outcome
1 A membership change read as missed writes: repairs bad, keeper "inexact", backlog "+" for one routine event Fixed. membershipGain matches the last walk to lastResync (node-list why, same repairs, walk inside the resync). Now one watch, "N rows gained after a membership change".
2 "Behind" false-alarmed on small denominators Fixed. Judged only past 100 late rows.
3 In flight silently short when a node's overview lacks leases Fixed. It prefers the keepers' complete sum; else "+" with the node named.
4 "Queue keeper watch" lasted the whole range after a routine restart Fixed. Not-live snapshots up to the number of loads read as "restarted N× in range".
5 Health and Queue disagreed on repairs Fixed. One shared repairsReading.
6 A 0.92 node was judged differently by scope Fixed. "no keeper" (watch) at cluster scope, node scope and for a whole 0.92 cluster. "Management API is disabled" stays no answer.
7 The Queue page hid the keeper card when the overview failed, with a wrong message Fixed. The card is kept, and the message names the overview.
8 At node scope, one fault gave two BADs (keeper plus queueUnavailable backlog) Fixed. Backlog is n/a and points at the Queue keeper check.

Nits

Nit Outcome
drain used Σdue − Σleases Fixed: per node, Σ max(0, due − leases).
Dead merged fields Fixed: byRoute and pausedOn dropped from the merge.
"wedged keys" counted holds Fixed: relabelled Holds / Render holds.
"every second" Fixed: publishInterval, skipped when unchanged, forced every 10 s.
Hardcoded "hourly" Fixed: reads queue.keeper.verifyInterval.
scan.* titled "Claim scan" Fixed: "Scan budgets".
Wrong scan-cap text Fixed.
Behind sub on edgesDiverge Fixed.
"x% of slots" next to a cluster sum Fixed: "fullest node x%".
Class worstNode naming Fixed: host:port like the rest.
Repairs coverage Fixed: "last walk on n/N nodes".

Test gaps: the fanOut test, flow In/Out/Net arithmetic, lateness bin values, and findings 1, 3 and 5 are all covered.

Not covered

  • A real multi-node cluster. The "two nodes" were one instance under two origins, so the per-node rows were identical.
  • The Docker captures are 0.93.0. I did not re-run the e2e on 0.93.1. Its changes are tolerated: phase-first verdicts cover unready on a fresh node, in flight is read from overview.leases either way, and listsTruncated is shown when present (tested).
  • Keeper phases failed, waiting-for-peers and a stale-state 503, and a real keeper_repaired or membership resync: synthetic tests only.
  • A browser. The rendering checks ran through the DOM shim; I did not view it in a browser or take screenshots.
  • Plugin gaps. lateness.listsTruncated is fixed in 0.93.1 (fix(plugin): queue-state counts in-flight exactly, a fresh node reads unready; v0.93.1 #220), and the console reads it when present. Follow-ups:
    • keeper_repaired / exact do not separate rows gained in a membership change from missed writes; the console infers it from lastResync.
    • No metric for lease refusals, the grants a full lease table turns down.
    • Class rows carry no carried flag.
    • explain does not show a key's hold state.
    • No distinct wedged-key count; claim_wedged counts holds.
    • keeper_live rides the backlog snapshot's cadence and carries no reason.

🤖 Generated with Claude Code

harper-joseph and others added 3 commits September 28, 2026 12:28
…ealth, leases; v0.18.0

Catches the console up with plugin 0.93.0, where the queue keeper replaced the
nextRenderTime index, the claim floor and the ready-set sweep.

- queue-state is proxied and merged per node. A 503 is an answer: the keeper's
  reason, stats and live `now` fields are kept per node, and the cluster total
  is withheld (never floored) until every node's keeper can vouch for it.
- Queue: a Queue keeper card (per-node phase and verdict, due, in flight, next
  hour, oldest due, behind, load/publish/verify) with range tiles for repairs,
  wedged keys, stale skips, publish p95, verify and loads; lateness bins and
  the most-behind classes; queue flow for the last hour. Due now and the next
  24h come from the keepers when all are live, else the labelled snapshot.
  In flight is overview.leases' exact slot walk, with slot pressure.
- Health: a Queue group (backlog, queue keeper worst node, keeper repairs,
  publish p95, wedged renders, stale claims, lease slots). A failed queue-state
  read is a bad input check; `unready` is its own warning.
- Removed: claim-floor tiles and lag, below-floor alarms, the prioritisation
  panel, the claim-scan check and chart, reset-claim-floor guidance, the Deep
  recompute cap.
- Guard: every queue_health series the console reads must be in the plugin
  catalog and every catalog series read or waived (QUEUE_HEALTH in queue.js).
- Fixtures: real plugin 0.93.0 payloads (overview, analytics, queue-state live
  and both 503 phases) captured from harperfast/harper:5.2.13.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… a watch, not an outage

During a canary rollout one node runs 0.93.0 and the rest 0.92.x. The old
nodes answer queue-state 404, and they still claim from their own index, so
calling them "no answer" (bad) and raising the grants-no-claims alarm was
wrong. They now read "no keeper" (watch); the cluster total stays withheld.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…nothing silently short

- Rows a membership change hands a node are counted by the plugin's
  verification walk as repairs and turn `exact` false. Told apart from missed
  writes by `trust.keeper.lastResync` (a node-list resync whose walk repaired
  the same rows): Keeper repairs reads "N rows gained after a membership
  change" at watch, the keeper stays ok, and the backlog is not marked "+".
- Repairs come from one shared reading (range sum and each node's last walk)
  on Health and Queue; the sub says how many nodes have walked.
- Behind is judged only past 100 late rows.
- In flight never reads short in silence: a lease sum missing a node yields
  to the keepers' complete sum, else is marked "+" with the node named.
- A restart's not-live snapshot (at most one per load) is explained, not a
  range-long watch.
- A node on a plugin before v0.93.0 ("Unknown route") is "no keeper", a watch,
  at cluster and node scope and for a whole 0.92 cluster; "Management API is
  disabled" stays no answer.
- The Queue page keeps the keeper card when the overview fails, and says so.
- A snapshot with no count because a keeper is not live is n/a on Backlog,
  pointing at the Queue keeper check.
- Waiting rows are summed per node (Σ max(0, due − leases)).
- Holds, not "wedged keys"; publish wording; verifyInterval read from config;
  the scan group is "Scan budgets"; the class list names its node by host;
  unread byRoute/pausedOn dropped from the merge; 0.93.1's listsTruncated shown
  when present.
- Tests: fan-out through PrerenderConsole.clusterGet (fails without
  errorBody), flow arithmetic, lateness bin values, each finding.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces significant updates to the prerender management console, primarily focused on integrating the new queue keeper (plugin v0.93.0). Key changes include replacing the 'claim floor' mechanism with a 'queue keeper' that provides exact, node-local queue state, updating the UI to reflect these changes, and refining how cluster-wide metrics are aggregated. The documentation and tests have been updated to support these architectural shifts. There are no review comments to address.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant