Skip to content

Latest commit

 

History

History
444 lines (381 loc) · 38 KB

File metadata and controls

444 lines (381 loc) · 38 KB

Node maintenance

This document covers day-to-day operation of a running QNet Super node: which endpoints to poll and what their fields mean, how to read the node's consensus-position diagnostics, how logging works, how disk usage grows and what the node prunes on its own, what must be backed up, how to upgrade, the coordinated-restart procedure, and how to work through common failure states. Installation and first start are in running-a-node.md; every environment variable named here is described in configuration.md.

What to monitor

All monitoring endpoints are served by the node's single HTTP server on QNET_API_PORT (default 8001, bound on 0.0.0.0, plain HTTP — TLS is terminated upstream). The full reference is rpc-api.md.

Endpoint Use it for Notes
GET /healthz container liveness probe Returns ok h={height} build={build} from a single atomic load and takes no lock; build is the image's QNET_BUILD_ID. This is the probe to wire into the container runtime; GET /health returns a bare OK.
GET /api/v1/node/health the main dashboard record Rich, but touches blockchain, P2P and mempool state; do not use it as a liveness probe.
GET /api/v1/sync/status catch-up progress local_height, network_height, is_syncing, is_ahead, blocks_behind, blocks_ahead, sync_progress, estimated_sync_time.
GET /api/v1/debug/consensus-position finality health See below — the most useful endpoint during an incident.
GET /api/v1/producer/status leadership is_producer, current_producer, leadership_round, next_rotation_height, blocks_until_rotation, computed for the next block.
GET /api/v1/peers connectivity Full peer list plus statistics; appends up to two genesis bootstrap peers when the node holds fewer than three.
GET /api/v1/diagnostics/network transport Peer counts and a QUIC statistics block.
GET /api/v1/mempool/status backlog size is the live mempool count.
GET /api/v1/blocks/stats production cadence current_height, macroblock_height, next_macroblock, blocks_until_macroblock, pending_transactions.
GET /api/v1/failovers failover history Takes limit and from_height; GET /api/v1/network/failovers is an alias onto the same handler.
GET /api/v1/reputation/history?node_id= reputation current_reputation is read from the latest macroblock snapshot, so every node reports the same value.

GET /api/v1/node/health reports status as one of healthy, isolated (zero peers), syncing, degraded (fewer than four validated peers on a non-genesis node), checking (peers present but network height undeterminable). A genesis node with no network-height reading reports sync_status bootstrap and status healthy. Alongside the obvious counters it carries the runtime consensus and clock observability fields: clock_drift_ema_secs, clock_drift_peak_secs, current_timeout_round (0 in steady state, above 0 during BFT failover), max_slot_delay_secs, max_timeout_round_seen, failover_count and timestamp_rejections.

Two API behaviours shape how you scrape a node. Rate-limit rejections come back as HTTP 200, carrying {"success":false,"error":"Rate limit exceeded","retry_after_seconds":…} in the body, so a scraper must read the body. The read_only bucket allows max(QNET_API_RATE_LIMIT × 3, 300) requests per 60 s and blocks for 30 s once exceeded, while 127.0.0.1, ::1 and anything in QNET_WHITELIST_IPS bypass rate limiting — scraping over loopback is the reliable option. The client IP comes from the raw socket, so behind a reverse proxy every request is attributed to the proxy address.

Reading the consensus-position diagnostic

GET /api/v1/debug/consensus-position is the node's own answer to "am I keeping up with finality":

Field Meaning
height, tip_hash local chain tip
own_window height / 90 — the macroblock window this node believes it is in
last_sealed_mb_index index of the newest macroblock this node holds sealed
sealed_lag_windows own_window − last_sealed_mb_index; the headline finality-lag number
finalized_height last finalized height
tc_window_floor observed timeout-certificate window floor
floor_above_window true when the observed floor is ahead of this node's own window
certified_round_current_window highest certified round seen for the current window

A healthy node holds sealed_lag_windows small and stable; a steadily rising value means blocks are still being produced while finality is not advancing. Cross-check current_timeout_round on /api/v1/node/health, where a non-zero value means producer failover is in progress. floor_above_window true on one node while its peers disagree means that node is behind the fleet, not that the fleet is stalled — compare across at least three operators before concluding. Production is bounded while finality is stuck: a node parks once its next block would exceed its seal base — the greater of the last sealed macroblock's height and the QC-verified frontier — by MAX_DERIVED_ROSTER_WINDOWS × MACROBLOCK_INTERVAL = 32 × 90 = 2880 blocks, logging the throttle reason roster_derivation_horizon. See consensus.md for the machinery behind these fields.

Logging

The node writes everything to stdout and stderr. Capture output through the container runtime (docker logs) or your supervisor. Lines are structured as [LEVEL][MODULE] message key=value key2=value2, for example [INFO][BLOCK] produced height=1234 txs=50.

RUST_LOG initialises env_logger at startup and is set to "info" when unset; it governs output emitted through the log crate. The node's own [LEVEL][MODULE] lines are gated by an in-process level that runs at INFO on the scale 0=OFF, 1=ERROR, 2=WARN, 3=INFO, 4=DEBUG, 5=TRACE. Plan on INFO verbosity. High-frequency events are sampled by height rather than printed per block: the common helper logs every 100th block, a second helper every 10th, with heights 0-5 always logged; both fall back to logging every block once the level is raised to DEBUG and TRACE respectively.

Lines worth alerting on directly:

Line Meaning
[FATAL][RESTART] malformed_manifest the release's restart manifest failed its well-formedness check; the node refuses to start
[CRIT][NODE] identity_anchor_mismatch the derived identity key does not match this node's chain anchor; startup is aborted
[FATAL][GEN] WS restart pin active … refusing to mint a restart pin is set but the local chain is empty; the node halts rather than minting a fresh genesis
[CRIT][MEMORY] … OOM_IMMINENT graceful_shutdown memory stayed above the fatal threshold 30 s after an emergency cleanup; the node flushes and exits 137 for the supervisor to restart
[CRIT][STORAGE] … state=critically_full action=admin_required the internal storage budget is at or above 95 %
[WARN][MONITOR] no_peers_connected emitted by the 30-second monitor loop
[INFO][HALT] Reached halt_height=… coordinated-upgrade stop reached
[CRIT][STATE] escalate=halt_signal then [CRIT][NODE] halt_requested the error ladder reached its terminal stage; the node exits 1 for the orchestrator to restart
[CRIT][FAILOVER] … action=self_restart the stuck-height watchdog is spending one of its three restart attempts
[CRIT][FAILOVER] … action=stay_up_degraded reason=restart_budget_exhausted the restart budget is spent; the node stays up and keeps syncing, and the cause is structural
[CRIT][WATCHDOG] chain_stuck … the chain-stuck watchdog fired; alert only, the process keeps running
[CRIT][WATCHDOG] chain_halted … the best height known to this node, its own or the network's, has not moved for 300 s — the whole network is stopped; alert only
[CRIT][WATCHDOG] runtime_stalled … the async runtime missed its heartbeat for 2 s or more; logged once per stall from a separate OS thread with tokio's worker, task and queue counts and RocksDB's write-stop, compaction, flush, memtable and L0 state, which [WARN][PIPELINE] slow_storage_write also carries; runtime_recovered closes the episode
[WARN][MEMORY] rss_floor_rising … the lowest RSS of the last hour exceeds the previous hour's lowest by more than 256 MB — memory that is not released, as opposed to a periodic peak that the next five-minute sample no longer shows

Three of these describe how a node handles its own failure. The error ladder counts consecutive transitions into a recoverable error state and resets on any other transition: at 10 cycles it requests a background resync, at 30 it drops and rediscovers peers, and at 120 (about two minutes) it sets a halt flag that the production loop consumes at the top of its next tick with exit(1), before doing any work. Every stage is signal-based; none touches consensus state, which is why nodes hitting the ladder at different moments still derive the same producer. The stuck-height self-restart fires when a node has been unable to obtain a block for more than 600 seconds while the network holds it: the attempt counter is persisted in the data directory, so the loop cannot reset its own budget by restarting, and past MAX_STUCK_SELF_RESTARTS = 3 the node stays up rather than wiping the RAM consensus state that recovery needs to accumulate. The chain-stuck watchdog deliberately never kills the process: it ticks every 60 s, treats fewer than one block in 300 s as stuck but only while the network is at least 30 blocks ahead, and throttles itself to one alert per stuck window. An operator decision, not a restart, is the intended response.

Disk growth and pruning

A Super node is archival by design: it keeps macroblocks, block hashes, snapshots and full account state for the whole chain, while Light nodes store no chain data at all. Storage is RocksDB across 34 column families with use_fsync enabled, WAL capped at 512 MB, memtables capped at 1 GB in total with up to four per family, RocksDB's own LOG files bounded to 64 MB × 10, one shared 512 MB LRU block cache, and Lz4 compression with Zstd for the cold blocks and snapshots families.

Two derived databases sit inside the same data directory and share that block cache, so a container volume needs no change:

  • state_tree/ holds the account tree (3 families) and is kept across restarts: about 2–3.3 GB at 10M accounts. Its memtables are bounded by 256 MB and its WAL by 256 MB, with fsync on.
  • state_aux/ holds the account leaf preimages and every contract storage tree (5 families). It is deleted and rebuilt at every start, written without a WAL: about 0.9–1.3 GB of preimages at 10M accounts plus about 380 bytes per contract storage slot (about 3.8 GB per 10M token balances). Its memtables are bounded by 128 MB; its write queue by 256 MB.
  • Certified proof views hold at most five RocksDB snapshots per derived DB, about 5–11 GB of superseded rows at 13k transfers per second, and never more than three times the two DBs' live data. They are released, without compaction, when cached usage reaches 95% and resume once a fresh measurement is below 90%.
  • Serving certified proofs (?mb= on the balance-proof routes) adds clamp(cores / 4, 2, 8) threads named qnet-proof-N and a 64 MB answer cache. Proof reads skip the block cache, and the [INFO][PROOFVIEW] stats line every 5 minutes counts what they served, refused and found.
  • The first start of this binary rebuilds state_tree/ through its normal snapshot restore or replay and then empties the main DB's old merkle_leaves and merkle_nodes families. A downgrade before that point finds them intact; after it, the older binary rebuilds them on its own next restore or replay, and a node that can neither restore nor replay (no snapshot, bodies pruned) resyncs from a peer snapshot, as it already does. The two directories stay on disk after a downgrade until the operator deletes them; an upgrade back finds the old families repopulated and rebuilds state_tree/.

The node prunes on two independent schedules and never deletes chain history to free space:

  • Hourly maintenance pass (PRUNE_RUNS_PER_HOUR = 1): ping history and attestations by timestamp; consensus rounds down to the last 1000; failover events on a 24-hour cutoff; snapshots down to the newest SNAPSHOT_KEEP_COUNT = 3 plus the height-90 anchor (the same rule also runs after every frame is written); and transactions, tx_index and tx_by_address below current_height − TX_INDEX_RETENTION_BLOCKS (100,000). Each index sweep resumes from a persisted cursor under a per-run row budget, so retention catches up across runs rather than in one pass. Compaction afterwards is selective: only column families that shed at least COMPACT_MIN_ROWS = 1000 rows are compacted. The pass also strips the committee signatures from every macroblock at least QC_SIG_RETENTION_MB = 14,880 windows below the tip's window, keeping its checkpoint, signer list and sig_merkle_root; it resumes from a persisted cursor and rewrites at most 512 macroblocks per run. Independently of this pass, every failover-event write trims the family to the newest 10,000 rows.
  • Microblock-body prune at every 14,400-block boundary, on the apply path of every Super node, not only the producer. Bodies older than MICROBLOCK_BODY_RETENTION_BLOCKS = 6 × 14,400 = 86,400 blocks are deleted along with their ancestry rows and per-block WASM log rows; macroblock objects, the height→hash aliases, snapshots and account state are kept, and block 0 is never pruned. Each run is bounded by a body_prune_watermark and is idempotent, so a single applied boundary reclaims the whole window. A compile-time assertion guarantees this retention window always exceeds both the snapshot-switch gap and the retained-snapshot span, so a cold or lagging node can never need a body that has been pruned.

History archive

A node started with QNET_ARCHIVE=1 keeps the history the prune would otherwise drop. At the middle of every 14,400-block epoch it writes each epoch that is finalized and not yet archived (normally the previous one, at most six per pass; one right after start, so a restart shows at once whether the archive can write) into <data dir>/archive/seg_{epoch}.qarc, with a seg_{epoch}.json meta beside it (blocks, heights whose body was already gone, macroblocks carried, size, SHA3-256). A block enters a segment only when it hashes to its slot's committed hash, the macroblock's QC-certified window list names it, and its transactions rebuild the merkle root; any mismatch refuses the whole segment ([ERR][HISTORY] segment_refused) and the pass stops there. Producer signatures, VRF proofs and timeout proofs are left out: none is part of the block hash, and an empty block drops from ~7 KB to a few hundred bytes. The macroblocks that certify the epoch go in with their committee signatures and the signers' public keys, so a segment proves finality on its own after the node strips QC signatures from its database (about 20 KB per macroblock with five signers: ~3.3 MB an epoch, ~7 GB a year). A macroblock whose signatures were already stripped when its segment was written is counted in macroblocks_unsigned. Writes are file, fsync, rename, then the meta the same way; a segment without its meta is deleted at the next start.

While the archive owes an epoch, the body prune and the transaction-index prune stop below that epoch, so the archive can still rebuild its blocks. The hold is bounded by ARCHIVE_HOLD_MAX_BLOCKS = 7 days: past it both prunes resume ([ERR][HISTORY] hold_released) and later segments list the lost heights as missing. A rollback that retracts the chain position also deletes every segment reaching above the target. Segments are a node-local, off-consensus copy: apply, sync and finality never read them; the RPC block and header reads fall back to them, and the explorer refills pruned bodies from them. The explorer needs n - quorum + 1 of its endpoints to serve the same body (3 of the five genesis nodes), so enable it on all five: QNET_SET_ENV=QNET_ARCHIVE=1 scripts/deploy-genesis.sh 005 001 002 003 004.

Expect this in API answers: /api/v1/logs, /api/v1/logs/proof and the token-transfer feeds return oldest_available, pruned_below or window_pruned, so an empty result below the prune floor is distinguishable from "no events". Registry and total-supply seals are pruned one 14,400-block window below the head (REGISTRY_SEAL_RETENTION = 14400). On a miss, a registry_root seal is recomputed from scratch, while the checkpoint reader defers on a missing total-supply seal.

The node's own storage monitor runs hourly and needs care. It measures the directory named by QNET_DATA_DIR (./node_data when unset, which may not be the directory the node chose at boot) against QNET_MAX_STORAGE_GB, default 2000 GB for a Super node. That is a configured budget, not the filesystem's free space, and the 70 %, 85 % and 95 % thresholds are percentages of it. All three cleanup tiers clear caches and the transaction pool and force compaction; chain data is kept. So: set QNET_DATA_DIR explicitly, set QNET_MAX_STORAGE_GB near the real volume size, and monitor actual free space externally with df regardless.

Memory is sampled every 300 seconds, logging RSS, virtual size, the delta since the last sample, the allocator's heap_alloc_mb, heap_resident_mb and heap_retained_mb, RocksDB's db_cache_mb, db_memtable_mb and db_readers_mb, and the sizes of the major in-process structures; a second [INFO][MEMORY] census line lists each in-RAM holder's entry count (its size, for _mb names). The limit is derived automatically — the cgroup memory limit at 85 % if visible, otherwise 70 % of total RAM with a 2000 MB floor — and the thresholds are fixed fractions of it: 60 % warn, 75 % emergency (clears both sync queues and the producer cache, forces a transaction-pool cleanup, flushes RocksDB), 90 % fatal. A fatal reading is re-measured 30 s later, and only if it still holds does the node flush and exit(137); a boot within 10 minutes of such an exit waits 15 s, doubling with each further memory exit within 10 minutes of the previous one, up to 120 s. A Super node below 4 GB of RAM (MIN_RAM_SERVER_MB = 4000) refuses to start unless QNET_SKIP_RAM_CHECK is set, and then logs [CRIT][MEMORY] INSUFFICIENT_RAM.

Public HTTPS endpoint

The node's RPC is plain HTTP on :8001. A phone platform that refuses cleartext to a public host — iOS does, and the Android build will once the exception is dropped — can only reach a node through a TLS name. scripts/node-tls.sh puts that in front of each genesis node: a Caddy container in host network mode terminating https://node1.aiqnet.io … node5.aiqnet.io on 443 and proxying to 127.0.0.1:8001, with a Let's Encrypt certificate Caddy obtains and renews itself. The node container is not touched.

The node takes the client address from X-Forwarded-For — the last entry, the one the proxy appended — but only when the socket peer is loopback; a request from any other peer, one that arrives on :8001 directly included, is judged by its socket address. An entry that names the host itself (127.0.0.1, ::1, 0.0.0.0) or is no address was not written by the terminator, and the request counts as 0.0.0.0: never loopback, never whitelisted. Whether the terminator's connections reach the node as loopback, and whether it writes the entry itself rather than passing a client's on, is checked per host before a roll: genesis-host-checks.md, section 1. The genesis-only internal routes never answer loopback.

Order: the node's A record must resolve to its server first (ACME validates over 80/443 of the name; the script refuses to start a terminator the name does not point at), then ./node-tls.sh 001 … 005. A light-node push names no address to answer at, so nothing in the node's environment names its HTTPS address.

Backup and restore

The identity secret is the 12- or 24-word recovery phrase (the mnemonic), and nothing else. The node's ML-DSA-65 keypair is derived deterministically from the mnemonic at every boot. Back the mnemonic up offline. Supply it through QNET_WALLET_SEED_FILE (a file readable only by the node, mode 0600) rather than QNET_WALLET_SEED: a value passed as a container environment variable is readable through docker inspect, through /proc/<pid>/environ and by most log-shipping agents, and the node emits a one-time [WARN][SECURITY] wallet_seed_from_env when you do it that way. A malformed mnemonic fails the structural check and refuses boot rather than deriving a valid-but-wrong identity.

A Super node also needs its activation code (QNET_ACTIVATION_CODE), plus the burn transaction hash and amount on first activation. The code is persisted encrypted with AES-256-GCM under a key derived from the code itself, which is never stored, so the on-disk copy is worthless without the code — keep code and mnemonic together. See node-activation.md. If a node's derived public key does not match its chain anchor, startup aborts with [CRIT][NODE] identity_anchor_mismatch and a hash of both keys for comparison; the remedy is to restore the correct mnemonic.

For chain data:

docker stop qnet-node        # SIGTERM; the node flushes RocksDB and persists certificates
tar -C /path/to/datadir -czf qnet-data-$(date +%F).tar.gz .
docker start qnet-node

Stop the node first. docker stop sends SIGTERM, which the node handles by running storage.flush_all() (WAL to SST) and persisting certificate state before exiting; a hot copy of a live RocksDB directory is not a consistent backup. Restoring is the reverse: stop, replace the directory contents, start. For most incidents a chain-data backup is an optimisation, not a requirement — a wiped node re-joins from peers, snapshot-jumping when it is more than SNAPSHOT_SYNC_SWITCH_GAP = 1500 blocks behind and block-replaying the tail. The one case where it is not optional is a coordinated restart: at least one node must retain state at or below the chosen resume point. A controlled remote stop is available at POST /api/v1/shutdown, which requires an internal caller IP, a configured QNET_ADMIN_SECRET and a matching admin_secret in the body; it flushes and exits 0.

Upgrading a node

Rolling upgrade (no consensus-visible change). Stop the node, replace the image or binary, start it again with QNET_HALT_HEIGHT and the one-shot recovery variables (QNET_ROLLBACK_TO_LAST_SEALED, QNET_ROLLBACK_TO_HEIGHT) unset. It catches up on its own, snapshot-jumping if it fell far enough behind. Do one node at a time and wait until it reports healthy on /api/v1/node/health with blocks_behind at zero on /api/v1/sync/status before touching the next. Track validated_peers while you work: taking down more of the committee than the fault bound tolerates turns a maintenance window into a liveness incident. /healthz answers with the new image's build= once the replacement has taken. scripts/deploy-genesis.sh runs this pass for the genesis fleet: it recreates each container from its own docker inspect output, leaving QNET_ROLLBACK_*, QNET_RECOVERY_HALTED and QNET_BUILD_ID out of the carried environment (QNET_SET_ENV=KEY=VALUE[,KEY=VALUE] adds or replaces entries). The container it replaces is renamed aside rather than deleted, and its log is gzipped to /root/qnet-logs/<container>-pre-<timestamp>.log.gz in the background (newest five kept) before it is dropped, so the hours before a roll can still be read afterwards. It touches the next node only after the previous one answers /healthz, is fewer than 10 blocks below the network height with at least one validated peer, has taken part in a checkpoint as a validator and seen a full-quorum seal at or above that window above its restart height, and has run 30 blocks past its restart with a failover-free metrics window. QNET_RECOVERY_HALTED=1 skips the last two gates for a roll that repairs a halted chain.

Gated rule change (rolling). A consensus rule that ships behind a feature gate rolls like an ordinary upgrade. The release carries the rule dormant with an activation height compiled into the binary, every node flips it at that height, and the operator's whole job is to have the new binary deployed fleet-wide before the height arrives — so treat the activation height as the deadline for the rolling pass above. The current gates and their heights are in consensus.md; a release note that names one is telling you when the deployment must be finished. A crossed gate never moves: only a gate the fleet has not reached may be moved to a later epoch boundary, before the first roll of its build, and scripts/deploy-genesis.sh refuses to move one whose old height is not GATE_MARGIN (28,800 blocks) above the fleet tip ("gate already crossed, its height is frozen").

Protocol-breaking upgrade. A change that alters what any node considers valid and is not carried by a gate cannot be rolled. Use the halt-height mechanism: set the same QNET_HALT_HEIGHT on every node. The 30-second monitor loop compares the current height against it, and at or above that height the node flushes storage and exits 0, logging [INFO][HALT] Reached halt_height=…. Operators then swap binaries and restart with the variable removed. Publish the halt height, the release commit and the restart wall-clock time ahead of time.

Un-barring identities from a restart manifest is also a coordinated cut-over, never a rolling change. The exclusion list filters the derived producer and committee set, which feeds consensus_committee and from there the QC-bound epoch_commitment. If some nodes un-bar a still-heartbeating identity before others, the two groups derive different committees for the crossover windows and disagree on those checkpoints until the fleet converges. Publish a wall-clock cut-over, and prefer un-barring only identities already heartbeat-absent, where the removal is a no-op.

Coordinated restart

This is the recovery procedure for a halted or forged chain: the fleet agrees on the newest macroblock everyone holds, pins it in a new release, and resumes production from it. Rehearse it on a test network before you need it — an unrehearsed runbook is not a recovery plan.

When to restart

Restart only when both of these hold. Anything less is a bug to fix, not an incident:

  1. Finality has not advanced for more than two hours. Production stops on its own once the roster_derivation_horizon (2880 blocks past the last seal) is reached.
  2. There is no software fix that restores liveness without abandoning chain data, and rolling the fleet back to its last sealed macroblock (see Fleet stalled or forked above its last seal below) does not restore it either.

For a forged finality incident — a committee majority certified a bad state_root — the trigger is different: restart as soon as it is confirmed, and pick K strictly below the first bad macroblock.

Procedure

  1. Freeze and gather. Stop producing. From every reachable operator collect the last macroblock index held and its MacroBlock::hash(), and the consensus_committee of the last sealed macroblock.
  2. Choose K — the newest macroblock that is full-quorum sealed and that every surviving operator agrees on by hash; never one that only some nodes hold. Record hash(MB_K) and the committee-fields digests for K and K-1; the release needs all four values.
  3. Build the exclusion list — the identities that were in the committee at the stall and did not vote. Assemble it off-chain from operator logs and publish it with the evidence so anyone can disagree before the release ships. List only identities multiple independent operators observed as silent, prefer under-listing (a restart that leaves a few passive identities in still recovers if the survivors clear quorum), and keep the list sorted and deduplicated — the build check fails otherwise.
  4. Cut the release. In development/qnet-integration/src/genesis_constants.rs set together WS_CHECKPOINT = (K, hash_of_MB_K), WS_CHECKPOINT_DIGEST_ANCHOR, WS_CHECKPOINT_DIGEST_PRED, and RESTART_MANIFEST { resume_from_mb: K, resume_mb_hash: hash_of_MB_K, excluded: &[…] }. restart_manifest_is_wellformed() runs before storage opens and refuses to start the node if the manifest disagrees with the pin index or hash, if either digest is zero, or if the list is unsorted or contains duplicates — a malformed manifest is a broken release, not a runtime condition.
  5. Publish before executing, and confirm a retained-state source. Post in one place: K, hash(MB_K), both digests, the exclusion list with per-entry evidence, the release commit hash and build instructions, and the wall-clock start time — a restart without a published record is indistinguishable from an attack on the chain. Mandatory precondition: at least one reachable node must retain chain state at or below K and serve its snapshot, because balances at K > 0 survive only as retained chain data. Designate that archival node and confirm it answers a K-height snapshot request before anyone wipes. If none retains state at or below K, the ledger is unrecoverable; do not proceed.
  6. Execute. Every operator, genesis included, stops their node and wipes chain data above K only. Do not wipe a data directory entirely unless the archival node is confirmed serving state at or below K: a full wipe everywhere destroys the ledger, and the boot guard halts an empty node under the pin rather than minting a fresh genesis. A node that keeps stale data above K fails closed on the pin path (v2_ws_pin_mismatch, v2_below_ws) rather than forking. Then run the new release — archival node and genesis nodes first, then the rest.
  7. Verify recovery. Blocks are produced and finality advances (last_sealed_mb_index rises on several nodes); the eligible_producers of the first newly sealed macroblock contains none of the barred identities; no node logs [FATAL][RESTART] malformed_manifest; the tip hash matches across at least three operators.
  8. Retire the manifest. Once the chain has been stable for a full epoch, the next release keeps the bumped WS_CHECKPOINT and clears RESTART_MANIFEST back to an empty exclusion list only if the barred identities should be allowed to re-register. Leaving them barred is a policy choice — state it publicly either way, and follow the coordinated cut-over rule above.

Rehearsal checklist

Run the whole thing on a test network, timed, before you need it in production:

  • Halt a test network deliberately by stopping more than one third of the committee.
  • Confirm production stops at the expected horizon (2880 blocks past the last seal).
  • Choose K and verify the hash agrees across operators.
  • Cut a release with a non-empty exclusion list.
  • Confirm a node with stale data above K fails closed instead of forking.
  • Confirm a malformed manifest refuses to start.
  • Measure the wall-clock time from decision to restored finality, and publish it.

Common failure states

Node stuck syncing. Check blocks_behind on /api/v1/sync/status and whether it is falling. If flat, check /api/v1/peers and /api/v1/node/health: isolated means zero peers (firewall or discovery — see networking.md and the port table in configuration.md); degraded means fewer than four validated peers, enough to sync but not to participate safely. A node more than 1500 blocks behind should snapshot-jump rather than block-replay; if it does not, check that a peer actually serves one (GET /api/v1/snapshot/latest against that peer) — snapshot serving is capped at 16 concurrent transfers node-wide and answers snapshot serve busy beyond that, so a fleet-wide restart can starve joiners temporarily. Restarting the stuck node is safe and is the usual first move; it resumes from its persisted height.

Node isolated after moving to a new address. Peers bind a registered identity to the IP of the endpoint it committed on chain, and inbound connections arriving from anywhere else are refused before any signature work. Confirm QNET_PUBLIC_IP on the new host names the new public address, then reactivate (POST /api/v1/node-reactivation/submit) — applying that transaction refreshes the committed endpoint, and peers pick up the new binding as the block applies. Each peer keeps the map on disk and rebuilds it at boot, so a peer that restarts before the reactivation lands still holds the old address until it does. See running-a-node.md.

Node not producing. Confirm it should be: GET /api/v1/producer/status reports is_producer for the next block and names current_producer. Eligibility comes from the eligible_producers snapshot of an earlier macroblock, so a freshly registered Super is ineligible until ACTIVATION_WARMUP_BLOCKS = 180 blocks have passed, and reputation must be at or above the consensus minimum of 70 (/api/v1/reputation/history). If the node is elected but no blocks appear, look at current_timeout_round and failover_count on /api/v1/node/health — the network is rotating away from it — and at clock_drift_ema_secs and timestamp_rejections. If sealed_lag_windows is climbing past the horizon the node has parked on roster_derivation_horizon, and the problem is network-wide finality, not this node.

Fleet stalled or forked above its last seal. When every branch agrees up to the last sealed macroblock and only the unsealed tail differs or stops, roll the fleet back to that point instead of cutting a restart release. Set QNET_ROLLBACK_TO_LAST_SEALED=1 (or QNET_ROLLBACK_TO_HEIGHT=<h>) on every node, start them together, and remove the variable before the next start. No rollback goes below what a node holds certified: LAST_SEALED means that certified floor, and an explicit height below it is refused ([ERR][ROLLBACK] refused … reason=certified_checkpoint_irrevocable, the node starts unchanged), because a certificate is n−f signatures and any surviving copy pulls the fleet back to it. At boot, before state recovery, each node truncates to that height, retracts the macroblocks and certified pairs above it, drops its own vote commitments above it and lowers its anti-double-sign mark to the height it ends at, so it can sign the re-produced windows. Run with either variable, scripts/deploy-genesis.sh sets it on the containers it recreates and strips it on its next roll; until then a restart of such a container repeats the rollback. A recovery decree prunes from one host: on each genesis node node_decreeEndorse (params seq, target_height) returns that node's consensus signature over the decree, and node_decreeSubmit (the same params plus sigs) with signatures from a quorum of the genesis consensus keys and a seq above the last applied one gossips it; every node that verifies it deletes the blocks above target_height, retracts the macroblocks and certified pairs above it, records the seq and exits for a clean boot. A decree whose target is below a node's certified floor is refused by that node (and by node_decreeSubmit), which records its seq so it is not gossiped again. Both methods answer internal callers only. For a fork in which no branch holds a quorum, QNET_ROLLBACK_TO_HEIGHT brings every stored marker back to one height at boot, never below the node's certified floor, and the node rejoins from the network.

Storage full. Distinguish the two meanings. If the filesystem is full the node cannot flush and should be stopped before it is starved; free space outside the data directory, then restart. If the node logs storage_warn_85pct_full or critically_full, that is the internal budget against QNET_MAX_STORAGE_GB, whose cleanups touch caches only. The real levers: confirm the body prune is running ([INFO][STORAGE] microblock_bodies_pruned appears at 14,400-block boundaries — a node that has not crossed one since starting has not pruned yet), confirm QNET_DATA_DIR points at the directory the node really uses, then provision more disk. Never hand-delete files from the RocksDB directory.

Reward-epoch commitment deferring. The checkpoint's reward_epoch_root folds every certified epoch root up to the N-2 macroblock, resuming from a persisted prefix, and defers rather than seal a shorter set. Two log lines name what it is waiting on. [WARN][REWARDS] epoch_root_gap … action=defer+repair means the macroblock is simply absent and the node has already fired a targeted repair fetch — it resolves itself; watch that the named missing_mb stops recurring. [ERR][REWARDS] epoch_root_mb_no_usable_qc … action=operator_resync means the macroblock is on this node's disk but unreadable, and it needs you: the object is QC-certified and sits at or below that N-2 macroblock, so the node keeps it rather than deleting it, and forward sync never revisits that height. Resync this node from a snapshot; do not hand-delete anything from the RocksDB directory. A third line, epoch_root_target_unknown … action=defer_no_repair, means a storage read itself failed — treat it as a disk fault on this node.

Node refuses to start. Read the first fatal line. malformed_manifest is a bad release — rebuild it, do not work around it. identity_anchor_mismatch is a wrong or lost mnemonic — restore the correct one, never "fix" it by editing the anchor. WS restart pin active … refusing to mint means the node has a restart pin but no local chain and must cold-join from the resume macroblock rather than mint a fresh genesis. A repeated exit 137 is the memory ceiling, not a crash: give the container more memory or reduce what else runs on the host.

Choosing K, deciding who goes on the exclusion list, and judging whether a stall is a bug or an incident are operator decisions taken off-chain. Say so to the other operators and publish the evidence: getting them wrong bars an honest operator or abandons real state.