This document covers day-to-day operation of a running QNet Super node: which endpoints to poll and what their fields mean, how to read the node's consensus-position diagnostics, how logging works, how disk usage grows and what the node prunes on its own, what must be backed up, how to upgrade, the coordinated-restart procedure, and how to work through common failure states. Installation and first start are in running-a-node.md; every environment variable named here is described in configuration.md.
All monitoring endpoints are served by the node's single HTTP server on QNET_API_PORT (default 8001, bound on
0.0.0.0, plain HTTP — TLS is terminated upstream). The full reference is rpc-api.md.
| Endpoint | Use it for | Notes |
|---|---|---|
GET /healthz |
container liveness probe | Returns ok h={height} build={build} from a single atomic load and takes no lock; build is the image's QNET_BUILD_ID. This is the probe to wire into the container runtime; GET /health returns a bare OK. |
GET /api/v1/node/health |
the main dashboard record | Rich, but touches blockchain, P2P and mempool state; do not use it as a liveness probe. |
GET /api/v1/sync/status |
catch-up progress | local_height, network_height, is_syncing, is_ahead, blocks_behind, blocks_ahead, sync_progress, estimated_sync_time. |
GET /api/v1/debug/consensus-position |
finality health | See below — the most useful endpoint during an incident. |
GET /api/v1/producer/status |
leadership | is_producer, current_producer, leadership_round, next_rotation_height, blocks_until_rotation, computed for the next block. |
GET /api/v1/peers |
connectivity | Full peer list plus statistics; appends up to two genesis bootstrap peers when the node holds fewer than three. |
GET /api/v1/diagnostics/network |
transport | Peer counts and a QUIC statistics block. |
GET /api/v1/mempool/status |
backlog | size is the live mempool count. |
GET /api/v1/blocks/stats |
production cadence | current_height, macroblock_height, next_macroblock, blocks_until_macroblock, pending_transactions. |
GET /api/v1/failovers |
failover history | Takes limit and from_height; GET /api/v1/network/failovers is an alias onto the same handler. |
GET /api/v1/reputation/history?node_id= |
reputation | current_reputation is read from the latest macroblock snapshot, so every node reports the same value. |
GET /api/v1/node/health reports status as one of healthy, isolated (zero peers), syncing, degraded (fewer
than four validated peers on a non-genesis node), checking (peers present but network height undeterminable).
A genesis node with no network-height reading reports sync_status bootstrap and status healthy. Alongside the obvious counters it carries the runtime consensus and clock observability fields:
clock_drift_ema_secs, clock_drift_peak_secs, current_timeout_round (0 in steady state, above 0 during BFT
failover), max_slot_delay_secs, max_timeout_round_seen, failover_count and timestamp_rejections.
Two API behaviours shape how you scrape a node. Rate-limit rejections come back as HTTP 200, carrying
{"success":false,"error":"Rate limit exceeded","retry_after_seconds":…} in the body, so a scraper must read the body.
The read_only bucket allows max(QNET_API_RATE_LIMIT × 3, 300) requests per 60 s and blocks for 30 s once exceeded,
while 127.0.0.1, ::1 and anything in QNET_WHITELIST_IPS bypass rate limiting — scraping over loopback is the
reliable option. The client IP comes from the raw socket, so behind a reverse proxy every request is attributed to
the proxy address.
GET /api/v1/debug/consensus-position is the node's own answer to "am I keeping up with finality":
| Field | Meaning |
|---|---|
height, tip_hash |
local chain tip |
own_window |
height / 90 — the macroblock window this node believes it is in |
last_sealed_mb_index |
index of the newest macroblock this node holds sealed |
sealed_lag_windows |
own_window − last_sealed_mb_index; the headline finality-lag number |
finalized_height |
last finalized height |
tc_window_floor |
observed timeout-certificate window floor |
floor_above_window |
true when the observed floor is ahead of this node's own window |
certified_round_current_window |
highest certified round seen for the current window |
A healthy node holds sealed_lag_windows small and stable; a steadily rising value means blocks are still being
produced while finality is not advancing. Cross-check current_timeout_round on /api/v1/node/health, where a
non-zero value means producer failover is in progress. floor_above_window true on one node while its peers disagree
means that node is behind the fleet, not that the fleet is stalled — compare across at least three operators before
concluding. Production is bounded while finality is stuck: a node parks once its next block would exceed its seal
base — the greater of the last sealed macroblock's height and the QC-verified frontier — by
MAX_DERIVED_ROSTER_WINDOWS × MACROBLOCK_INTERVAL = 32 × 90 = 2880 blocks, logging the throttle reason
roster_derivation_horizon. See consensus.md for the machinery behind these fields.
The node writes everything to stdout and stderr. Capture output through the container runtime (docker logs) or your
supervisor. Lines are structured as [LEVEL][MODULE] message key=value key2=value2, for example
[INFO][BLOCK] produced height=1234 txs=50.
RUST_LOG initialises env_logger at startup and is set to "info" when unset; it governs output emitted through the
log crate. The node's own [LEVEL][MODULE] lines are gated by an in-process level that runs at INFO on the scale
0=OFF, 1=ERROR, 2=WARN, 3=INFO, 4=DEBUG, 5=TRACE. Plan on INFO verbosity. High-frequency events are sampled by height
rather than printed per block: the common helper logs every 100th block, a second helper every 10th, with heights 0-5
always logged; both fall back to logging every block once the level is raised to DEBUG and TRACE respectively.
Lines worth alerting on directly:
| Line | Meaning |
|---|---|
[FATAL][RESTART] malformed_manifest |
the release's restart manifest failed its well-formedness check; the node refuses to start |
[CRIT][NODE] identity_anchor_mismatch |
the derived identity key does not match this node's chain anchor; startup is aborted |
[FATAL][GEN] WS restart pin active … refusing to mint |
a restart pin is set but the local chain is empty; the node halts rather than minting a fresh genesis |
[CRIT][MEMORY] … OOM_IMMINENT graceful_shutdown |
memory stayed above the fatal threshold 30 s after an emergency cleanup; the node flushes and exits 137 for the supervisor to restart |
[CRIT][STORAGE] … state=critically_full action=admin_required |
the internal storage budget is at or above 95 % |
[WARN][MONITOR] no_peers_connected |
emitted by the 30-second monitor loop |
[INFO][HALT] Reached halt_height=… |
coordinated-upgrade stop reached |
[CRIT][STATE] escalate=halt_signal then [CRIT][NODE] halt_requested |
the error ladder reached its terminal stage; the node exits 1 for the orchestrator to restart |
[CRIT][FAILOVER] … action=self_restart |
the stuck-height watchdog is spending one of its three restart attempts |
[CRIT][FAILOVER] … action=stay_up_degraded reason=restart_budget_exhausted |
the restart budget is spent; the node stays up and keeps syncing, and the cause is structural |
[CRIT][WATCHDOG] chain_stuck … |
the chain-stuck watchdog fired; alert only, the process keeps running |
[CRIT][WATCHDOG] chain_halted … |
the best height known to this node, its own or the network's, has not moved for 300 s — the whole network is stopped; alert only |
[CRIT][WATCHDOG] runtime_stalled … |
the async runtime missed its heartbeat for 2 s or more; logged once per stall from a separate OS thread with tokio's worker, task and queue counts and RocksDB's write-stop, compaction, flush, memtable and L0 state, which [WARN][PIPELINE] slow_storage_write also carries; runtime_recovered closes the episode |
[WARN][MEMORY] rss_floor_rising … |
the lowest RSS of the last hour exceeds the previous hour's lowest by more than 256 MB — memory that is not released, as opposed to a periodic peak that the next five-minute sample no longer shows |
Three of these describe how a node handles its own failure. The error ladder counts consecutive
transitions into a recoverable error state and resets on any other transition: at 10 cycles it requests a
background resync, at 30 it drops and rediscovers peers, and at 120 (about two minutes) it sets a halt flag
that the production loop consumes at the top of its next tick with exit(1), before doing any work. Every
stage is signal-based; none touches consensus state, which is why nodes hitting the ladder at different
moments still derive the same producer. The stuck-height self-restart fires when a node has been unable
to obtain a block for more than 600 seconds while the network holds it: the attempt counter is persisted in
the data directory, so the loop cannot reset its own budget by restarting, and past MAX_STUCK_SELF_RESTARTS = 3 the node stays up rather than wiping the RAM consensus state that recovery needs to accumulate. The
chain-stuck watchdog deliberately never kills the process: it ticks every 60 s, treats fewer than one
block in 300 s as stuck but only while the network is at least 30 blocks ahead, and throttles itself to one
alert per stuck window. An operator decision, not a restart, is the intended response.
A Super node is archival by design: it keeps macroblocks, block hashes, snapshots and full account state for the whole
chain, while Light nodes store no chain data at all. Storage is RocksDB across 34 column families with use_fsync
enabled, WAL capped at 512 MB, memtables capped at 1 GB in total with up to four per family, RocksDB's own LOG files
bounded to 64 MB × 10, one shared 512 MB LRU block cache, and Lz4 compression with Zstd for the cold blocks and
snapshots families.
Two derived databases sit inside the same data directory and share that block cache, so a container volume needs no change:
state_tree/holds the account tree (3 families) and is kept across restarts: about 2–3.3 GB at 10M accounts. Its memtables are bounded by 256 MB and its WAL by 256 MB, with fsync on.state_aux/holds the account leaf preimages and every contract storage tree (5 families). It is deleted and rebuilt at every start, written without a WAL: about 0.9–1.3 GB of preimages at 10M accounts plus about 380 bytes per contract storage slot (about 3.8 GB per 10M token balances). Its memtables are bounded by 128 MB; its write queue by 256 MB.- Certified proof views hold at most five RocksDB snapshots per derived DB, about 5–11 GB of superseded rows at 13k transfers per second, and never more than three times the two DBs' live data. They are released, without compaction, when cached usage reaches 95% and resume once a fresh measurement is below 90%.
- Serving certified proofs (
?mb=on the balance-proof routes) addsclamp(cores / 4, 2, 8)threads namedqnet-proof-Nand a 64 MB answer cache. Proof reads skip the block cache, and the[INFO][PROOFVIEW] statsline every 5 minutes counts what they served, refused and found. - The first start of this binary rebuilds
state_tree/through its normal snapshot restore or replay and then empties the main DB's oldmerkle_leavesandmerkle_nodesfamilies. A downgrade before that point finds them intact; after it, the older binary rebuilds them on its own next restore or replay, and a node that can neither restore nor replay (no snapshot, bodies pruned) resyncs from a peer snapshot, as it already does. The two directories stay on disk after a downgrade until the operator deletes them; an upgrade back finds the old families repopulated and rebuildsstate_tree/.
The node prunes on two independent schedules and never deletes chain history to free space:
- Hourly maintenance pass (
PRUNE_RUNS_PER_HOUR = 1): ping history and attestations by timestamp; consensus rounds down to the last 1000; failover events on a 24-hour cutoff; snapshots down to the newestSNAPSHOT_KEEP_COUNT = 3plus the height-90 anchor (the same rule also runs after every frame is written); and transactions,tx_indexandtx_by_addressbelowcurrent_height − TX_INDEX_RETENTION_BLOCKS (100,000). Each index sweep resumes from a persisted cursor under a per-run row budget, so retention catches up across runs rather than in one pass. Compaction afterwards is selective: only column families that shed at leastCOMPACT_MIN_ROWS = 1000rows are compacted. The pass also strips the committee signatures from every macroblock at leastQC_SIG_RETENTION_MB = 14,880windows below the tip's window, keeping its checkpoint, signer list andsig_merkle_root; it resumes from a persisted cursor and rewrites at most 512 macroblocks per run. Independently of this pass, every failover-event write trims the family to the newest 10,000 rows. - Microblock-body prune at every 14,400-block boundary, on the apply path of every Super node, not only the
producer. Bodies older than
MICROBLOCK_BODY_RETENTION_BLOCKS = 6 × 14,400 = 86,400blocks are deleted along with their ancestry rows and per-block WASM log rows; macroblock objects, the height→hash aliases, snapshots and account state are kept, and block 0 is never pruned. Each run is bounded by abody_prune_watermarkand is idempotent, so a single applied boundary reclaims the whole window. A compile-time assertion guarantees this retention window always exceeds both the snapshot-switch gap and the retained-snapshot span, so a cold or lagging node can never need a body that has been pruned.
A node started with QNET_ARCHIVE=1 keeps the history the prune would otherwise drop. At the middle of every
14,400-block epoch it writes each epoch that is finalized and not yet archived (normally the previous one, at most
six per pass; one right after start, so a restart shows at once whether the archive can write) into
<data dir>/archive/seg_{epoch}.qarc, with a seg_{epoch}.json meta beside it (blocks, heights whose body was already
gone, macroblocks carried, size, SHA3-256). A block enters a segment only when it hashes to its slot's committed
hash, the macroblock's QC-certified window list names it, and its transactions rebuild the merkle root; any
mismatch refuses the whole segment ([ERR][HISTORY] segment_refused) and the pass stops there. Producer signatures,
VRF proofs and timeout proofs are left out: none is part of the block hash, and an empty block drops from ~7 KB to a
few hundred bytes. The macroblocks that certify the epoch go in with their committee signatures and the signers'
public keys, so a segment proves finality on its own after the node strips QC signatures from its database
(about 20 KB per macroblock with five signers: ~3.3 MB an epoch, ~7 GB a year). A macroblock whose signatures were
already stripped when its segment was written is counted in macroblocks_unsigned. Writes are file, fsync, rename, then the meta the same way; a segment without its meta is deleted
at the next start.
While the archive owes an epoch, the body prune and the transaction-index prune stop below that epoch, so the
archive can still rebuild its blocks. The hold is bounded by ARCHIVE_HOLD_MAX_BLOCKS = 7 days: past it both prunes
resume ([ERR][HISTORY] hold_released) and later segments list the lost heights as missing. A rollback that retracts
the chain position also deletes every segment reaching above the target. Segments are a node-local, off-consensus
copy: apply, sync and finality never read them; the RPC block and header reads fall back to them, and the explorer
refills pruned bodies from them. The explorer needs n - quorum + 1 of its endpoints to serve the same body (3 of the five
genesis nodes), so enable it on all five: QNET_SET_ENV=QNET_ARCHIVE=1 scripts/deploy-genesis.sh 005 001 002 003 004.
Expect this in API answers: /api/v1/logs, /api/v1/logs/proof and the token-transfer feeds return
oldest_available, pruned_below or window_pruned, so an empty result below the prune floor is distinguishable
from "no events". Registry and total-supply seals are pruned one 14,400-block window below the head
(REGISTRY_SEAL_RETENTION = 14400). On a miss, a registry_root seal is recomputed from scratch, while the
checkpoint reader defers on a missing total-supply seal.
The node's own storage monitor runs hourly and needs care. It measures the directory named by QNET_DATA_DIR
(./node_data when unset, which may not be the directory the node chose at boot) against QNET_MAX_STORAGE_GB,
default 2000 GB for a Super node. That is a configured budget, not the filesystem's free space, and the 70 %, 85 % and
95 % thresholds are percentages of it. All three cleanup tiers clear caches and the transaction pool and force
compaction; chain data is kept. So: set QNET_DATA_DIR explicitly, set QNET_MAX_STORAGE_GB near the real volume
size, and monitor actual free space externally with df regardless.
Memory is sampled every 300 seconds, logging RSS, virtual size, the delta since the last sample, the allocator's
heap_alloc_mb, heap_resident_mb and heap_retained_mb, RocksDB's db_cache_mb, db_memtable_mb and
db_readers_mb, and the sizes of the major in-process structures; a second [INFO][MEMORY] census line lists each
in-RAM holder's entry count (its size, for _mb names). The limit is derived automatically — the cgroup memory limit
at 85 % if visible, otherwise 70 % of total RAM with a 2000 MB floor — and the thresholds are fixed fractions of it:
60 % warn, 75 % emergency (clears both sync queues and the producer cache, forces a transaction-pool cleanup, flushes
RocksDB), 90 % fatal. A fatal reading is re-measured 30 s later, and only if it still holds does the node flush and
exit(137); a boot within 10 minutes of such an exit waits 15 s, doubling with each further memory exit within 10
minutes of the previous one, up to 120 s. A Super node below 4 GB of RAM (MIN_RAM_SERVER_MB = 4000) refuses to
start unless QNET_SKIP_RAM_CHECK is set, and then logs [CRIT][MEMORY] INSUFFICIENT_RAM.
The node's RPC is plain HTTP on :8001. A phone platform that refuses cleartext to a public host — iOS
does, and the Android build will once the exception is dropped — can only reach a node through a TLS
name. scripts/node-tls.sh puts that in front of each genesis node: a Caddy container in host network
mode terminating https://node1.aiqnet.io … node5.aiqnet.io on 443 and proxying to 127.0.0.1:8001,
with a Let's Encrypt certificate Caddy obtains and renews itself. The node container is not touched.
The node takes the client address from X-Forwarded-For — the last entry, the one the proxy appended —
but only when the socket peer is loopback; a request from any other peer, one that arrives on :8001
directly included, is judged by its socket address. An entry that names the host itself (127.0.0.1, ::1,
0.0.0.0) or is no address was not written by the terminator, and the request counts as 0.0.0.0: never
loopback, never whitelisted. Whether the terminator's connections reach the node as loopback, and whether it
writes the entry itself rather than passing a client's on, is checked per host before a roll:
genesis-host-checks.md, section 1. The genesis-only internal routes never answer
loopback.
Order: the node's A record must resolve to its server first (ACME validates over 80/443 of the name; the
script refuses to start a terminator the name does not point at), then ./node-tls.sh 001 … 005. A
light-node push names no address to answer at, so nothing in the node's environment names its HTTPS
address.
The identity secret is the 12- or 24-word recovery phrase (the mnemonic), and nothing else. The node's ML-DSA-65 keypair is derived
deterministically from the mnemonic at every boot. Back the mnemonic up offline. Supply it through
QNET_WALLET_SEED_FILE (a file readable only by the node, mode 0600) rather than QNET_WALLET_SEED: a value passed as
a container environment variable is readable through docker inspect, through /proc/<pid>/environ and by most
log-shipping agents, and the node emits a one-time [WARN][SECURITY] wallet_seed_from_env when you do it that way. A
malformed mnemonic fails the structural check and refuses boot rather than deriving a valid-but-wrong identity.
A Super node also needs its activation code (QNET_ACTIVATION_CODE), plus the burn transaction hash and amount on
first activation. The code is persisted encrypted with AES-256-GCM under a key derived from the code itself, which is
never stored, so the on-disk copy is worthless without the code — keep code and mnemonic together. See
node-activation.md. If a node's derived public key does not match its chain anchor,
startup aborts with [CRIT][NODE] identity_anchor_mismatch and a hash of both keys for comparison; the remedy is to
restore the correct mnemonic.
For chain data:
docker stop qnet-node # SIGTERM; the node flushes RocksDB and persists certificates
tar -C /path/to/datadir -czf qnet-data-$(date +%F).tar.gz .
docker start qnet-nodeStop the node first. docker stop sends SIGTERM, which the node handles by running storage.flush_all() (WAL to SST)
and persisting certificate state before exiting; a hot copy of a live RocksDB directory is not a consistent backup.
Restoring is the reverse: stop, replace the directory contents, start. For most incidents a chain-data backup is an
optimisation, not a requirement — a wiped node re-joins from peers, snapshot-jumping when it is more than
SNAPSHOT_SYNC_SWITCH_GAP = 1500 blocks behind and block-replaying the tail. The one case where it is not optional is
a coordinated restart: at least one node must retain state at or below the chosen resume point. A controlled remote
stop is available at POST /api/v1/shutdown, which requires an internal caller IP, a configured QNET_ADMIN_SECRET
and a matching admin_secret in the body; it flushes and exits 0.
Rolling upgrade (no consensus-visible change). Stop the node, replace the image or binary, start it again with
QNET_HALT_HEIGHT and the one-shot recovery variables (QNET_ROLLBACK_TO_LAST_SEALED, QNET_ROLLBACK_TO_HEIGHT)
unset. It catches up on its own, snapshot-jumping if it fell far enough behind. Do one node at a
time and wait until it reports healthy on /api/v1/node/health with blocks_behind at zero on
/api/v1/sync/status before touching the next. Track validated_peers while you work: taking down more of the
committee than the fault bound tolerates turns a maintenance window into a liveness incident. /healthz answers
with the new image's build= once the replacement has taken. scripts/deploy-genesis.sh runs this pass for the
genesis fleet: it recreates each container from its own docker inspect output, leaving QNET_ROLLBACK_*,
QNET_RECOVERY_HALTED and QNET_BUILD_ID out of the carried environment (QNET_SET_ENV=KEY=VALUE[,KEY=VALUE] adds or
replaces entries). The container it replaces is renamed aside rather than deleted, and its log is gzipped to
/root/qnet-logs/<container>-pre-<timestamp>.log.gz in the background (newest five kept) before it is dropped, so
the hours before a roll can still be read afterwards. It touches the next node only after the
previous one answers /healthz, is fewer than 10 blocks below the network height with at least one validated peer,
has taken part in a checkpoint as a validator and seen a full-quorum seal at or above that window above its restart
height, and has run 30 blocks past its restart with a failover-free metrics window. QNET_RECOVERY_HALTED=1 skips the
last two gates for a roll that repairs a halted chain.
Gated rule change (rolling). A consensus rule that ships behind a feature gate rolls like an ordinary upgrade.
The release carries the rule dormant with an activation height compiled into the binary, every node flips it at that
height, and the operator's whole job is to have the new binary deployed fleet-wide before the height arrives — so
treat the activation height as the deadline for the rolling pass above. The current gates and their heights are in
consensus.md; a release note that names one is telling you
when the deployment must be finished. A crossed gate never moves: only a gate the fleet has not reached may be moved
to a later epoch boundary, before the first roll of its build, and scripts/deploy-genesis.sh refuses to move one
whose old height is not GATE_MARGIN (28,800 blocks) above the fleet tip ("gate already crossed, its height is
frozen").
Protocol-breaking upgrade. A change that alters what any node considers valid and is not carried by a gate cannot
be rolled. Use the halt-height mechanism: set the same QNET_HALT_HEIGHT on every node. The 30-second monitor loop compares the
current height against it, and at or above that height the node flushes storage and exits 0, logging
[INFO][HALT] Reached halt_height=…. Operators then swap binaries and restart with the variable removed. Publish the
halt height, the release commit and the restart wall-clock time ahead of time.
Un-barring identities from a restart manifest is also a coordinated cut-over, never a rolling change. The
exclusion list filters the derived producer and committee set, which feeds consensus_committee and from there the
QC-bound epoch_commitment. If some nodes un-bar a still-heartbeating identity before others, the two groups derive
different committees for the crossover windows and disagree on those checkpoints until the fleet converges. Publish a
wall-clock cut-over, and prefer un-barring only identities already heartbeat-absent, where the removal is a no-op.
This is the recovery procedure for a halted or forged chain: the fleet agrees on the newest macroblock everyone holds, pins it in a new release, and resumes production from it. Rehearse it on a test network before you need it — an unrehearsed runbook is not a recovery plan.
Restart only when both of these hold. Anything less is a bug to fix, not an incident:
- Finality has not advanced for more than two hours. Production stops on its own once the
roster_derivation_horizon(2880 blocks past the last seal) is reached. - There is no software fix that restores liveness without abandoning chain data, and rolling the fleet back to its last sealed macroblock (see Fleet stalled or forked above its last seal below) does not restore it either.
For a forged finality incident — a committee majority certified a bad state_root — the trigger is different:
restart as soon as it is confirmed, and pick K strictly below the first bad macroblock.
- Freeze and gather. Stop producing. From every reachable operator collect the last macroblock index held and its
MacroBlock::hash(), and theconsensus_committeeof the last sealed macroblock. - Choose
K— the newest macroblock that is full-quorum sealed and that every surviving operator agrees on by hash; never one that only some nodes hold. Recordhash(MB_K)and the committee-fields digests forKandK-1; the release needs all four values. - Build the exclusion list — the identities that were in the committee at the stall and did not vote. Assemble it off-chain from operator logs and publish it with the evidence so anyone can disagree before the release ships. List only identities multiple independent operators observed as silent, prefer under-listing (a restart that leaves a few passive identities in still recovers if the survivors clear quorum), and keep the list sorted and deduplicated — the build check fails otherwise.
- Cut the release. In
development/qnet-integration/src/genesis_constants.rsset togetherWS_CHECKPOINT = (K, hash_of_MB_K),WS_CHECKPOINT_DIGEST_ANCHOR,WS_CHECKPOINT_DIGEST_PRED, andRESTART_MANIFEST { resume_from_mb: K, resume_mb_hash: hash_of_MB_K, excluded: &[…] }.restart_manifest_is_wellformed()runs before storage opens and refuses to start the node if the manifest disagrees with the pin index or hash, if either digest is zero, or if the list is unsorted or contains duplicates — a malformed manifest is a broken release, not a runtime condition. - Publish before executing, and confirm a retained-state source. Post in one place:
K,hash(MB_K), both digests, the exclusion list with per-entry evidence, the release commit hash and build instructions, and the wall-clock start time — a restart without a published record is indistinguishable from an attack on the chain. Mandatory precondition: at least one reachable node must retain chain state at or belowKand serve its snapshot, because balances atK > 0survive only as retained chain data. Designate that archival node and confirm it answers aK-height snapshot request before anyone wipes. If none retains state at or belowK, the ledger is unrecoverable; do not proceed. - Execute. Every operator, genesis included, stops their node and wipes chain data above
Konly. Do not wipe a data directory entirely unless the archival node is confirmed serving state at or belowK: a full wipe everywhere destroys the ledger, and the boot guard halts an empty node under the pin rather than minting a fresh genesis. A node that keeps stale data aboveKfails closed on the pin path (v2_ws_pin_mismatch,v2_below_ws) rather than forking. Then run the new release — archival node and genesis nodes first, then the rest. - Verify recovery. Blocks are produced and finality advances (
last_sealed_mb_indexrises on several nodes); theeligible_producersof the first newly sealed macroblock contains none of the barred identities; no node logs[FATAL][RESTART] malformed_manifest; the tip hash matches across at least three operators. - Retire the manifest. Once the chain has been stable for a full epoch, the next release keeps the bumped
WS_CHECKPOINTand clearsRESTART_MANIFESTback to an empty exclusion list only if the barred identities should be allowed to re-register. Leaving them barred is a policy choice — state it publicly either way, and follow the coordinated cut-over rule above.
Run the whole thing on a test network, timed, before you need it in production:
- Halt a test network deliberately by stopping more than one third of the committee.
- Confirm production stops at the expected horizon (2880 blocks past the last seal).
- Choose
Kand verify the hash agrees across operators. - Cut a release with a non-empty exclusion list.
- Confirm a node with stale data above
Kfails closed instead of forking. - Confirm a malformed manifest refuses to start.
- Measure the wall-clock time from decision to restored finality, and publish it.
Node stuck syncing. Check blocks_behind on /api/v1/sync/status and whether it is falling. If flat, check
/api/v1/peers and /api/v1/node/health: isolated means zero peers (firewall or discovery — see
networking.md and the port table in configuration.md);
degraded means fewer than four validated peers, enough to sync but not to participate safely. A node more than 1500
blocks behind should snapshot-jump rather than block-replay; if it does not, check that a peer actually serves one
(GET /api/v1/snapshot/latest against that peer) — snapshot serving is capped at 16 concurrent transfers node-wide
and answers snapshot serve busy beyond that, so a fleet-wide restart can starve joiners temporarily. Restarting the
stuck node is safe and is the usual first move; it resumes from its persisted height.
Node isolated after moving to a new address. Peers bind a registered identity to the IP of the endpoint it
committed on chain, and inbound connections arriving from anywhere else are refused before any signature work. Confirm
QNET_PUBLIC_IP on the new host names the new public address, then reactivate
(POST /api/v1/node-reactivation/submit) — applying that transaction refreshes the committed endpoint, and peers pick
up the new binding as the block applies. Each peer keeps the map on disk and rebuilds it at boot, so a peer that
restarts before the reactivation lands still holds the old address until it does. See
running-a-node.md.
Node not producing. Confirm it should be: GET /api/v1/producer/status reports is_producer for the next block
and names current_producer. Eligibility comes from the eligible_producers snapshot of an earlier macroblock, so a
freshly registered Super is ineligible until ACTIVATION_WARMUP_BLOCKS = 180 blocks have passed, and reputation must
be at or above the consensus minimum of 70 (/api/v1/reputation/history). If the node is elected but no blocks
appear, look at current_timeout_round and failover_count on /api/v1/node/health — the network is rotating away
from it — and at clock_drift_ema_secs and timestamp_rejections. If sealed_lag_windows is climbing past the
horizon the node has parked on roster_derivation_horizon, and the problem is network-wide finality, not this node.
Fleet stalled or forked above its last seal. When every branch agrees up to the last sealed macroblock and only the
unsealed tail differs or stops, roll the fleet back to that point instead of cutting a restart release. Set
QNET_ROLLBACK_TO_LAST_SEALED=1 (or QNET_ROLLBACK_TO_HEIGHT=<h>) on every node, start them together, and remove the
variable before the next start. No rollback goes below what a node holds certified: LAST_SEALED means that certified
floor, and an explicit height below it is refused ([ERR][ROLLBACK] refused … reason=certified_checkpoint_irrevocable,
the node starts unchanged), because a certificate is n−f signatures and any surviving copy pulls the fleet back to it.
At boot, before state recovery, each node truncates to that height, retracts the
macroblocks and certified pairs above it, drops its own vote commitments above it and lowers its anti-double-sign mark
to the height it ends at, so it can sign the re-produced windows. Run with either variable, scripts/deploy-genesis.sh
sets it on the containers it recreates and strips it on its next roll; until then a restart of such a container
repeats the rollback. A recovery decree prunes from one host: on each genesis node node_decreeEndorse (params seq,
target_height) returns that node's consensus signature over the decree, and node_decreeSubmit (the same params plus
sigs) with signatures from a quorum of the genesis consensus keys and a seq above the last applied one gossips it;
every node that verifies it deletes the blocks above target_height, retracts the macroblocks and certified pairs
above it, records the seq and exits for a clean boot. A decree whose target is below a node's certified floor is
refused by that node (and by node_decreeSubmit), which records its seq so it is not gossiped again. Both methods
answer internal callers only. For a fork in which no branch holds a quorum, QNET_ROLLBACK_TO_HEIGHT brings every
stored marker back to one height at boot, never below the node's certified floor, and the node rejoins from the network.
Storage full. Distinguish the two meanings. If the filesystem is full the node cannot flush and should be stopped
before it is starved; free space outside the data directory, then restart. If the node logs storage_warn_85pct_full
or critically_full, that is the internal budget against QNET_MAX_STORAGE_GB, whose cleanups touch caches only. The
real levers: confirm the body prune is running ([INFO][STORAGE] microblock_bodies_pruned appears at 14,400-block
boundaries — a node that has not crossed one since starting has not pruned yet), confirm QNET_DATA_DIR points at the
directory the node really uses, then provision more disk. Never hand-delete files from the RocksDB directory.
Reward-epoch commitment deferring. The checkpoint's reward_epoch_root folds every certified epoch
root up to the N-2 macroblock, resuming from a persisted prefix, and defers rather than seal a shorter
set. Two log lines name what it is waiting on. [WARN][REWARDS] epoch_root_gap … action=defer+repair
means the macroblock is simply absent and the node has already fired a targeted repair fetch — it
resolves itself; watch that the named missing_mb stops recurring. [ERR][REWARDS] epoch_root_mb_no_usable_qc … action=operator_resync means the macroblock is on this node's disk but
unreadable, and it needs you: the object is QC-certified and sits at or below that N-2 macroblock,
so the node keeps it rather than deleting it, and forward sync never revisits that height. Resync this
node from a snapshot; do not hand-delete anything from the RocksDB directory. A third line,
epoch_root_target_unknown … action=defer_no_repair, means a storage read itself failed — treat it as
a disk fault on this node.
Node refuses to start. Read the first fatal line. malformed_manifest is a bad release — rebuild it, do not work
around it. identity_anchor_mismatch is a wrong or lost mnemonic — restore the correct one, never "fix" it by editing
the anchor. WS restart pin active … refusing to mint means the node has a restart pin but no local chain and must
cold-join from the resume macroblock rather than mint a fresh genesis. A repeated exit 137 is the memory ceiling, not a
crash: give the container more memory or reduce what else runs on the host.
Choosing K, deciding who goes on the exclusion list, and judging whether a stall is a bug or an incident are operator
decisions taken off-chain. Say so to the other operators and publish the evidence: getting them wrong bars an honest
operator or abandons real state.