Skip to content

Latest commit

 

History

History
347 lines (250 loc) · 18 KB

File metadata and controls

347 lines (250 loc) · 18 KB

XChain Platform Decoder: Operations

Prerequisites

  • Node.js 22 (22.x LTS), the platform's canonical runtime, pinned in .nvmrc
  • MariaDB server
  • A running coin node (bitcoind, litecoind, or dogecoind) with JSON-RPC enabled

Running the Decoder

npm run api
# or directly:
node ./src/api.js

On startup, the decoder:

  1. Loads environment variables from .env
  2. Starts the Express JSON-RPC API server on DECODER_API_PORT
  3. Validates the database name (alphanumeric + underscores only)
  4. Creates the database if it doesn't exist
  5. Creates all 9 tables if they don't exist
  6. Waits for the coin node to reach 99% verification progress
  7. Begins parsing from the configured start block (or last parsed block + 1)

Docker

The decoder includes a docker-compose.yml for containerized deployment:

docker-compose up --build

The Dockerfile copies the source into the container and runs npm run api. Environment variables can be passed via docker-compose.yml or a mounted .env file.

Stopping

The decoder handles graceful shutdown via SIGTERM and SIGINT signals:

  1. Sets the stopFlag to true
  2. The main polling loop exits after the current iteration completes
  3. Mempool updates are cancelled
  4. All database connections are released

In Docker, docker stop sends SIGTERM, triggering the graceful shutdown path.

API

The decoder exposes a minimal JSON-RPC API for health monitoring:

ping

Basic health check.

Request:

{
    "jsonrpc": "2.0",
    "method": "ping",
    "id": 1
}

Response:

{
    "jsonrpc": "2.0",
    "result": { "status": "success" },
    "id": 1
}

health

Detailed health status including decoder state.

Request:

{
    "jsonrpc": "2.0",
    "method": "health",
    "id": 1
}

Response:

{
    "jsonrpc": "2.0",
    "result": {
        "status": "healthy",
        "phase": "running",
        "synced": true,
        "last_processed_block": 900123,
        "node_height": 900124,
        "lag": 1,
        "lastProcessedBlock": 900123,
        "chainTipBlock": 900124,
        "blockLag": 1,
        "lag_blocks": 1,
        "rpc_errors": 0,
        "parse_errors": 0,
        "error": null
    },
    "id": 1
}
Field Type Description
status string "healthy" when the decoder is running and MariaDB is reachable; "unhealthy" otherwise
phase string "running" when the DB probe succeeds; "starting" while the decoder is still connecting to MariaDB; "db-unreachable" when the probe fails
synced boolean Whether the decoder is within 3 blocks of the chain tip
last_processed_block integer|null Last block index written to the decoder DB (from getSyncStatus())
node_height integer|null Current tip reported by the coin node (from getSyncStatus())
lag integer|null Blocks behind tip from getSyncStatus()
lastProcessedBlock integer|null Alias for last_processed_block (convenience copy)
chainTipBlock integer|null Alias for node_height (convenience copy)
blockLag integer|null Alias for lag (convenience copy)
lag_blocks integer|null Live lag computed from internal decoder state; null when either height is still unknown (before the first getBlockchainInfo, or nothing processed yet) rather than a misleading 0. May differ slightly from lag during rapid catch-up
node_height_stale boolean Present and true when the last successful node-tip poll is more than two refresh intervals old (node outage): node_height is then frozen, so a zero lag does not mean caught-up
reorg_halted boolean true while the database carries a live durable REORG_HALT marker (see Decoder halted after a deep reorg). The decoder keeps parsing and status stays "healthy"; the next reorg parks it (see reorg_halt_parked)
reorg_halt_reason string|null Why the halt was written
reorg_halted_at string|null When the halt was written (ISO 8601)
reorg_halt_cleared_at string|null When an operator cleared the last halt with clear-reorg-halt; null while a halt is live or none was recorded
reorg_halt_cleared_reason string|null The reason the operator recorded with that clear
reorg_halt_checked_at integer|null Epoch ms of the last marker probe (cached for one minute)
reorg_halt_parked boolean true once the parse loop has stopped on the halt and is waiting for the clear. false on a decoder that carries the marker but is still parsing forward, which is the only thing that separates the two
reorg_halt_parked_at string|null When the parse loop parked (ISO 8601); null when it is not parked
rpc_errors integer Combined RPC error count from the decoder and its BlockchainConnector
parse_errors integer Number of transactions quarantined due to parse failures
error string|null Error message if the decoder crashed, otherwise null

GET /status (REST)

Returns HTTP 200 with {status: "healthy", db, running} when the decoder is running and MariaDB is reachable, or HTTP 503 otherwise. Distinct from the JSON-RPC health method so load balancers and uptime monitors can rely on the HTTP status code directly (a plain GET against the JSON-RPC root always answers 200). Point load balancers and uptime monitors here; point Docker HEALTHCHECKs at GET /live.

GET /live (REST)

The liveness probe, and the route the Docker HEALTHCHECK runs. Everything /status reports, plus stalled, poll_silent, last_poll_at, last_processed_block, node_height, lag, parse_errors, and rpc_errors. Returns HTTP 503 when the decoder is not running, MariaDB is unreachable, the block loop is stalled, or the block loop has stopped iterating.

Stalled means the loop is alive and retrying but no longer making progress the chain is waiting on: either one height has failed to fetch on many consecutive attempts, or nothing advanced for DECODER_STALL_ALERT_MS (default 900000 ms, 15 minutes) while the node tip was fresh and visibly ahead. The block loop never skips a block on a fetch or parse fault, because skipping would corrupt the index, so a deterministic fault at one height retries forever with the process alive and the DB answering. That case is invisible to /status, which is why autoheal probes /live instead.

A frozen node tip (node outage) is deliberately not stalled: both sides stop and a restart fixes nothing. Tune the window per host with DECODER_STALL_ALERT_MS, and the consecutive-fetch-failure threshold with DECODER_STALL_FETCH_ATTEMPTS (default 20 attempts, about one minute).

Stopped iterating is a third failure, and the only one none of the fields above can see. stalled asks whether the CHAIN is advancing, so a caught-up decoder is never stalled by construction; a loop that dies or hangs while caught up therefore left running and db true and stalled false, and this probe answered 200 indefinitely while nothing parsed. last_poll_at is stamped at the top of every loop iteration whether or not a block arrived, and poll_silent goes true once it is older than DECODER_POLL_SILENT_MS (default twice the stall window, 30 minutes at stock settings). Unlike a frozen node tip this one gates health, because a restart does fix it. The outage retry path re-enters the loop top every three seconds, so a node outage keeps the heartbeat ticking and stays 200 here.

getlatestblock

Returns the decoder's latest parsed block alongside the coin-node's tip, useful for monitoring decoder-to-node lag in a single call.

Request:

{
    "jsonrpc": "2.0",
    "method": "getlatestblock",
    "id": 1
}

Response:

{
    "jsonrpc": "2.0",
    "result": {
        "block_index": 900123,
        "node_block_index": 900124,
        "is_synced": true
    },
    "id": 1
}
Field Type Description
block_index integer|null Last block index written to the decoder DB
node_block_index integer|null Current tip reported by the coin node
is_synced boolean Whether the decoder is within 3 blocks of the chain tip

Security

The API includes:

  • Helmet: sets secure HTTP headers
  • CORS: enabled for cross-origin requests
  • Rate limiting: 100 requests per minute per IP
  • Body size limit: 100kb maximum request body

Schema Migrations

The decoder ships with a migration system that tracks and applies schema changes to existing databases. Two migration modes exist:

  • Auto migrations (tagged mode=auto) are applied automatically at every startup. These are additive and idempotent (guarded with IF NOT EXISTS).
  • Manual migrations (tagged mode=manual) require an explicit operator run. These cover destructive or data-backfill changes that must not run unattended.

To apply pending manual migrations:

node src/db/migrate.js
# or: npm run migrate

The run holds a database-scoped advisory lock so concurrent processes cannot apply the same migration twice. Each applied migration is recorded in the schema_migrations table with its filename, SHA-256 checksum, mode, and timestamp. Re-running the command is safe; only pending migrations are applied.

Take a backup and stop the decoder before running manual migrations, since some involve full table rebuilds.

Reorg Handling

Chain reorganizations are detected automatically during the block polling loop:

  1. Before writing a new block, the decoder compares the previous_block_hash from the coin node with the hash stored in the database for the previous block
  2. If they don't match, a reorganization has occurred
  3. The decoder deletes the invalid block (and all its transactions) from the database
  4. A REORG event is recorded in the events table with the affected block height
  5. The polling loop resumes, re-parsing from the corrected chain

The indexer monitors the decoder's blocks table and independently handles reorg rollback of its own state.

Two cases are deliberately not treated as a reorg:

  • The node is still catching up. While getblockchaininfo reports initialblockdownload: true and the node's tip sits below the decoder's stored tip, the decoder waits (one warning, then a poll every 5 seconds) instead of rolling back. A node that has not yet validated blocks the decoder already holds is not a rolled-back node; once it passes the stored tip, the normal hash compare decides.
  • The gap is deeper than the safe-depth window. When the node's tip sits more than DISPENSER_EXPIRE_SAFE_DEPTH (126) blocks below the stored tip, the decoder refuses before deleting anything and keeps polling. Nothing is rolled back and no halt marker is written, so no resync is owed; either the node catches up on its own or the operator acts on a real rollback.

Deep reorg halt (REORG_HALT)

Soft-expired dispensers are hard-purged once they are DISPENSER_EXPIRE_SAFE_DEPTH blocks deep. A rollback that crosses that window could not resurrect them, so the decoder stops the rollback at the ceiling and writes a durable marker: a row in the events table with code = 'REORG_HALT'. The decoder keeps parsing forward and reports reorg_halted: true on health, GET /status and GET /live; the next reorg refuses to roll back and stops the process. The marker survives restarts and is only ever superseded, never deleted. See Troubleshooting for recovery.

Mempool Tracking

Mempool tracking activates when the decoder is synced (within 3 blocks of the tip):

  • Poll interval: every 60 seconds
  • Batch size: 1000 transactions per RPC batch
  • Comparison method: binary search against sorted txid lists
  • Cleanup: stale mempool entries (no longer in node's mempool) are deleted each cycle

Mempool tracking pauses if the decoder falls more than 3 blocks behind the tip, and resumes automatically when caught up.

Connection Resilience

Coin Node (JSON-RPC)

  • All RPC calls have a 30-second HTTP timeout by default (overridable via NODE_RPC_TIMEOUT)
  • Failed calls are retried up to 10 times with 500ms backoff
  • HTTP 429 (rate limited) triggers a longer 5-second backoff
  • Connection aborts (ECONNABORTED) are retried with timeout warnings logged

MariaDB

  • Connection pool of 10 concurrent connections
  • 30-second query execution timeout (configurable via DB_QUERY_TIMEOUT)
  • Transaction locking via a queue-based mutex prevents concurrent block commits and mempool updates from interleaving
  • On connection failure, errors are logged with e.code (not full error object, to avoid credential leakage)

Troubleshooting

Decoder won't start parsing

  • Verify the coin node is accessible at NODE_URL:NODE_PORT
  • Check that verificationprogress is >= 0.99 (run bitcoin-cli getblockchaininfo)
  • If using Dogecoin, ensure AUX_POW is set in the environment

Decoder is stuck / not advancing

  • Check coin node connectivity (the decoder logs RPC timeout warnings)
  • Verify MariaDB is accessible and the connection pool isn't exhausted
  • Check for reorg loops, if the chain is continuously reorganizing, the decoder may repeatedly delete and re-parse the same block

Decoder halted after a deep reorg (REORG_HALT)

Symptoms: the log printed LATENT REORG_HALT MARKER PRESENT once at startup, health reports reorg_halted: true, or xchain-node ps shows running REORG_HALT on the decoder. The decoder still parses forward; the next reorg parks it.

A halted decoder stays up. When a reorg reaches the marker the parse loop parks instead of exiting: the container keeps running at a stable restart count, the API keeps answering, and health, /status and /live report reorg_halt_parked: true beside reorg_halted: true, which is what tells a parked decoder apart from one still parsing forward on a dormant marker. xchain-node ps shows REORG_HALT against that steady count rather than a climbing one. The parked loop re-reads the marker every 60 seconds, so once the audited clear below lands, the decoder logs one line and resumes parsing from its stored tip on its own, with no restart.

Confirm it (this is the decoder's own marker; the sync_halt table in the same database belongs to xchain-sync and is a different signal):

SELECT id, time, code, data FROM events
 WHERE code IN ('REORG_HALT', 'REORG_HALT_CLEARED')
 ORDER BY id DESC LIMIT 5;

The newest row decides: a REORG_HALT newer than every REORG_HALT_CLEARED is live.

Two recoveries:

  1. Full resync from a known-good snapshot. Always correct. Required when the database has held dispenser state that a purge may have dropped.

  2. Audited clear, when the database is known to be intact:

    xchain-node clear-reorg-halt <chain> <network> --reason "<why this database is known good>"
    # standalone, inside the decoder's environment:
    npm run clear-reorg-halt -- --reason "<why this database is known good>"

    The clear checks that every rolled-back block above the tip has been re-parsed (cannot be forced) and that the database holds no dispenser rows and never decoded a DISPENSER action (so the purge could not have lost anything). A database that has held dispensers is refused unless you pass --force after comparing its dispensers table against a known-good replica; the clear is then recorded as forced. --dry-run reports the verdict without writing and needs no --reason, so run it first to see what a clear would do.

    When the re-parse check refuses, reorg_halt_parked decides what to do next. A decoder still parsing forward on a dormant marker (reorg_halt_parked: false) re-parses the range on its own: wait for it to pass the halt height and run the clear again. A parked decoder (reorg_halt_parked: true) parses nothing and never re-parses the range, so the clear cannot pass and recovery 1, a full resync, is the path.

    The clear writes a REORG_HALT_CLEARED event carrying the reason, the check results and the halt it supersedes. The halt row stays for the audit trail, health reports reorg_halted: false with reorg_halt_cleared_at set on its next probe, and the bootstrap health gate accepts the database again.

Never delete the REORG_HALT row by hand: that erases the evidence the clear records and leaves nothing for the next operator to read.

One marker is rarely alone: a deep reorg on one chain often coincides with halts on the other decoders of the same box. After finding one, run xchain-node ps and check every decoder it lists for REORG_HALT before moving on.

Database name rejected

  • Database names must match /^[A-Za-z0-9_]+$/. No spaces, backticks, or special characters
  • Follow the naming convention: XChain_{CHAIN}_{NETWORK}_Decoder

Mempool not updating

  • Mempool tracking only runs when the decoder is synced (within 3 blocks of tip)
  • Check that the 60-second interval hasn't been interrupted by a long block parse
  • Verify getRawMempool RPC is accessible

High memory usage

  • Large mempool batches (>10,000 unconfirmed txs) can cause temporary memory spikes during the batch-fetch phase
  • The 1000-tx chunk size limits peak memory per batch

Monitoring

The health API endpoint provides the key monitoring signals:

Condition status synced
Normal operation, caught up "healthy" true
Normal operation, catching up "healthy" false
Decoder crashed "unhealthy" false

Monitor the events table for REORG events, which indicate chain instability.


Copyright © 2025–2026 Dankest, LLC

Based on XChain Platform by Dankest, LLC – https://dankest.llc

Licensed under the GNU Affero General Public License v3.0 (AGPL-3.0-or-later) with a commercial license available for proprietary use.

You may use, modify, and distribute this material under the terms of the License. See LICENSE and NOTICE for full terms. See the licensing overview.