During sustained bulk ingest (snapshot backfill at ~100 MB/s), the write-ahead log can grow faster than it flushes. On a Kubernetes deployment with an S3/object-storage backend, the local PVC holds only the WAL and cache, and the chart's default persistence.size: 10Gi (helm/arc/values.yaml) fills in minutes at that rate. The resulting failure chain is worse than a full disk:
- WAL fills the volume with no backpressure. Writes keep being accepted while the WAL grows unboundedly until the volume is full.
- Boot deadlock. After the writer goes down, it cannot start again: creating the initial WAL file on startup requires free space on the already-full volume, so the pod can never come back without manual intervention (growing the PVC or deleting WAL segments by hand).
- Liveness kills recovery. Once space is available, WAL replay after an unclean stop can exceed the liveness budget in
helm/arc/templates/deployment.yaml (periodSeconds: 10, failureThreshold: 3, no startupProbe), so the pod is killed mid-recovery and crash-loops repeatedly before it manages to drain.
Observed in the field during a TB-scale migration; the operator's workaround was ARC_WAL_ENABLED=false for the migration window (throughput was unchanged, WAL stayed flat), which works but gives up durability for live writes during that window.
Proposed fixes
Related (probably a separate issue)
With very wide rows (tens of KB of line protocol per row), the ingest buffer's row-count bound (100k rows) and the WAL's payload byte cap (WAL payload exceeds maximum allowed size, 104,857,600 bytes) cannot be reconciled by tuning: a full flush of wide rows exceeds the cap, so those flushes never reach the WAL at all — the WAL directory holds only a 7-byte header file while ingest runs. Chunking WAL record writes, or bounding the ingest buffer by bytes as well as rows, would fix it. Happy to split this into its own issue.
During sustained bulk ingest (snapshot backfill at ~100 MB/s), the write-ahead log can grow faster than it flushes. On a Kubernetes deployment with an S3/object-storage backend, the local PVC holds only the WAL and cache, and the chart's default
persistence.size: 10Gi(helm/arc/values.yaml) fills in minutes at that rate. The resulting failure chain is worse than a full disk:helm/arc/templates/deployment.yaml(periodSeconds: 10,failureThreshold: 3, nostartupProbe), so the pod is killed mid-recovery and crash-loops repeatedly before it manages to drain.Observed in the field during a TB-scale migration; the operator's workaround was
ARC_WAL_ENABLED=falsefor the migration window (throughput was unchanged, WAL stayed flat), which works but gives up durability for live writes during that window.Proposed fixes
startupProbe(chart): add astartupProbewith a generousfailureThreshold × periodSecondswindow (minutes, not seconds) so WAL recovery is never killed by the liveness probe; keep the current livenessProbe for steady state.10Giis a dev-scale default, and consider raising it or adding a loud comment invalues.yaml.Related (probably a separate issue)
With very wide rows (tens of KB of line protocol per row), the ingest buffer's row-count bound (100k rows) and the WAL's payload byte cap (
WAL payload exceeds maximum allowed size, 104,857,600 bytes) cannot be reconciled by tuning: a full flush of wide rows exceeds the cap, so those flushes never reach the WAL at all — the WAL directory holds only a 7-byte header file while ingest runs. Chunking WAL record writes, or bounding the ingest buffer by bytes as well as rows, would fix it. Happy to split this into its own issue.