Summary
During a full-copy resync with source-unavailable blobs, recordBlobReplicationFailure() fires (observed 14× in the repro) and the receiver logs advancing the resume cursor past it (7×), but cluster_status on both nodes reports blobReplicationFailures: 0. The log message itself cites cluster_status.blobReplicationFailures as the authoritative cumulative total — so an operator monitoring that field would never see blob-replication failures that ARE happening.
Likely cause
The auditStore lookup in clusterStatus.ts:47-52 finds nothing at query time (or reads a freshly-zeroed shared-buffer slot), so the metric-augmentation block is skipped. (harper-pro 282a0bc)
Severity
Low-medium — operator-facing observability gap, not data integrity. (The underlying advance-past behavior is separately tracked; this issue is specifically that the metric stays 0.)
Repro
integrationTests/cluster/qa-scratch/qa339-blob-resync-wedge.test.mjs (from QA-339).
— from Harper exploratory QA (KrAIs)
Summary
During a full-copy resync with source-unavailable blobs,
recordBlobReplicationFailure()fires (observed 14× in the repro) and the receiver logsadvancing the resume cursor past it(7×), butcluster_statuson both nodes reportsblobReplicationFailures: 0. The log message itself citescluster_status.blobReplicationFailuresas the authoritative cumulative total — so an operator monitoring that field would never see blob-replication failures that ARE happening.Likely cause
The
auditStorelookup inclusterStatus.ts:47-52finds nothing at query time (or reads a freshly-zeroed shared-buffer slot), so the metric-augmentation block is skipped. (harper-pro 282a0bc)Severity
Low-medium — operator-facing observability gap, not data integrity. (The underlying advance-past behavior is separately tracked; this issue is specifically that the metric stays 0.)
Repro
integrationTests/cluster/qa-scratch/qa339-blob-resync-wedge.test.mjs(from QA-339).— from Harper exploratory QA (KrAIs)