Background
Every deploy_component call stores the uploaded tarball as a Blob on the hdb_deployment record. That blob is replicated to all peers so they can run the local install without a separate upload channel. After replication completes the blob is retained indefinitely — there's no automatic cleanup.
For apps with large payloads (multiple deploys per day, multi-peer clusters) this accumulates fast: each 16 MB payload becomes 16 MB × N peers, and orphaned blobs from superseded deployments are only removed by an explicit delete_deployment_payload call or a full orphan scan. In practice, clusters accumulate hundreds of orphaned payload blobs totaling several GB.
Proposed behavior
1. Threshold-based auto-delete after deploy completes
Once a deploy reaches terminal status (success or failed) and all peers have confirmed receipt, automatically delete the payload_blob from the deployment record if the payload exceeds 10 MB (configurable). The deployment metadata (status, event_log, timestamps, deployment_id, payload_size) is retained; only the blob bytes are dropped.
- Trigger point: after the replicate phase settles (all
onPeerResult callbacks have fired, or the response.replicated fallback path completes)
- Guard: skip deletion if the blob was never stored (small deploys below the threshold keep their blob for audit convenience)
- The existing
delete_deployment_payload operation already implements the deletion; the new code just calls it automatically at the right moment
2. Storage reclamation hook under disk pressure
Add a reclamation callback to the hdb_deployment table that Harper can invoke when available disk space falls below a threshold. The callback should evict payload_blob from the oldest deployment records first (by deploy timestamp), stopping once enough space has been reclaimed or all blobs are gone. This mirrors the pattern used for audit-log pruning.
- Harper already has hooks for storage pressure in the blob layer; the deployment table should register with that mechanism
- Per-node reclamation is sufficient — no cross-cluster coordination needed, since each peer holds its own copy
Implementation notes
- Deploy logic:
components/operations.js — the terminal write happens after emit('phase', { phase: 'replicate', status: 'done' }) at ~line 526; that's the right place to trigger auto-delete
- Blob deletion primitive:
resources/blob.ts deleteBlob() / the blob attribute .delete() path already used by delete_deployment_payload
- Size is already recorded as
payload_size on the row — no need to re-stat the blob to apply the threshold
Acceptance
- Deploys ≥ 10 MB:
payload_blob is null on the deployment record after the deploy settles; payload_size and all metadata remain
- Deploys < 10 MB: blob is retained (small enough that storage cost is negligible, audit value is higher)
- Under disk pressure: reclamation hook evicts oldest payload blobs first; system remains operational
delete_deployment_payload remains available for manual/programmatic cleanup regardless of threshold
Background
Every
deploy_componentcall stores the uploaded tarball as a Blob on thehdb_deploymentrecord. That blob is replicated to all peers so they can run the local install without a separate upload channel. After replication completes the blob is retained indefinitely — there's no automatic cleanup.For apps with large payloads (multiple deploys per day, multi-peer clusters) this accumulates fast: each 16 MB payload becomes 16 MB × N peers, and orphaned blobs from superseded deployments are only removed by an explicit
delete_deployment_payloadcall or a full orphan scan. In practice, clusters accumulate hundreds of orphaned payload blobs totaling several GB.Proposed behavior
1. Threshold-based auto-delete after deploy completes
Once a deploy reaches terminal status (
successorfailed) and all peers have confirmed receipt, automatically delete thepayload_blobfrom the deployment record if the payload exceeds 10 MB (configurable). The deployment metadata (status, event_log, timestamps, deployment_id, payload_size) is retained; only the blob bytes are dropped.onPeerResultcallbacks have fired, or theresponse.replicatedfallback path completes)delete_deployment_payloadoperation already implements the deletion; the new code just calls it automatically at the right moment2. Storage reclamation hook under disk pressure
Add a reclamation callback to the
hdb_deploymenttable that Harper can invoke when available disk space falls below a threshold. The callback should evictpayload_blobfrom the oldest deployment records first (by deploy timestamp), stopping once enough space has been reclaimed or all blobs are gone. This mirrors the pattern used for audit-log pruning.Implementation notes
components/operations.js— the terminal write happens afteremit('phase', { phase: 'replicate', status: 'done' })at ~line 526; that's the right place to trigger auto-deleteresources/blob.tsdeleteBlob()/ the blob attribute.delete()path already used bydelete_deployment_payloadpayload_sizeon the row — no need to re-stat the blob to apply the thresholdAcceptance
payload_blobis null on the deployment record after the deploy settles;payload_sizeand all metadata remaindelete_deployment_payloadremains available for manual/programmatic cleanup regardless of threshold