Skip to content

Auto-delete large deployment payload blobs after deploy; storage reclamation for older payloads #1495

Description

@kriszyp

Background

Every deploy_component call stores the uploaded tarball as a Blob on the hdb_deployment record. That blob is replicated to all peers so they can run the local install without a separate upload channel. After replication completes the blob is retained indefinitely — there's no automatic cleanup.

For apps with large payloads (multiple deploys per day, multi-peer clusters) this accumulates fast: each 16 MB payload becomes 16 MB × N peers, and orphaned blobs from superseded deployments are only removed by an explicit delete_deployment_payload call or a full orphan scan. In practice, clusters accumulate hundreds of orphaned payload blobs totaling several GB.

Proposed behavior

1. Threshold-based auto-delete after deploy completes

Once a deploy reaches terminal status (success or failed) and all peers have confirmed receipt, automatically delete the payload_blob from the deployment record if the payload exceeds 10 MB (configurable). The deployment metadata (status, event_log, timestamps, deployment_id, payload_size) is retained; only the blob bytes are dropped.

  • Trigger point: after the replicate phase settles (all onPeerResult callbacks have fired, or the response.replicated fallback path completes)
  • Guard: skip deletion if the blob was never stored (small deploys below the threshold keep their blob for audit convenience)
  • The existing delete_deployment_payload operation already implements the deletion; the new code just calls it automatically at the right moment

2. Storage reclamation hook under disk pressure

Add a reclamation callback to the hdb_deployment table that Harper can invoke when available disk space falls below a threshold. The callback should evict payload_blob from the oldest deployment records first (by deploy timestamp), stopping once enough space has been reclaimed or all blobs are gone. This mirrors the pattern used for audit-log pruning.

  • Harper already has hooks for storage pressure in the blob layer; the deployment table should register with that mechanism
  • Per-node reclamation is sufficient — no cross-cluster coordination needed, since each peer holds its own copy

Implementation notes

  • Deploy logic: components/operations.js — the terminal write happens after emit('phase', { phase: 'replicate', status: 'done' }) at ~line 526; that's the right place to trigger auto-delete
  • Blob deletion primitive: resources/blob.ts deleteBlob() / the blob attribute .delete() path already used by delete_deployment_payload
  • Size is already recorded as payload_size on the row — no need to re-stat the blob to apply the threshold

Acceptance

  • Deploys ≥ 10 MB: payload_blob is null on the deployment record after the deploy settles; payload_size and all metadata remain
  • Deploys < 10 MB: blob is retained (small enough that storage cost is negligible, audit value is higher)
  • Under disk pressure: reclamation hook evicts oldest payload blobs first; system remains operational
  • delete_deployment_payload remains available for manual/programmatic cleanup regardless of threshold

Activity

  1. self-assigned this
    on Jun 25, 2026
  2. kriszyp commented on Jul 6, 2026

    @kriszyp
    MemberAuthor

    Follow-up from exploratory QA (harper 228eacc0f): Part 1 (success-path payload-blob reclaim) works — 9× 3MB deploys held blob-store growth flat (0.00 MB, payload_blob_present→false, file removed). But the failure path leaks the payload blob permanently (the "Part 2" gap):

    A failed deploy_component's hdb_deployment row stays payload_blob_present: true, payload_size: <N> forever. deployComponent's success branch (components/operations.js:591-596) is the only call to dropPayload(); the catch block (~601+) never calls it. There's no TTL/cap on hdb_deployment rows, and cleanupOrphans only sweeps unreferenced blobs — this one stays referenced. So N repeated failed deploys (e.g. a broken CI loop) leak N × payload_size of multi-MB blobs permanently.

    Fix: call dropPayload() in the catch branch (or add a TTL/cap on retained failure payloads).

    Repro: local qa-scratch/qa474-deploy-blob-reclaim.test.ts.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Fields

Priority

None yet

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions