You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Failed deploy payloads are retained forever; add age/count expiry so retried rollouts stop consuming tenant storage quota #2681
Add a retention policy for the deployment payload tarballs that Harper keeps for failed (and other non-successful) deploys. Today those blobs are retained forever unless an operator manually calls delete_deployment_payload for each deployment_id.
Problem This Solves
Since #1496 (v5.1.15), deployComponent drops payload_blob after a successful deploy when the payload exceeds deployment.payloadRetention.maxSize (default 10 MiB). The drop is deliberately skipped for failed deploys, deploys with failed peers, and staged deploys (components/operations.js, the getPayloadRetentionMaxSize() guard around recorder.dropPayload()), because the tarball is the artifact you would debug or retry with. That is the right call for the most recent failure. It is the wrong outcome once a rollout has been retried a dozen times.
Nothing ever reclaims those retained failure artifacts:
v5.1 and v5.2 (through 5.2.13): only deployment.payloadRetention.maxSize exists, and it only applies on the success path.
main: adds deployment_stagingRetention_maxCount (components/Application.ts, getStagingRetentionMaxCount), which prunes dormant staged builds. Failed deploys are still untouched.
Observed on a multi-tenant instance (harper 5.1.24) with a 10 GB storage quota:
Those 14 payloads totalled 3.6 GB, 36% of the tenant's quota. Twelve were the same ~326 MB bundle failing repeatedly during one rollout attempt; the oldest was ~12 weeks old. The instance was at 95% of quota.
The tenant has no filesystem access, so the only remedy is to discover list_deployments → payload_blob_present and call delete_deployment_payload 14 times. Nothing in the product surfaces that this is needed.
Because failed payloads are the only class of deploy payload that persists, a project whose deploys keep failing (broken restart hook, bad dependency, oversized bundle) accumulates storage at payload_size per attempt with no ceiling. Repeated failures are exactly the situation where an operator retries a lot.
Proposed Solution
Bound retained failure payloads by age and count, keeping the metadata (size, hash, event_log) exactly as dropPayload() / handleDeleteDeploymentPayload already do:
deployment.payloadRetention.failedMaxAge (duration, default e.g. 7d): a terminal failed row older than this has its payload_blob set to null via the same put() path, so the null replicates and peers unlink their copies too. Append event_log entry payload_dropped with reason: 'retention'.
deployment.payloadRetention.failedMaxCount (per project, default e.g. 2): keep only the N most recent failed payloads per project, regardless of age. This is the one that stops a retry loop from accumulating unboundedly within a single afternoon.
Always keep the payload of the most recent failed deploy per project, so the debug/retry value the current guard protects is preserved.
Run the sweep from the same place the staged-build prune runs on main (or on startup plus the existing quota-pressure hook), so there is one deployment-retention pass rather than two.
Never touch non-terminal rows (same 409 guard as delete_deployment_payload), and never touch staged rows (those have their own retention).
Setting either value to 0/unset can preserve today's keep-forever behaviour for operators who want it.
User Stories
As a Fabric tenant, I want failed deploy attempts to stop consuming my storage quota after a reasonable window so that a bad rollout day does not push my instance to its quota ceiling.
As an operator debugging a failing deploy, I want the most recent failed payload kept so that I can get_deployment_payload and reproduce.
As a support engineer, I want a config knob to point customers at instead of a per-deployment_id cleanup procedure.
Alternatives Considered
Status quo, document delete_deployment_payload. Works, but requires the customer to know the bytes are there. list_deployments exposes payload_blob_present but nothing summarises retained bytes or warns near quota.
Drop failed payloads immediately like successful ones. Loses the debug/retry artifact that the current guard intentionally protects. Rejected.
Surface a retained_payload_bytes total in list_deployments / system_information. Useful as a complement, does not fix the accumulation on its own.
Feature Summary
Add a retention policy for the deployment payload tarballs that Harper keeps for failed (and other non-successful) deploys. Today those blobs are retained forever unless an operator manually calls
delete_deployment_payloadfor eachdeployment_id.Problem This Solves
Since #1496 (v5.1.15),
deployComponentdropspayload_blobafter a successful deploy when the payload exceedsdeployment.payloadRetention.maxSize(default 10 MiB). The drop is deliberately skipped for failed deploys, deploys with failed peers, and staged deploys (components/operations.js, thegetPayloadRetentionMaxSize()guard aroundrecorder.dropPayload()), because the tarball is the artifact you would debug or retry with. That is the right call for the most recent failure. It is the wrong outcome once a rollout has been retried a dozen times.Nothing ever reclaims those retained failure artifacts:
v5.1andv5.2(through 5.2.13): onlydeployment.payloadRetention.maxSizeexists, and it only applies on the success path.main: addsdeployment_stagingRetention_maxCount(components/Application.ts,getStagingRetentionMaxCount), which prunes dormant staged builds. Failed deploys are still untouched.Observed on a multi-tenant instance (harper 5.1.24) with a 10 GB storage quota:
hdb_deploymentheld 271 rows. Exactly 14 still hadpayload_blob_present: true, and every one of them wasstatus: failed(12 atphase: restart, 2 atphase: prepare). Zero successful deploys retained bytes, so the feat(deploy): auto-drop large deployment payload blobs after a successful deploy #1496 drop is working as designed.list_deployments→payload_blob_presentand calldelete_deployment_payload14 times. Nothing in the product surfaces that this is needed.Because failed payloads are the only class of deploy payload that persists, a project whose deploys keep failing (broken restart hook, bad dependency, oversized bundle) accumulates storage at
payload_sizeper attempt with no ceiling. Repeated failures are exactly the situation where an operator retries a lot.Proposed Solution
Bound retained failure payloads by age and count, keeping the metadata (size, hash,
event_log) exactly asdropPayload()/handleDeleteDeploymentPayloadalready do:deployment.payloadRetention.failedMaxAge(duration, default e.g.7d): a terminalfailedrow older than this has itspayload_blobset tonullvia the sameput()path, so the null replicates and peers unlink their copies too. Appendevent_logentrypayload_droppedwithreason: 'retention'.deployment.payloadRetention.failedMaxCount(per project, default e.g.2): keep only the N most recent failed payloads per project, regardless of age. This is the one that stops a retry loop from accumulating unboundedly within a single afternoon.main(or on startup plus the existing quota-pressure hook), so there is one deployment-retention pass rather than two.delete_deployment_payload), and never touch staged rows (those have their own retention).Setting either value to
0/unset can preserve today's keep-forever behaviour for operators who want it.User Stories
get_deployment_payloadand reproduce.deployment_idcleanup procedure.Alternatives Considered
delete_deployment_payload. Works, but requires the customer to know the bytes are there.list_deploymentsexposespayload_blob_presentbut nothing summarises retained bytes or warns near quota.retained_payload_bytestotal inlist_deployments/system_information. Useful as a complement, does not fix the accumulation on its own.Priority/Impact
Medium - important but not urgent
Additional Context
handleDeleteDeploymentPayloadincomponents/deploymentOperations.ts.origin/v5.1,origin/v5.2, andorigin/mainon 2026-09-18: no age or count bound on failed payloads in any of them.