Summary
On Fabric instances, the npm download cache (<rootPath>/.npm/_cacache) grows unbounded across deploys and is counted against the instance storage quota. For projects with large/frequent deploys it can consume nearly the entire quota, fill it, and write-wedge the instance — even when the actual database is tiny.
Live incident (where this was found)
wild cluster, origin node zne-us-west2-a-1 (free 5 GB instance, harper-pro 5.1.6), Bailey's personal project (Phaser open-world game, large bundled client + generated content, frequent deploys).
Quota usage breakdown at time of incident:
| Path |
Size |
.npm/_cacache |
4.3 GB |
blobs/ |
213 MB |
database/ (actual data) |
205 MB |
| total |
4.7 GB / 5.0 GB quota → full |
With the quota full, every system-db transaction-log write aborted:
[main/0] [error]: unhandledRejection [Error: Operation aborted: Failed to write
transaction log entries to file: .../database/system/transaction_logs/local/3.txnlog] { code: 'ERR_ABORTED' }
…looping continuously on the main thread. deploy_component then failed at prepare with -122 (EDQUOT) — the instance was up and serving reads but could not stage any payload.
Remediation applied: removed .npm/_cacache (4.7 GB → 459 MB) and restarted the container. Quota recovered, abort loop cleared, deploys unblocked.
Secondary fallout: corrupt txnlog tail on restart
The aborted (partially-written) txnlog writes left torn entries at the tail of the local replication logs. On restart, replay stopped at the corrupt boundary (handled gracefully, [warn] not crash):
[warn]: Stopping transaction log "local" at a corrupt entry during replay
RangeError: Corrupt transaction log entry at position cc06d of log 6:
declared length 1688494450 overruns the log (limit=842946)
So a quota-full condition can produce corrupt local txnlog tails, not just blocked writes — worth confirming the truncate-on-replay path is always safe.
Root cause
core/utility/npmUtilities.ts (and the auto-install-on-deploy path) invoke npm install with ['install', '--force', '--omit=dev', '--json'] and no --cache override. npm therefore defaults its cache to $HOME/.npm/_cacache, and $HOME inside the Harper container is the rootPath (/home/harperdb/harper) — which is exactly the directory the Fabric quota meters. --force aggravates it by forcing re-download/re-store. Nothing ever prunes the cache.
Suggested fixes (any one helps; first two are the real fix)
- Point the npm cache outside the quota-counted rootPath (e.g.
--cache to an ephemeral/container-scratch dir, or set npm_config_cache).
- Prune the cache after each deploy (
npm cache clean / cap size), so it can't grow unbounded.
- Consider dropping
--force unless there's a specific reason — it defeats cache reuse and inflates writes.
- Defensive: when a txnlog write aborts on EDQUOT, surface a clear "storage quota exceeded" error rather than an unhandledRejection loop, and ensure replay-truncation of torn tail entries is always safe.
Impact / why it matters now
This is independent of the in-flight replication-wedge fix (#443 / #1425). Any customer doing large or frequent deploys will eventually hit EDQUOT this way regardless of actual data size — directly relevant to the Macy's 5.1.6 promotion question, since they deploy large projects.
Summary
On Fabric instances, the npm download cache (
<rootPath>/.npm/_cacache) grows unbounded across deploys and is counted against the instance storage quota. For projects with large/frequent deploys it can consume nearly the entire quota, fill it, and write-wedge the instance — even when the actual database is tiny.Live incident (where this was found)
wildcluster, origin nodezne-us-west2-a-1(free 5 GB instance, harper-pro 5.1.6), Bailey's personal project (Phaser open-world game, large bundled client + generated content, frequent deploys).Quota usage breakdown at time of incident:
.npm/_cacacheblobs/database/(actual data)With the quota full, every
system-db transaction-log write aborted:…looping continuously on the main thread.
deploy_componentthen failed atpreparewith-122(EDQUOT) — the instance was up and serving reads but could not stage any payload.Remediation applied: removed
.npm/_cacache(4.7 GB → 459 MB) and restarted the container. Quota recovered, abort loop cleared, deploys unblocked.Secondary fallout: corrupt txnlog tail on restart
The aborted (partially-written) txnlog writes left torn entries at the tail of the local replication logs. On restart, replay stopped at the corrupt boundary (handled gracefully,
[warn]not crash):So a quota-full condition can produce corrupt local txnlog tails, not just blocked writes — worth confirming the truncate-on-replay path is always safe.
Root cause
core/utility/npmUtilities.ts(and the auto-install-on-deploy path) invokenpm installwith['install', '--force', '--omit=dev', '--json']and no--cacheoverride. npm therefore defaults its cache to$HOME/.npm/_cacache, and$HOMEinside the Harper container is the rootPath (/home/harperdb/harper) — which is exactly the directory the Fabric quota meters.--forceaggravates it by forcing re-download/re-store. Nothing ever prunes the cache.Suggested fixes (any one helps; first two are the real fix)
--cacheto an ephemeral/container-scratch dir, or setnpm_config_cache).npm cache clean/ cap size), so it can't grow unbounded.--forceunless there's a specific reason — it defeats cache reuse and inflates writes.Impact / why it matters now
This is independent of the in-flight replication-wedge fix (#443 / #1425). Any customer doing large or frequent deploys will eventually hit EDQUOT this way regardless of actual data size — directly relevant to the Macy's 5.1.6 promotion question, since they deploy large projects.