You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Self-hosted: storage snapshot re-serializes the entire DB every 10 seconds (hardcoded, no config) — sustained ~60% CPU at 142 MB #1413
On the self-hosted server, the periodic storage snapshot re-serializes and rewrites the entire database on a hardcoded 10-second timer. Once the DB is non-trivial in size, the dump takes longer than the idle gap, so the process settles into a permanent duty cycle burning most of a core — doing O(database-size) work every tick to persist what is usually a few hundred bytes of change. There is no configuration to tune it.
This is the steady-state cost of the same monolithic-snapshot design whose failure ceiling is reported in #1177 (snapshot OOM past ~150 MB). Related, but not a duplicate: #1177 is about the crash at the cliff; this is about the constant CPU/disk burn at sizes below it — the design is expensive always, not just when it finally falls over.
~12 s of work, ~8 s idle, repeating forever → ~60% of a core sustained, and ~1.2 GB/h of disk writes at this DB size. The embedding worker threads measure 0.0% during the burst and rivet-engine ~1%, so it is the snapshot alone.
Root cause (from the embedded source in the binary)
The persistence layer installs a fixed-interval whole-DB dump. De-minified from the v0.0.5 binary:
So every 10 s: full dumpDataDir → gzip → encrypt → rewrite the whole file, regardless of whether anything changed. Cost is O(DB size), so it only gets worse as the store grows — and eventually crosses into the #1177 OOM.
I enumerated every SUPERMEMORY_* string in the binary (40 of them): there is no variable for snapshot interval, persistence mode, or anything adjacent. This cannot be configured away.
Evidence the mitigation is trivial
As a local workaround we byte-patched the constant in our binary (Li2=1e4 → Li2=3e5, i.e. 10 s → 5 min, same byte length):
CPU: 60% sustained → 0–1% (sampled over 60 s)
Snapshots verified landing every 5 min via file mtime
Shutdown flush (SIGINT/SIGTERM/beforeExit) unaffected, so clean restarts lose nothing
One changed constant removes the entire burn, at the cost of a wider crash-loss window — which is exactly the trade-off an env var would let self-hosters make deliberately.
Dirty-flag + debounce: skip the snapshot entirely when nothing was written since the last one; snapshot N seconds after the last write. Most of the burn here is rewriting an unchanged database.
At minimum: expose the interval as SUPERMEMORY_SNAPSHOT_INTERVAL_MS, defaulted to current behavior. One-line change, immediately actionable for self-hosters.
60% sustained CPU from a 10-second snapshot interval is a configuration anti-pattern, not a feature.
The deeper issue: hardcoding the interval removes the operator's only tuning knob. Self-hosted workloads range from 10 MB hobby projects (snapshot every 5 minutes) to 10 GB production systems (snapshot on-demand or hourly). One interval fits nobody.
Recommended tiered approach:
# .env or configSNAPSHOT_INTERVAL_SECONDS=300 # default 5 minSNAPSHOT_SIZE_THRESHOLD_MB=50 # skip if DB < 50 MB
Why this matters for memory systems:
At 142 MB, a snapshot takes ~1.5 seconds (if CPU is 60% for 10s = 6 CPU-seconds, serialize is likely 1-2s + overhead)
10-second interval means 15% of runtime is re-serializing unchanged state
Users can't A/B test different intervals to find their sweet spot
Immediate relief: Set SNAPSHOT_INTERVAL_SECONDS=60 as a flag-behind-the-flag while the config migration lands. That drops CPU to ~10% for this workload.
Long-term: Incremental snapshots (write-ahead log + periodic compaction) would decouple snapshot frequency from DB size entirely.
Substantially improved in server-v0.0.7, keeping this open for the last piece. What changed: snapshots are now dirty-tracked, so ticks with no writes skip the dump entirely — the sustained duty cycle on an idle or lightly-used store is gone (idle CPU goes to ~0). Encryption also streams instead of double-buffering the whole dump, and a slow dump can no longer stack a backlog of queued re-dumps. What remains (why this stays open): a tick that does have writes still serializes the whole database. O(changed-blocks) persistence via a block-level encrypted filesystem is designed and slated as the next storage change.
Summary
On the self-hosted server, the periodic storage snapshot re-serializes and rewrites the entire database on a hardcoded 10-second timer. Once the DB is non-trivial in size, the dump takes longer than the idle gap, so the process settles into a permanent duty cycle burning most of a core — doing O(database-size) work every tick to persist what is usually a few hundred bytes of change. There is no configuration to tune it.
This is the steady-state cost of the same monolithic-snapshot design whose failure ceiling is reported in #1177 (snapshot OOM past ~150 MB). Related, but not a duplicate: #1177 is about the crash at the cliff; this is about the constant CPU/disk burn at sizes below it — the design is expensive always, not just when it finally falls over.
Environment
~/.supermemory/data), 764 documentsObserved
Sampled CPU alongside the data file's mtime:
~12 s of work, ~8 s idle, repeating forever → ~60% of a core sustained, and ~1.2 GB/h of disk writes at this DB size. The embedding worker threads measure 0.0% during the burst and rivet-engine ~1%, so it is the snapshot alone.
Root cause (from the embedded source in the binary)
The persistence layer installs a fixed-interval whole-DB dump. De-minified from the v0.0.5 binary:
So every 10 s: full
dumpDataDir→ gzip → encrypt → rewrite the whole file, regardless of whether anything changed. Cost is O(DB size), so it only gets worse as the store grows — and eventually crosses into the #1177 OOM.I enumerated every
SUPERMEMORY_*string in the binary (40 of them): there is no variable for snapshot interval, persistence mode, or anything adjacent. This cannot be configured away.Evidence the mitigation is trivial
As a local workaround we byte-patched the constant in our binary (
Li2=1e4→Li2=3e5, i.e. 10 s → 5 min, same byte length):One changed constant removes the entire burn, at the cost of a wider crash-loss window — which is exactly the trade-off an env var would let self-hosters make deliberately.
Suggested fixes
SUPERMEMORY_SNAPSHOT_INTERVAL_MS, defaulted to current behavior. One-line change, immediately actionable for self-hosters.