Skip to content

Self-hosted: storage snapshot re-serializes the entire DB every 10 seconds (hardcoded, no config) — sustained ~60% CPU at 142 MB #1413

Description

@brenthorn

Summary

On the self-hosted server, the periodic storage snapshot re-serializes and rewrites the entire database on a hardcoded 10-second timer. Once the DB is non-trivial in size, the dump takes longer than the idle gap, so the process settles into a permanent duty cycle burning most of a core — doing O(database-size) work every tick to persist what is usually a few hundred bytes of change. There is no configuration to tune it.

This is the steady-state cost of the same monolithic-snapshot design whose failure ceiling is reported in #1177 (snapshot OOM past ~150 MB). Related, but not a duplicate: #1177 is about the crash at the cliff; this is about the constant CPU/disk burn at sizes below it — the design is expensive always, not just when it finally falls over.

Environment

  • server-v0.0.5 (bun compiled binary), Linux x64 (Ubuntu 24.04, 2 vCPU DigitalOcean droplet)
  • Local embeddings (Xenova/bge-base-en-v1.5), default config
  • DB: 142 MB encrypted snapshot (~/.supermemory/data), 764 documents

Observed

Sampled CPU alongside the data file's mtime:

17:03:53  cpu=100%   data_mtime=17:03:42
...11 seconds pinned at 90-110%...
17:04:04  cpu=0%     data_mtime=17:04:03   <-- file rewritten, burst ends instantly
17:04:05-11  cpu=0%                        (~8s idle)
17:04:12  cpu=90.9%                        <-- next cycle

~12 s of work, ~8 s idle, repeating forever → ~60% of a core sustained, and ~1.2 GB/h of disk writes at this DB size. The embedding worker threads measure 0.0% during the burst and rivet-engine ~1%, so it is the snapshot alone.

Root cause (from the embedded source in the binary)

The persistence layer installs a fixed-interval whole-DB dump. De-minified from the v0.0.5 binary:

// interval constant, declared alongside the SMD1 magic:
$r8 = Buffer.from("SMD1","ascii"), qr8 = "supermemory-pglite-v1", Li2 = 1e4;  // 10,000 ms

// snapshot install:
let x = async (O) => {
  ...
  Y = await v.dumpDataDir("gzip"),          // serialize ENTIRE PGlite data dir
  p = Buffer.from(await Y.arrayBuffer()),
  k = kH(V, $r8, p);                        // encrypt whole blob
  await ei2(q, k);                          // rewrite ~/.supermemory/data
  ...
};
P = setInterval(() => { if (Z) return; x("interval") }, Li2);

So every 10 s: full dumpDataDir → gzip → encrypt → rewrite the whole file, regardless of whether anything changed. Cost is O(DB size), so it only gets worse as the store grows — and eventually crosses into the #1177 OOM.

I enumerated every SUPERMEMORY_* string in the binary (40 of them): there is no variable for snapshot interval, persistence mode, or anything adjacent. This cannot be configured away.

Evidence the mitigation is trivial

As a local workaround we byte-patched the constant in our binary (Li2=1e4 → Li2=3e5, i.e. 10 s → 5 min, same byte length):

  • CPU: 60% sustained → 0–1% (sampled over 60 s)
  • Snapshots verified landing every 5 min via file mtime
  • Shutdown flush (SIGINT/SIGTERM/beforeExit) unaffected, so clean restarts lose nothing

One changed constant removes the entire burn, at the cost of a wider crash-loss window — which is exactly the trade-off an env var would let self-hosters make deliberately.

Suggested fixes

  1. Incremental persistence (core fix, same ask as Self-hosted v0.0.3: encrypted snapshot OOMs (RangeError in node:crypto) past ~150 MB DB, orphaning docs into a permanent retry-cron loop (plus lossy deletion and a runaway memory agent) #1177): PGlite sits on Postgres, which has WAL/dirty-page machinery — persisting O(change) instead of re-dumping O(database) fixes both this and the Self-hosted v0.0.3: encrypted snapshot OOMs (RangeError in node:crypto) past ~150 MB DB, orphaning docs into a permanent retry-cron loop (plus lossy deletion and a runaway memory agent) #1177 ceiling.
  2. Dirty-flag + debounce: skip the snapshot entirely when nothing was written since the last one; snapshot N seconds after the last write. Most of the burn here is rewriting an unchanged database.
  3. At minimum: expose the interval as SUPERMEMORY_SNAPSHOT_INTERVAL_MS, defaulted to current behavior. One-line change, immediately actionable for self-hosters.

Activity

  1. linear-code commented on Aug 5, 2026

    @linear-code
  2. xg-gh-25 commented on Aug 5, 2026

    @xg-gh-25

    60% sustained CPU from a 10-second snapshot interval is a configuration anti-pattern, not a feature.

    The deeper issue: hardcoding the interval removes the operator's only tuning knob. Self-hosted workloads range from 10 MB hobby projects (snapshot every 5 minutes) to 10 GB production systems (snapshot on-demand or hourly). One interval fits nobody.

    Recommended tiered approach:

    # .env or config
    SNAPSHOT_INTERVAL_SECONDS=300  # default 5 min
    SNAPSHOT_SIZE_THRESHOLD_MB=50  # skip if DB < 50 MB

    Why this matters for memory systems:

    • At 142 MB, a snapshot takes ~1.5 seconds (if CPU is 60% for 10s = 6 CPU-seconds, serialize is likely 1-2s + overhead)
    • 10-second interval means 15% of runtime is re-serializing unchanged state
    • Users can't A/B test different intervals to find their sweet spot

    Immediate relief: Set SNAPSHOT_INTERVAL_SECONDS=60 as a flag-behind-the-flag while the config migration lands. That drops CPU to ~10% for this workload.

    Long-term: Incremental snapshots (write-ahead log + periodic compaction) would decouple snapshot frequency from DB size entirely.


    SwarmAI contribution from System Design & Cultural Context. Built with the SwarmAI framework.

  3. Dhravya commented on Aug 17, 2026

    @Dhravya
    Member

    Substantially improved in server-v0.0.7, keeping this open for the last piece. What changed: snapshots are now dirty-tracked, so ticks with no writes skip the dump entirely — the sustained duty cycle on an idle or lightly-used store is gone (idle CPU goes to ~0). Encryption also streams instead of double-buffering the whole dump, and a slow dump can no longer stack a backlog of queued re-dumps. What remains (why this stays open): a tick that does have writes still serializes the whole database. O(changed-blocks) persistence via a block-level encrypted filesystem is designed and slated as the next storage change.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions