Skip to content

Reduce small-write process and storage overhead with measured evidence #974

Description

@flyingrobots

id: "GW20-018"
title: "Reduce small-write process and storage overhead with measured evidence"
type: "bug"
status: "planned"
milestone: "v20.1.0"
category: "required"
workstream: "runtime"
issue: "#974"
baseline_commit: "ceb58e656ec5bc85f0ae5991bf2a55599693bbb3"
feedback_sections: [13]
prerequisites: []

Bug

LLM prompt

Complete GW20-018: Reduce small-write process and storage overhead with measured evidence.
Read this entire task and its current GitHub issue before implementation.
Verify the current mainline and all prerequisite integrations.
Read AGENTS.md and the three TypeScript policy documents.
Use runtime-backed domain values and injected ports.
Keep parsing and codecs at the adapter boundary.
Preserve v20 APIs, stored identities, CRDT semantics, and failure meanings.
Use one coherent issue, PR, and mainline integration. Do not leave a broken intermediate state.
Reproduce the symptom in Docker. Fill Diagnosis with evidence before fix work.
Test each hypothesis against the actual execution path.
Problem to resolve:
The report records 250 to 320 ms p50 per write and about 28.9 KB per unbatched write. It did not measure the subprocess count.
Acceptance checks:
Measure process count, p50/p95/p99 latency, object bytes, packed bytes, CPU, and RSS before selecting a fix.
Record a numeric improvement target from the baseline before measuring the candidate.
The same workload meets that target without worse retained correctness, cancellation, or same-writer CAS behavior.
Compare 1-intent and 100-intent writes, cold and warm sessions, and short-lived runtimes.
Document explicit batching and durability boundaries. Each returned receipt proves its own retained publication.
Preserve synchronous Lane.write completion. Optional batching cannot silently delay durability.
Read these source paths:
src/application/GitStorage.ts
src/infrastructure/adapters/GitTimelineHistoryAdapter.ts
src/domain/api/WriteRuntime.ts
docs/topics/git-perf.md
Prerequisites: None established. Recheck external readiness.
Scope exclusions: Claim ten subprocesses without a trace, make the 10k host result a CI threshold, or weaken retention for speed.
Run all tests and benchmarks in COPY-based Docker without host repository mounts.
Use the project guarded runner and the shared git-locks authority.
Lock host/heavy-work and the exact worker key together for expensive work.
Limit build caches to 20 GiB, test data to 4 GiB, and logs to 128 MiB.
Require 50 GiB free on host and Docker backing storage before heavy work.
Enforce disk accounting, CPU, memory, timeout, and child-process shutdown.
Do not start an unguarded workload. Do not bypass a resource refusal.
Require full touched-code coverage for refactors and an all-green manual SSJS scorecard.
Record source, image, command, fixtures, results, limits, and the regression witness.
Commit only your own files. Do not amend, rebase, force, push, or publish without authorization.
Report incomplete or blocked checks. Do not claim completion from a narrow green test.
Link the issue, coherent PR, and actual mainline integration when those actions are authorized.

1. Background Context

Source: FEEDBACK-git-warp.md, sections 13.
The report SHA-256 is 9b15209d51cd705059a9e6bfebf0f6681ff832924c1462b99ff3cfebde6bacae.
The report uses npm git-warp 20.0.0 and git-cas 6.5.11.
Its runtime results come from macOS host experiments. This planning task did not rerun them.
Source review baseline: ceb58e65.

Tracker: #974.

Read these source and test surfaces before work:

2. Observed

The report records 250 to 320 ms p50 per write and about 28.9 KB per unbatched write. It did not measure the subprocess count.

2b. Expected

Measure process count, p50/p95/p99 latency, object bytes, packed bytes, CPU, and RSS before selecting a fix.

2c. Reproduction

Source report: experiments/07-performance/run.ts. The timings are macOS host measurements, not Docker results.

Golden regression: Run bounded 100-write and 1000-write workloads with 1 and 100 intents per write. Compare baseline and candidate under the same image and budgets.

2d. Blast radius

Population: applications that issue many small atomic writes.
Observed population count: unknown. The report measures fixtures, not deployed users.
No user-count query exists in this task. Do not infer a count from fixture size.
Since: reported in v20.0.0 on 2026-10-06. The introducing commit remains unverified.
Downstream: The report records 250 to 320 ms p50 per write and about 28.9 KB per unbatched write. It did not measure the subprocess count.

2e. Hypothesis

Repeated one-shot history and ref operations dominate small writes. Attribute measured subprocess and retention costs. Reject the hypothesis if another phase dominates.

2f. Diagnosis

3. Prerequisites

None established as an internal issue dependency. Verify the current boundary and external readiness before activation.

4. Scope

In: The report records 250 to 320 ms p50 per write and about 28.9 KB per unbatched write. It did not measure the subprocess count.

Out: Claim ten subprocesses without a trace, make the 10k host result a CI threshold, or weaken retention for speed.

Safe intermediate state: one independently mergeable PR passes its relevant checks after its prerequisites.
Existing supported applications remain usable. Historical data remains intact.

5. Why now

This flaw blocks a supported Synapse store workflow or gives its operator incorrect guidance.
Resolve it within v20.1.0 without reducing the existing causal-history commitment.

6. Risks

Source inspection does not establish runtime or performance results.
Preserve historical identities, validation, and evidence. Do not weaken refusal to obtain a green result.
If the fix requires a breaking contract, record the conflict before activation. Do not hide it within a minor release.

7. Definition of Done

A reproducible regression fails on the baseline and passes on the candidate.

  • Measure process count, p50/p95/p99 latency, object bytes, packed bytes, CPU, and RSS before selecting a fix.
  • Record a numeric improvement target from the baseline before measuring the candidate.
  • The same workload meets that target without worse retained correctness, cancellation, or same-writer CAS behavior.
  • Compare 1-intent and 100-intent writes, cold and warm sessions, and short-lived runtimes.
  • Document explicit batching and durability boundaries. Each returned receipt proves its own retained publication.
  • Preserve synchronous Lane.write completion. Optional batching cannot silently delay durability.

Golden: Run bounded 100-write and 1000-write workloads with 1 and 100 intents per write. Compare baseline and candidate under the same image and budgets.

Edges: Normal GC; session close; cancellation; publication failure; reopen; multiple writers; incompressible data.

All acceptance checks have evidence from the exact candidate.
Relevant lint, typecheck, compatibility, and Docker checks pass.
The manual SSJS scorecard is green. Refactor coverage reaches 100% on touched code.
Record the issue, PR, and mainline integration commit. Record any external release gate.
A required check with missing evidence remains incomplete.

8. Stakeholders

James Ross: git-warp maintainer and acceptance owner.
Synapse store authors: applications need correct writes, reads, history, and operational guidance.
Library and CLI consumers: existing v20 contracts must remain usable.

9. Related Issues

Activity

  1. added this to the v20.1.0 milestone on Oct 7, 2026
  2. added
    type:bugDefect or incorrect behavior.
    priority:nextNext in line after active work.
    area:runtimePrimary work area: runtime.
    status:availableOpen and available for prioritization; not blocked or actively in progress.
    on Oct 7, 2026
  3. linear-code commented on Oct 7, 2026

    @linear-code
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:runtimePrimary work area: runtime.priority:nextNext in line after active work.status:availableOpen and available for prioritization; not blocked or actively in progress.type:bugDefect or incorrect behavior.

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions