Skip to content

perf(export): serialize HTTP exports with bounded event buffering - #677

Merged
ErikBjare merged 1 commit into
ActivityWatch:masterfrom
0xbrayo:perf/bounded-exports
Sep 15, 2026
Merged

ErikBjare merged 1 commit into
ActivityWatch:masterfrom
0xbrayo:perf/bounded-exports

Conversation

@0xbrayo

@0xbrayo 0xbrayo commented Sep 11, 2026

Copy link
Copy Markdown
Member

HTTP exports currently hold every event and then a complete JSON string in memory. Serialize event rows incrementally through the existing datastore worker into a private temporary file, then serve that file as the response body. This preserves visibility of acknowledged, uncommitted writes and allows database/serialization/write failures to return an HTTP error before a download begins.

Both single-bucket and all-bucket exports retain the JSON format and download naming. The export query explicitly uses the ordered starttime index, avoiding an unbounded sort even if additional indexes exist.

Part of #671 (HTTP export memory). ActivityWatch/aw-android#229 fixes WebView download saving; this PR addresses the separate Rust HTTP buffering path.

Validation:

  • cargo test -p aw-datastore --lib and cargo test -p aw-server: export serialization/error tests and all server/API tests pass.
  • Regression coverage compares JSON with the previous materialized format, including empty buckets, tied timestamps, clipping, nested/Unicode payloads, corrupt rows, writer failure, missing buckets, response headers, and acknowledged writes before commit.
  • Added export_memory, a reproducible example using only synthetic in-memory data. On 100,000 events with 256-character titles, /usr/bin/time -l measured peak RSS of 203,177,984 bytes for materialization versus 54,788,096 bytes for incremental serialization. This includes the in-memory input database and excludes temporary-file I/O; it measures the serialization memory reduction, not HTTP throughput.
  • No live database or running server was used for tests or profiling.

Tradeoffs: the completed JSON export occupies temporary disk space (plaintext, with private-file permissions) until the response is dropped. The datastore worker remains occupied while the snapshot is serialized; this PR does not introduce parallel readers or change the consistency contract discussed in #646.

@greptile-apps

greptile-apps Bot commented Sep 11, 2026

Copy link
Copy Markdown

RetriggerConfidence Score: 5/5

The PR appears safe to merge; no concrete correctness, security, or repository-rule violation remains.

Summary

  • Event rows are queried and serialized individually through the datastore worker.
  • Completed files are rewound and streamed as JSON responses while preserving export naming.
  • Single-bucket and all-bucket endpoints share the new export responder.
  • Regression tests cover pending writes, missing buckets, headers, serialization parity, corrupt rows, and writer failures.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[HTTP export request] --> B[Create private tempfile]
  B --> C[Send Export command to datastore worker]
  C --> D[Query bucket events using ordered index]
  D --> E[Serialize rows incrementally]
  E --> F[Flush and return completed file]
  F --> G[Seek to beginning]
  G --> H[Stream JSON response]
Loading

Reviews (1) · Last reviewed commit: "perf(export): serialize HTTP exports wit..."

@codecov

codecov Bot commented Sep 11, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 91.22807% with 10 lines in your changes missing coverage. Please review.
✅ Project coverage is 80.22%. Comparing base (656f3c9) to head (d1ad337).
⚠️ Report is 103 commits behind head on master.

Files with missing lines Patch % Lines
aw-datastore/src/export.rs 87.23% 6 Missing ⚠️
aw-datastore/src/worker.rs 81.81% 2 Missing ⚠️
aw-server/src/endpoints/util.rs 85.71% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##           master     #677      +/-   ##
==========================================
+ Coverage   70.81%   80.22%   +9.40%     
==========================================
  Files          51       68      +17     
  Lines        2916     6007    +3091     
==========================================
+ Hits         2065     4819    +2754     
- Misses        851     1188     +337     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@ErikBjare
ErikBjare merged commit b834a8f into ActivityWatch:master Sep 15, 2026
8 checks passed
TimeToBuildBob added a commit to TimeToBuildBob/aw-server-rust that referenced this pull request Sep 23, 2026
Serialize to a tempfile on the worker (disk-paced, same as ActivityWatch#677), then
copy to the response pipe from the export thread. Headers still go out
before the body; unread or slow downloads no longer block heartbeats.

ActivityWatch/aw-android#228

Git-Session-Id: 3f22e322-f840-50e9-aa06-6db541c7f001
ErikBjare pushed a commit that referenced this pull request Sep 23, 2026
* fix(export): send HTTP headers before serializing large exports

#677 still wrote the full JSON tempfile before responding, so a ~500k-event
export could sit silent until the 30s web UI timeout. Open a pipe, return
200 + Content-Disposition immediately, and serialize into the body.

Missing buckets still 404 before headers. Mid-stream failures truncate the
download instead of hanging the connection.

ActivityWatch/aw-android#228

Git-Session-Id: a180614b-5a5a-5d29-84d8-e1c4e076b990

* fix(export): don't let a slow client stall the datastore worker

Serialize to a tempfile on the worker (disk-paced, same as #677), then
copy to the response pipe from the export thread. Headers still go out
before the body; unread or slow downloads no longer block heartbeats.

ActivityWatch/aw-android#228

Git-Session-Id: 3f22e322-f840-50e9-aa06-6db541c7f001
TimeToBuildBob added a commit to TimeToBuildBob/aw-server-rust that referenced this pull request Sep 25, 2026
Extend `export_memory` with `csv` and `csv-union` modes so the CSV export
path has the same reproducible memory measurement that ActivityWatch#677 added for the
JSON export. `csv-union` gives every event a distinct data key to exercise
the MAX_CSV_DATA_COLUMNS fallback.

Measured with /usr/bin/time -v (release build, synthetic in-memory data,
includes the in-memory input database):

  csv              100k events    45 MB peak RSS,  31 MB CSV
  csv              500k events   213 MB peak RSS, 155 MB CSV
  csv-union        500k events   148 MB peak RSS,  35 MB CSV
  stream (JSON)    500k events   214 MB peak RSS
  materialized     500k events   913 MB peak RSS  (pre-ActivityWatch#677 JSON baseline)

Peak RSS for the CSV path is ~4x lower than the materialized baseline on the
same dataset, and the 500k-distinct-key case stays bounded (one JSON data
column) instead of emitting events x keys cells.

Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675
ErikBjare pushed a commit that referenced this pull request Sep 26, 2026
* feat(server): add streaming CSV export endpoint for bucket events

GET /api/0/buckets/{id}/export/csv streams events as CSV, sending HTTP
headers before serialization so large buckets don't look like a hung
connection on Android WebView or slow networks.

Mirrors the existing streaming JSON export (BucketsExportRocket / #721):
OS pipe + background thread serializes into a tempfile, then copies to
the client. The route is at /export/csv (parallel to /export for JSON)
to avoid any Rocket route collision with the single-event endpoint.

RFC-4180 CSV: id, timestamp, duration (fractional seconds), then all
data-map keys from the first event. Fields containing commas, quotes, or
newlines are double-quoted and internal quotes are doubled.

Test: csv_export_returns_csv_with_correct_headers_and_missing_bucket_errors
Git-Session-Id: cdc6

* fix(export): stream CSV rows, preserve ns duration, neutralize formulas

Address Greptile P1s on #722:
- Write CSV from SQL row-by-row on the datastore worker (no full Vec + String)
- Format duration from nanoseconds so 0.0015s is not truncated to 0.001
- Neutralize spreadsheet formula prefixes (=, +, -, @)
- Preflight LIMIT 1 before 200 so worker/SQL failures still return JSON

Git-Session-Id: 79e64905-13b9-5822-b880-3863140ba2d7

* fix(export): move CSV serialization off datastore worker

The ExportEventsCsv worker command blocked the single shared datastore
worker for the entire duration of the export — heartbeats and all other
requests queued until serialization finished.

Fix: remove ExportEventsCsv from the worker. The background thread now
calls get_events() (worker holds the DB lock only for the SQL read), then
serializes CSV via write_csv_from_events() off-worker. The worker is free
for other requests during the write.

write_csv_from_events() is the Vec<Event>-based counterpart to the
connection-based write_events_csv(); both produce identical RFC-4180 output
with ns-precision durations and formula neutralization.

Git-Session-Id: bb28

* fix(export): surface CSV staging flush errors before copy

BufWriter::drop swallows a final flush failure, so a full staging
filesystem still rewound and copied a truncated CSV under 200 OK.
Keep the writer in a local, flush explicitly, and flush inside
write_csv_from_events / write_events_csv so the error propagates.

Git-Session-Id: 01a0ce79

* fix(export): neutralize whitespace-prefixed formulas; union CSV keys across events

Git-Session-Id: 4a49419b-e1d7-5734-b196-44aca96d17ff

* fix(export): stream CSV export rows on the worker, bound column count

The CSV export endpoint fetched every matching event into a `Vec<Event>`
via `get_events`, serialized that slice into a second full-size CSV buffer,
and then copied the staging file to the client. For the 500k-event buckets
this endpoint exists to serve, the full event set plus the serialized CSV
were both live in memory at once — the server-side OOM the endpoint was
meant to remove, moved from the webview to the server.

It also built its column set from the union of every event's data keys, so
a bucket with many distinct keys produced `events × keys` cells (quadratic
when keys grow with events) written to disk before any body bytes were sent.

Changes:

- `aw-datastore`: add `Command::ExportCsv` / `Response::ExportCsv` and
  `Datastore::export_csv_to_file`, which run the existing row-at-a-time
  `write_events_csv` on the worker straight into the staging file.
- `aw-server`: `spawn_csv_export_stream` now uses that instead of
  `get_events` + `write_csv_from_events`, so peak memory is O(distinct
  keys) for the header pre-pass rather than O(events). The staging-file
  pattern (and its slow/dropped-download protection) is unchanged and now
  matches the JSON export.
- `aw-datastore`: cap the union of data keys at `MAX_CSV_DATA_COLUMNS`
  (32). Above the cap the export emits `id,timestamp,duration,data` with
  each event's full data object as JSON, bounding the worst case while
  still dropping no keys.
- Remove `write_csv_from_events` (the materializing variant) so the HTTP
  path cannot regress to it; its tests now exercise `write_events_csv`.

Verified: `cargo test -p aw-datastore --lib export` 10 passed (incl. new
`csv_falls_back_to_a_json_data_column_for_wide_schemas`), `cargo test
-p aw-server --test api csv_export` passed, `cargo fmt --check` and
`cargo clippy` clean.

Git-Session-Id: 366e2805-96f6-5ba1-a390-5e3501978b44

* perf(export): profile CSV export memory with the reproducible example

Extend `export_memory` with `csv` and `csv-union` modes so the CSV export
path has the same reproducible memory measurement that #677 added for the
JSON export. `csv-union` gives every event a distinct data key to exercise
the MAX_CSV_DATA_COLUMNS fallback.

Measured with /usr/bin/time -v (release build, synthetic in-memory data,
includes the in-memory input database):

  csv              100k events    45 MB peak RSS,  31 MB CSV
  csv              500k events   213 MB peak RSS, 155 MB CSV
  csv-union        500k events   148 MB peak RSS,  35 MB CSV
  stream (JSON)    500k events   214 MB peak RSS
  materialized     500k events   913 MB peak RSS  (pre-#677 JSON baseline)

Peak RSS for the CSV path is ~4x lower than the materialized baseline on the
same dataset, and the 500k-distinct-key case stays bounded (one JSON data
column) instead of emitting events x keys cells.

Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675

* fix(export): bound the CSV key pre-pass; sanitize Content-Disposition

In-band review findings on the CSV export:

- The key pre-pass collected the full union of data keys before deciding to
  fall back to the single JSON `data` column, so a bucket whose events each
  carry a distinct key still held every key in memory -- the unbounded term the
  PR claims to remove. Stop collecting at MAX_CSV_DATA_COLUMNS, where the
  fallback decision is already made. Measured on the 500k-distinct-key case:
  144 MiB -> 70 MiB peak RSS, 0.90 s -> 0.51 s.
- `duration_csv` went through `num_nanoseconds().unwrap_or(0)`, which silently
  exports `0.000000000` for a duration outside the i64 nanosecond range. Compose
  from whole seconds and the signed subsecond remainder instead.
- `Content-Disposition` interpolated the client-supplied bucket id verbatim.
  Ids are not restricted at creation and Rocket percent-decodes path segments,
  so CR/LF in an id would split the response headers. Header values now strip
  control characters and quote/backslash/semicolon -- applied to the CSV
  filename and to the pre-existing bucket-export filename (#721 had the same
  exposure).

Tests: streamed_csv_stops_key_pass_at_the_cap_without_truncating_rows,
sanitize_header_value_strips_header_metacharacters.

Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675

* fix(export): terminate CSV records with CRLF per RFC-4180

The endpoint advertises RFC-4180 CSV, which delimits each record with CRLF;
the writer emitted a bare LF. Readers accept both, but the format claim should
match the bytes. Pinned by streamed_csv_terminates_every_record_with_crlf.

Also applies to the CSV the staging-file path copies to the response.

Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675

* fix(export): quote the Content-Disposition filename

Bucket ids may contain spaces, so `filename=aw-events-export-my bucket.csv` was
a malformed parameter that clients could drop. Filenames are now emitted as a
quoted-string via a single `content_disposition()` helper, applied to the CSV
export and to the bucket export (which had the same unquoted form). The
sanitizer already removes `"`, `\` and control characters, so the quoted value
cannot be broken out of.

Tests: content_disposition_quotes_the_filename, plus the three existing
Content-Disposition header assertions updated to the quoted form.

Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675

* fix(export): correct negative CSV durations and saturate the limit

Two correctness fixes in the streamed CSV export path:

- `duration_csv` split a negative duration into `num_seconds()` and
  `subsec_nanos()` while both carried their own sign handling. chrono
  stores -1.5s as secs=-2, nanos=500_000_000, so the row rendered as
  "-2.500000000": a full second plus the fraction away from the event's
  real duration, corrupting any timeline built from the export. Take the
  magnitude once, before splitting.
- `Some(l) => l as i64` wrapped a `u64` limit above `i64::MAX` to a
  negative value, and SQLite reads a negative LIMIT as "no limit" — so
  `?limit=18446744073709551615` exported the whole bucket instead of
  capping it. Saturate via `i64::try_from` (new `sql_limit` helper).

Also create the CSV staging file during preflight rather than inside the
spawned writer thread, so a full staging filesystem returns a JSON 500
before the response commits `200 OK` instead of an empty body that looks
like a genuinely empty bucket.

Tests: `csv_duration_handles_negative_durations`,
`csv_limit_saturates_instead_of_wrapping`.

Git-Session-Id: a0ee77dc-cf96-5982-817e-85a571138c85
@0xbrayo
0xbrayo deleted the perf/bounded-exports branch October 8, 2026 11:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants