Repository navigation
perf(export): serialize HTTP exports with bounded event buffering - #677
Merged
Merged
Conversation
4 of 5 tasks
|
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## master #677 +/- ##
==========================================
+ Coverage 70.81% 80.22% +9.40%
==========================================
Files 51 68 +17
Lines 2916 6007 +3091
==========================================
+ Hits 2065 4819 +2754
- Misses 851 1188 +337 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
TimeToBuildBob
added a commit
to TimeToBuildBob/aw-server-rust
that referenced
this pull request
Sep 23, 2026
Serialize to a tempfile on the worker (disk-paced, same as ActivityWatch#677), then copy to the response pipe from the export thread. Headers still go out before the body; unread or slow downloads no longer block heartbeats. ActivityWatch/aw-android#228 Git-Session-Id: 3f22e322-f840-50e9-aa06-6db541c7f001
ErikBjare
pushed a commit
that referenced
this pull request
Sep 23, 2026
* fix(export): send HTTP headers before serializing large exports #677 still wrote the full JSON tempfile before responding, so a ~500k-event export could sit silent until the 30s web UI timeout. Open a pipe, return 200 + Content-Disposition immediately, and serialize into the body. Missing buckets still 404 before headers. Mid-stream failures truncate the download instead of hanging the connection. ActivityWatch/aw-android#228 Git-Session-Id: a180614b-5a5a-5d29-84d8-e1c4e076b990 * fix(export): don't let a slow client stall the datastore worker Serialize to a tempfile on the worker (disk-paced, same as #677), then copy to the response pipe from the export thread. Headers still go out before the body; unread or slow downloads no longer block heartbeats. ActivityWatch/aw-android#228 Git-Session-Id: 3f22e322-f840-50e9-aa06-6db541c7f001
TimeToBuildBob
added a commit
to TimeToBuildBob/aw-server-rust
that referenced
this pull request
Sep 25, 2026
Extend `export_memory` with `csv` and `csv-union` modes so the CSV export path has the same reproducible memory measurement that ActivityWatch#677 added for the JSON export. `csv-union` gives every event a distinct data key to exercise the MAX_CSV_DATA_COLUMNS fallback. Measured with /usr/bin/time -v (release build, synthetic in-memory data, includes the in-memory input database): csv 100k events 45 MB peak RSS, 31 MB CSV csv 500k events 213 MB peak RSS, 155 MB CSV csv-union 500k events 148 MB peak RSS, 35 MB CSV stream (JSON) 500k events 214 MB peak RSS materialized 500k events 913 MB peak RSS (pre-ActivityWatch#677 JSON baseline) Peak RSS for the CSV path is ~4x lower than the materialized baseline on the same dataset, and the 500k-distinct-key case stays bounded (one JSON data column) instead of emitting events x keys cells. Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675
ErikBjare
pushed a commit
that referenced
this pull request
Sep 26, 2026
* feat(server): add streaming CSV export endpoint for bucket events
GET /api/0/buckets/{id}/export/csv streams events as CSV, sending HTTP
headers before serialization so large buckets don't look like a hung
connection on Android WebView or slow networks.
Mirrors the existing streaming JSON export (BucketsExportRocket / #721):
OS pipe + background thread serializes into a tempfile, then copies to
the client. The route is at /export/csv (parallel to /export for JSON)
to avoid any Rocket route collision with the single-event endpoint.
RFC-4180 CSV: id, timestamp, duration (fractional seconds), then all
data-map keys from the first event. Fields containing commas, quotes, or
newlines are double-quoted and internal quotes are doubled.
Test: csv_export_returns_csv_with_correct_headers_and_missing_bucket_errors
Git-Session-Id: cdc6
* fix(export): stream CSV rows, preserve ns duration, neutralize formulas
Address Greptile P1s on #722:
- Write CSV from SQL row-by-row on the datastore worker (no full Vec + String)
- Format duration from nanoseconds so 0.0015s is not truncated to 0.001
- Neutralize spreadsheet formula prefixes (=, +, -, @)
- Preflight LIMIT 1 before 200 so worker/SQL failures still return JSON
Git-Session-Id: 79e64905-13b9-5822-b880-3863140ba2d7
* fix(export): move CSV serialization off datastore worker
The ExportEventsCsv worker command blocked the single shared datastore
worker for the entire duration of the export — heartbeats and all other
requests queued until serialization finished.
Fix: remove ExportEventsCsv from the worker. The background thread now
calls get_events() (worker holds the DB lock only for the SQL read), then
serializes CSV via write_csv_from_events() off-worker. The worker is free
for other requests during the write.
write_csv_from_events() is the Vec<Event>-based counterpart to the
connection-based write_events_csv(); both produce identical RFC-4180 output
with ns-precision durations and formula neutralization.
Git-Session-Id: bb28
* fix(export): surface CSV staging flush errors before copy
BufWriter::drop swallows a final flush failure, so a full staging
filesystem still rewound and copied a truncated CSV under 200 OK.
Keep the writer in a local, flush explicitly, and flush inside
write_csv_from_events / write_events_csv so the error propagates.
Git-Session-Id: 01a0ce79
* fix(export): neutralize whitespace-prefixed formulas; union CSV keys across events
Git-Session-Id: 4a49419b-e1d7-5734-b196-44aca96d17ff
* fix(export): stream CSV export rows on the worker, bound column count
The CSV export endpoint fetched every matching event into a `Vec<Event>`
via `get_events`, serialized that slice into a second full-size CSV buffer,
and then copied the staging file to the client. For the 500k-event buckets
this endpoint exists to serve, the full event set plus the serialized CSV
were both live in memory at once — the server-side OOM the endpoint was
meant to remove, moved from the webview to the server.
It also built its column set from the union of every event's data keys, so
a bucket with many distinct keys produced `events × keys` cells (quadratic
when keys grow with events) written to disk before any body bytes were sent.
Changes:
- `aw-datastore`: add `Command::ExportCsv` / `Response::ExportCsv` and
`Datastore::export_csv_to_file`, which run the existing row-at-a-time
`write_events_csv` on the worker straight into the staging file.
- `aw-server`: `spawn_csv_export_stream` now uses that instead of
`get_events` + `write_csv_from_events`, so peak memory is O(distinct
keys) for the header pre-pass rather than O(events). The staging-file
pattern (and its slow/dropped-download protection) is unchanged and now
matches the JSON export.
- `aw-datastore`: cap the union of data keys at `MAX_CSV_DATA_COLUMNS`
(32). Above the cap the export emits `id,timestamp,duration,data` with
each event's full data object as JSON, bounding the worst case while
still dropping no keys.
- Remove `write_csv_from_events` (the materializing variant) so the HTTP
path cannot regress to it; its tests now exercise `write_events_csv`.
Verified: `cargo test -p aw-datastore --lib export` 10 passed (incl. new
`csv_falls_back_to_a_json_data_column_for_wide_schemas`), `cargo test
-p aw-server --test api csv_export` passed, `cargo fmt --check` and
`cargo clippy` clean.
Git-Session-Id: 366e2805-96f6-5ba1-a390-5e3501978b44
* perf(export): profile CSV export memory with the reproducible example
Extend `export_memory` with `csv` and `csv-union` modes so the CSV export
path has the same reproducible memory measurement that #677 added for the
JSON export. `csv-union` gives every event a distinct data key to exercise
the MAX_CSV_DATA_COLUMNS fallback.
Measured with /usr/bin/time -v (release build, synthetic in-memory data,
includes the in-memory input database):
csv 100k events 45 MB peak RSS, 31 MB CSV
csv 500k events 213 MB peak RSS, 155 MB CSV
csv-union 500k events 148 MB peak RSS, 35 MB CSV
stream (JSON) 500k events 214 MB peak RSS
materialized 500k events 913 MB peak RSS (pre-#677 JSON baseline)
Peak RSS for the CSV path is ~4x lower than the materialized baseline on the
same dataset, and the 500k-distinct-key case stays bounded (one JSON data
column) instead of emitting events x keys cells.
Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675
* fix(export): bound the CSV key pre-pass; sanitize Content-Disposition
In-band review findings on the CSV export:
- The key pre-pass collected the full union of data keys before deciding to
fall back to the single JSON `data` column, so a bucket whose events each
carry a distinct key still held every key in memory -- the unbounded term the
PR claims to remove. Stop collecting at MAX_CSV_DATA_COLUMNS, where the
fallback decision is already made. Measured on the 500k-distinct-key case:
144 MiB -> 70 MiB peak RSS, 0.90 s -> 0.51 s.
- `duration_csv` went through `num_nanoseconds().unwrap_or(0)`, which silently
exports `0.000000000` for a duration outside the i64 nanosecond range. Compose
from whole seconds and the signed subsecond remainder instead.
- `Content-Disposition` interpolated the client-supplied bucket id verbatim.
Ids are not restricted at creation and Rocket percent-decodes path segments,
so CR/LF in an id would split the response headers. Header values now strip
control characters and quote/backslash/semicolon -- applied to the CSV
filename and to the pre-existing bucket-export filename (#721 had the same
exposure).
Tests: streamed_csv_stops_key_pass_at_the_cap_without_truncating_rows,
sanitize_header_value_strips_header_metacharacters.
Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675
* fix(export): terminate CSV records with CRLF per RFC-4180
The endpoint advertises RFC-4180 CSV, which delimits each record with CRLF;
the writer emitted a bare LF. Readers accept both, but the format claim should
match the bytes. Pinned by streamed_csv_terminates_every_record_with_crlf.
Also applies to the CSV the staging-file path copies to the response.
Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675
* fix(export): quote the Content-Disposition filename
Bucket ids may contain spaces, so `filename=aw-events-export-my bucket.csv` was
a malformed parameter that clients could drop. Filenames are now emitted as a
quoted-string via a single `content_disposition()` helper, applied to the CSV
export and to the bucket export (which had the same unquoted form). The
sanitizer already removes `"`, `\` and control characters, so the quoted value
cannot be broken out of.
Tests: content_disposition_quotes_the_filename, plus the three existing
Content-Disposition header assertions updated to the quoted form.
Git-Session-Id: 050403af-7294-5a68-bc97-2dab10c51675
* fix(export): correct negative CSV durations and saturate the limit
Two correctness fixes in the streamed CSV export path:
- `duration_csv` split a negative duration into `num_seconds()` and
`subsec_nanos()` while both carried their own sign handling. chrono
stores -1.5s as secs=-2, nanos=500_000_000, so the row rendered as
"-2.500000000": a full second plus the fraction away from the event's
real duration, corrupting any timeline built from the export. Take the
magnitude once, before splitting.
- `Some(l) => l as i64` wrapped a `u64` limit above `i64::MAX` to a
negative value, and SQLite reads a negative LIMIT as "no limit" — so
`?limit=18446744073709551615` exported the whole bucket instead of
capping it. Saturate via `i64::try_from` (new `sql_limit` helper).
Also create the CSV staging file during preflight rather than inside the
spawned writer thread, so a full staging filesystem returns a JSON 500
before the response commits `200 OK` instead of an empty body that looks
like a genuinely empty bucket.
Tests: `csv_duration_handles_negative_durations`,
`csv_limit_saturates_instead_of_wrapping`.
Git-Session-Id: a0ee77dc-cf96-5982-817e-85a571138c85
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
HTTP exports currently hold every event and then a complete JSON string in memory. Serialize event rows incrementally through the existing datastore worker into a private temporary file, then serve that file as the response body. This preserves visibility of acknowledged, uncommitted writes and allows database/serialization/write failures to return an HTTP error before a download begins.
Both single-bucket and all-bucket exports retain the JSON format and download naming. The export query explicitly uses the ordered starttime index, avoiding an unbounded sort even if additional indexes exist.
Part of #671 (HTTP export memory). ActivityWatch/aw-android#229 fixes WebView download saving; this PR addresses the separate Rust HTTP buffering path.
Validation:
cargo test -p aw-datastore --libandcargo test -p aw-server: export serialization/error tests and all server/API tests pass.export_memory, a reproducible example using only synthetic in-memory data. On 100,000 events with 256-character titles,/usr/bin/time -lmeasured peak RSS of 203,177,984 bytes for materialization versus 54,788,096 bytes for incremental serialization. This includes the in-memory input database and excludes temporary-file I/O; it measures the serialization memory reduction, not HTTP throughput.Tradeoffs: the completed JSON export occupies temporary disk space (plaintext, with private-file permissions) until the response is dropped. The datastore worker remains occupied while the snapshot is serialized; this PR does not introduce parallel readers or change the consistency contract discussed in #646.