Skip to content

explore(stream): binary frame format negotiated via Accept, alongside SSE #467

Description

@EricAndrechek

Area: streaming (api + sdk) · exploratory — do not start before #465 has numbers.

GET /v1/stream speaks text/event-stream, whose framing is UTF-8 and line-oriented by definition. Now that the SDK reads the response body itself over fetch + ReadableStream (#203) instead of delegating to EventSource, a second, binary wire format becomes possible for the first time — negotiated via Accept, with SSE remaining the default. This issue is to track and experiment with that, not a commitment to ship it.

Why it is newly possible

EventSource could only ever consume text/event-stream; the browser API gave us no say. A fetch-based transport reads raw bytes, so the client can accept any framing we define. Browsers handle binary response bodies fine — nothing about this is browser-blocked.

What binary would actually buy

  • No base64 for genuinely binary values. SSE cannot carry arbitrary bytes in data:; encoding costs +33%.
  • O(1) framing. wireFrame (internal/stream/hub.go) currently splits each payload on \n and emits one data: line per fragment, and the client re-joins them. A length-prefixed frame is a varint and a memcpy.
  • Native-typed values. Timestamps as int64, UUIDs as 16 bytes, numerics unboxed — rather than JSON strings re-parsed per event. For numeric-heavy, high-frequency tables this is where the real win would be, well beyond what compression recovers.
  • Per-frame control. Type tags for event / keepalive / replay would replace the current convention of inferring replay-vs-live from context (relevant to feat(sdk): let liveQuery consumers distinguish NATS replay from genuinely-live events #98).

Sketch

Accept: application/vnd.wavehouse.stream+binary, text/event-stream;q=0.9

Server content-negotiates; absent or unsupported, it answers text/event-stream exactly as today. Frame shape, roughly:

varint length | u8 type (event|keepalive|replay) | u64 event id / received_ts | payload

Resumption stays transport-agnostic — Last-Event-ID is a request header and works identically; only the response-side id moves from an id: line into the frame header.

What it costs, and why the bar is high

  • curl stops working. Today curl -N .../v1/stream is a first-class debugging path, documented and used in dogfooding. A binary stream needs a tool we'd have to write.
  • "Any spec-compliant SSE client" is lost for consumers who opt in — other-language SSE libraries, hand-rolled EventSource, observability tooling that speaks SSE.
  • Two server paths and two client paths, permanently. Once negotiated and adopted, the SSE path can never be dropped, so this is strictly additive maintenance forever — including for every future streaming feature (feat(sdk): SSE replay semantics across multiple subscriptions #204 multiplexing, feat(sdk): let liveQuery consumers distinguish NATS replay from genuinely-live events #98 replay marking), each of which then has to be designed twice.
  • The Hub's one-serialization-per-(topic, role) optimization would need a second serialization per format, doubling that work when both are in use.

Sequencing — deliberate

#465 (gzip) first. Compression is available today, needs no client change, preserves curl and every SSE client, and plausibly captures most of the bandwidth win. Only if its measurements show the remaining gap is large — and specifically if the win is in parse/CPU cost or native typing rather than bytes on the wire — is binary worth opening. Framed the other way: #465 measures whether this issue has a business case.

Note also that POST-with-a-body does not require binary. The multiplexing that #204 needs is unblocked by fetch alone, and a POST response carrying text/event-stream is still fully compliant SSE and still curl-able — so "we need binary for multi-subscription streams" is not a valid argument for this. See the note on #204.

Suggested experiment (if picked up)

  1. Capture a representative frame stream from a numeric-heavy table.
  2. Compare four ways: raw SSE, gzipped SSE (perf(stream): evaluate gzip on /v1/stream with per-frame flush #465), binary frames with JSON payloads, binary frames with natively-typed payloads.
  3. Measure bytes on the wire and client-side parse CPU — the second is the axis compression can't help with and the only one that would justify the maintenance cost.

Related: #465 (compression — do first), #203 (made this possible), #204 (multiplexing — does not need this), #98 (replay-vs-live marking), #152 (per-subscriber buffers).

Activity

  1. added
    enhancementNew feature or request
    area/apiHTTP handlers, routing, middleware
    area/sdkTypeScript SDK (clients/ts/)
    area/streamingSSE / live-query delivery path (/v1/stream)
    on Aug 13, 2026
  2. coderabbitai commented on Aug 13, 2026

    @coderabbitai
    🔗 Related PRs

    #124 - chore(api)!: drop hub wildcard fan-out; SSE/WS use ?table= (closes #100) [merged]
    #353 - perf(stream): project SSE frames once per role, not per subscriber [merged]
    #381 - fix(stream): apply policy row-filter per subscriber on SSE [open]


    🧪 Issue enrichment is currently in open beta.

    You can configure auto-planning by selecting labels in the issue_enrichment configuration.

    To disable automatic issue enrichment, add the following to your .coderabbit.yaml:

    issue_enrichment:
      auto_enrich:
        enabled: false

    💬 Have feedback or questions? Drop into our discord!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/apiHTTP handlers, routing, middlewarearea/sdkTypeScript SDK (clients/ts/)area/streamingSSE / live-query delivery path (/v1/stream)enhancementNew feature or request

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions