Skip to content

Build a V3 streaming inline event ingestion pipeline #2368

Description

@ejsmith

Summary

Build a new V3 event ingestion API and pipeline optimized for minimum cost, bounded memory, low allocations, and horizontal scale-out.

V3 will:

  • Use ASP.NET Core Minimal APIs only. Do not add a controller.
  • Accept an asynchronous stream of independent JSON event objects instead of a giant JSON array.
  • Read and deserialize directly from the request PipeReader with System.Text.Json source generation.
  • Process events inline and acknowledge only after the primary event data is durable.
  • Remove the V3 object-storage payload handoff, ingestion queue round trip, worker scheduling delay, second payload read, and second deserialization.
  • Let clients submit a raw stack trace; the server owns stack-trace parsing and structured error construction.
  • Identify events belonging to discarded stacks before billable quota reservation, full error materialization, event indexing, stack statistics, notifications, or other expensive work.
  • Keep notifications, webhooks, archival, and other secondary effects asynchronous through durable, idempotent outbox work.
  • Extract reusable batch-oriented services that V2 can adopt where doing so improves V2 without adding compatibility cost to V3.

This is a new API version and may make breaking contract changes.

Why

The current V2 ingestion path performs several expensive handoffs:

  1. EventController accepts the complete request.
  2. EventPostService writes metadata and the request payload to object storage.
  3. EventPostService enqueues a pointer to the stored payload.
  4. EventPostsJob downloads the complete payload into a byte array.
  5. Compressed payloads are expanded into another complete byte array.
  6. The complete body is parsed into an event collection.
  7. EventPipeline repeatedly filters and materializes context lists between actions.
  8. Stack assignment, event persistence, statistics, notifications, and post-processing run afterward.

Relevant current entry points include:

  • src/Exceptionless.Web/Controllers/EventController.cs
  • src/Exceptionless.Core/Services/EventPostService.cs
  • src/Exceptionless.Core/Jobs/EventPostsJob.cs
  • src/Exceptionless.Core/Pipeline/Base/PipelineBase.cs
  • src/Exceptionless.Core/Pipeline/010_AssignToStackAction.cs
  • src/Exceptionless.Core/Pipeline/040_SaveEventAction.cs
  • src/Exceptionless.Core/Pipeline/060_UpdateStatsAction.cs
  • src/Exceptionless.Core/Pipeline/070_QueueNotificationAction.cs

Discarded stacks are currently detected during AssignToStackAction, after the normal event-processing plugins have run. Plan-limit checks also occur before stack assignment, which means the system can block an event before learning that it belongs to a free discarded stack.

V3 should make the common path direct and make the discarded path exceptionally cheap.

Goals

  • Reduce CPU, allocation rate, memory pressure, network round trips, and storage operations per event.
  • Keep memory bounded by configured microbatch and request limits, independent of total stream duration.
  • Scale API instances horizontally without sticky sessions or instance-owned correctness state.
  • Apply natural backpressure instead of accumulating unbounded server-side work.
  • Preserve at-least-once delivery with deterministic idempotency.
  • Make new clients easy to implement using ordinary streaming HTTP and newline-delimited JSON.
  • Move raw stack-trace parsing and structured frame generation to the server.
  • Do not charge for discarded events.
  • Detect discarded events before full materialization and persistence.
  • Preserve safe JSON encoding and use System.Text.Json only.
  • Provide enough telemetry and benchmarks to prove cost and scaling improvements.

Non-goals and explicit decisions

  • V3 will not use an MVC controller.
  • V3 will not route through the existing reflection/action-based PipelineBase hot path.
  • V3 will not write the incoming payload to object storage before processing.
  • V3 will not enqueue the primary event payload during normal operation.
  • V3 will not acknowledge an event before its primary durable write succeeds.
  • V3 will not automatically enqueue an event after an ambiguous inline failure. The client retries using the same event identifier.
  • Notifications, webhooks, email, archival, and nonessential enrichment are not moved into the request critical path.
  • V2 remains supported and contract-compatible.
  • V3 performance will not be reduced to preserve V2 behavior. Shared improvements must be extracted below the transport boundary.
  • A durable server-side spillover mode is not part of the initial implementation. It can be designed separately if accepting events while Elasticsearch is unavailable becomes a product requirement.
  • Do not introduce Newtonsoft.Json, NEST application dependencies, or unsafe relaxed JSON escaping.

Proposed V3 API contract

Routes

Map two Minimal API routes to one static handler:

  • POST /api/v3/events
    • Resolves the default project from the API credential.
  • POST /api/v3/projects/{projectId}/events
    • Uses the explicit project and rejects a credential/project mismatch.

Both routes require AuthorizationRoles.ClientPolicy.

Place route mapping in a focused endpoint module, for example:

src/Exceptionless.Web/Endpoints/EventIngestionV3Endpoints.cs

The endpoint must remain thin: authenticate, resolve request context, invoke the ingestion service, and map the result to HTTP.

Request framing

Primary media type:

application/x-ndjson

Each line is one complete JSON object. No top-level JSON array is required or accepted by the streaming route.

Example:

{"id":"01J...","type":"error","date":"2026-07-13T12:00:00Z","message":"Operation failed","exception_type":"System.InvalidOperationException","stack_trace":"   at Example.Service.Run() in Service.cs:line 42"}
{"id":"01J...","type":"log","date":"2026-07-13T12:00:01Z","message":"Retry scheduled","source":"Example.Service"}

The initial DTO should contain only fields needed by supported event types and should avoid the current client requirement to construct the complete nested error and stack-frame model.

Proposed common fields:

  • id: required stable client event identifier used for idempotency.
  • type: required event type.
  • date: optional; defaults to server receipt time according to documented rules.
  • source: optional.
  • message: optional except where required by a specific type.
  • reference_id: optional business reference; separate from transport idempotency.
  • tags: optional bounded collection.
  • value: optional.
  • data: optional bounded extension data.
  • exception_type: optional error type.
  • stack_trace: raw stack trace string; the server parses it.
  • user, request, environment, and other supported metadata should be represented by small typed V3 contracts rather than the V2 dynamic error graph.

Contract rules:

  • Project and organization identifiers come from the authenticated route/context, not the event payload.
  • Unknown fields may be ignored for additive forward compatibility, but malformed JSON and invalid required fields are rejected.
  • Every string, collection, object depth, per-event byte size, and total segment size has an explicit limit.
  • A top-level array is rejected with a clear V3 contract error.
  • Support one top-level JSON value at a time using the .NET 10 System.Text.Json top-level-values streaming path.
  • Use snake_case and safe Exceptionless JSON encoding.

Streaming implementation

Use source-generated System.Text.Json metadata and deserialize from HttpRequest.BodyReader using the PipeReader overload of JsonSerializer.DeserializeAsyncEnumerable with topLevelValues enabled.

The implementation must:

  • Pass HttpContext.RequestAborted through every asynchronous operation.
  • Avoid copying the request into a string, byte array, MemoryStream, or JsonDocument.
  • Avoid reflection-based serialization metadata on the hot path.
  • Support chunked requests without requiring Content-Length.
  • Support streaming gzip and Brotli request decompression.
  • Enforce compressed-byte, decompressed-byte, per-event, nesting-depth, event-count, and elapsed-segment limits to prevent decompression bombs and unbounded requests.
  • Bypass or replace the current OverageMiddleware behavior for V3 because it requires Content-Length and checks billable quota before discard classification.
  • Configure endpoint-specific request limits rather than increasing global limits without bounds.

Acknowledgement boundary

Do not encourage a literally endless unacknowledged HTTP request. Official clients should stream events immediately into a bounded request segment and close/reopen the request when any configured threshold is reached, for example:

  • Maximum event count.
  • Maximum uncompressed bytes.
  • Maximum elapsed time since the segment opened.
  • Explicit client flush/shutdown.

This keeps the client API streaming and array-free while providing a finite acknowledgement boundary. Clients retain only the current unacknowledged segment.

A lost connection or 5xx response causes the client to replay the segment with the same event ids. Create-only/deterministic persistence makes this safe.

Response contract

Return 200 OK after every event in the segment has reached a terminal result:

{
  "received": 100,
  "persisted": 80,
  "discarded": 17,
  "duplicate": 2,
  "blocked": 1,
  "invalid": 0
}

Discarded and duplicate events are successful acknowledgements and must be removed from the client retry buffer.

Use ProblemDetails for request-level failures:

  • 400 for malformed JSON/framing.
  • 401/403 for authentication/authorization.
  • 404 for an inaccessible or nonexistent project.
  • 413 for compressed, decompressed, segment, or individual-event size limits.
  • 415 for unsupported media type or content encoding.
  • 422 for a segment that cannot produce valid event records.
  • 429 for protective concurrency/rate limits.
  • 503 for disabled ingestion, open downstream circuit, timeout, or unavailable durable storage.

If an earlier microbatch became durable before a later request-level failure, returning a failure is safe because replay uses stable event ids. Document this explicitly.

Target architecture

Core abstractions

Create a purpose-built batch processor in Exceptionless.Core. Names may be refined, but responsibilities should remain narrow:

  • IEventIngestionProcessor
    • Coordinates a bounded microbatch through classification, quota, persistence, and result aggregation.
  • IStackFingerprintService
    • Produces the canonical grouping signature from the minimal V3 event/raw stack trace.
  • IStackRouteResolver
    • Bulk resolves signature to stack id and stack status.
  • IIngestionQuotaService
    • Atomically reserves, commits, and releases billable usage for non-discarded events.
  • IEventMaterializer
    • Builds PersistentEvent and structured Error/StackFrame objects only for surviving events.
  • IEventBatchWriter
    • Performs deterministic bulk stack creation, event writes, and durable outbox writes.
  • IIngestionSideEffectDispatcher or outbox processor
    • Handles noncritical work outside the request.

Keep Web dependent on Core; infrastructure implementations belong in Exceptionless.Insulation where appropriate.

Do not reproduce the current generic action chain for V3. Prefer direct, explicit calls over reflection, repeated LINQ filtering, and per-action context-list allocation.

Processing sequence

For each bounded microbatch:

  1. Resolve immutable request context

    • Authenticate once.
    • Resolve project, organization, configuration, retention, and client metadata once per request.
    • Reject mismatched explicit project routes.
    • Capture configuration needed for canonical stack grouping.
  2. Apply protective admission

    • Enforce concurrent request and in-flight event limits.
    • Count all submitted bytes/events for abuse protection, including discarded events.
    • Reject quickly with 429 or 503 when capacity/downstream health is unavailable.
    • Protective admission is not customer billing.
  3. Deserialize incrementally

    • Read one top-level JSON value at a time.
    • Validate cheap structural requirements immediately.
    • Accumulate only until the configured microbatch count or byte threshold.
    • Stop reading while the bounded microbatch is being processed so Kestrel/TCP backpressure propagates to the client.
    • Consider a channel capacity of one only after benchmarks prove overlap between reading and processing improves throughput without harmful memory growth.
  4. Perform minimal normalization

    • Normalize only fields required for event validity, date/retention checks, manual stacking, and fingerprinting.
    • Reject invalid/expired events before billable work.
    • Do not construct the complete PersistentEvent data graph yet.
  5. Compute canonical stack fingerprints

    • Parse enough of a raw stack trace to determine the same grouping target used by normal stack assignment.
    • Include event type/source fallback and project grouping settings where applicable.
    • Deduplicate identical signatures within the microbatch.
    • Preserve V2/V3 grouping compatibility where it does not impose transport cost on V3.
  6. Bulk resolve stack routes

    • Resolve every unique signature to a lightweight StackRoute containing stack id, status, and any minimal version needed for consistency.
    • Use a microbatch-local dictionary first.
    • Use ICacheClient for distributed routing entries.
    • Bulk query only cache misses from the stack repository with a source projection.
    • Do not load the complete Stack document merely to test discarded status.
    • Short negative caching is allowed, but stack creation must immediately replace the negative entry.
  7. Terminate discarded events

    • Mark every event whose StackRoute status is Discarded as a successful discard.
    • Increment discarded usage/telemetry once per microbatch, not once per event.
    • Release raw event and parser working memory immediately.
    • Do not reserve billable quota.
    • Do not create an event id/document, materialize frames, index an event, increment stack occurrences, create notifications, write outbox work, perform geolocation, or run post-processing.
  8. Reserve billable quota

    • Atomically reserve quota only for active-stack and new-stack survivors.
    • Replace the current read-then-process GetEventsLeftAsync pattern for this path with a scale-out-safe reservation/settlement operation.
    • Split a microbatch deterministically when only part of the remaining quota is available.
    • Count nonadmitted survivors as blocked, not discarded.
    • Release reservations for invalid events or failed durable writes.
    • Commit usage only for successfully persisted events.
  9. Fully parse and materialize survivors

    • Parse the complete stack trace and construct structured Error/StackFrame data on the server.
    • Apply privacy filtering, truncation, typed metadata normalization, and final validation.
    • Reuse the fingerprint parse result where practical.
    • If retaining parse state across the asynchronous route lookup is worthwhile, use a bounded batch-owned pooled representation containing offsets/compact frame structs, not per-frame object graphs.
    • If that complexity is not justified, compute the fingerprint first and reparse only survivors. Benchmark both approaches.
    • Return pooled memory on every success, rejection, cancellation, and exception path.
  10. Resolve/create stacks in bulk

    • Reuse known StackRoute ids.
    • Group true misses by unique signature.
    • Use distributed locking/single-flight only for true new-stack races.
    • Recheck after acquiring the lock.
    • Persist new stacks before events reference them.
    • Update the distributed route cache immediately after creation.
    • Avoid sequential repository reads and writes per event.
  11. Persist events in bulk

    • Generate deterministic storage ids from project plus client event id, or otherwise enforce create-only idempotency.
    • Group repository writes into the configured microbatch.
    • Treat duplicate/create-conflict results as successful duplicate acknowledgements.
    • Classify partial bulk failures precisely.
    • Retry only clearly transient operations within a small bounded policy.
    • On an ambiguous failure, return failure and let the client replay the segment. Do not enqueue after processing has started.
  12. Persist durable side-effect intent

    • Write deterministic, idempotent outbox records for required downstream work.
    • Notifications, webhooks, archive export, email, geolocation that is not required for grouping, and derived/repairable statistics run asynchronously.
    • Avoid per-event queue operations in the request. Aggregate or bulk outbox work where semantics allow.
    • The request may acknowledge once the event and required repair/outbox intent are durable.
    • Ensure retries cannot create duplicate user-visible notifications.
  13. Settle quota and acknowledge

    • Commit reservations for persisted events.
    • Release reservations for failures.
    • Aggregate received/persisted/discarded/duplicate/blocked/invalid counts.
    • Return 200 only after durable success for acknowledged events.

Discarded-stack fast path

Discard performance and correctness are first-class requirements.

Fingerprint before expansion

Extract canonical grouping logic from ErrorPlugin/ErrorSignature and stack assignment into a reusable service. The same signature calculation must drive:

  • V3 lightweight classification.
  • V3 full materialization.
  • V2 grouping when V2 is later adapted.
  • Regression tests using known existing event/stack fixtures.

For raw stack traces:

  • Build parser fixtures for the stack formats officially supported by Exceptionless clients.
  • Use span-based scanning and compact value representations.
  • Avoid substring creation for individual frames while fingerprinting.
  • Normalize unstable data such as line numbers/offsets according to existing grouping behavior.
  • Fall back to a documented normalized raw-trace signature when a format cannot be parsed.
  • Emit a metric for parser fallback so unsupported formats are visible.
  • Never log raw stack traces or sensitive event payloads.

Stack route cache

Introduce a lightweight cache entry rather than caching a full Stack only for routing:

(project_id, signature_hash) -> (stack_id, status, route_version)

Required behavior:

  • One distributed lookup per unique signature per microbatch at most.
  • Bulk cache get/set operations.
  • Bulk projected repository lookup for misses.
  • Cache keys include a version so routing schema changes can be rolled out safely.
  • Marking a stack discarded updates/invalidates the shared route entry before the status-change request returns.
  • Reopening a discarded stack updates/invalidates the shared route entry before the status-change request returns.
  • Stack creation writes the positive route entry immediately.
  • Stack removal/reset invalidates affected route entries.
  • Multi-instance tests prove status changes are observed without sticky sessions.

Start with the distributed cache as authoritative. Do not add an L1 positive discard cache until invalidation/version semantics are proven, because a stale positive entry would silently discard an event after a stack is reopened. If later benchmarks justify L1 caching, add versioned entries, message-bus invalidation, bounded size, TTL, and a scale-out correctness test.

An event already in flight when a status changes may use the status snapshot it observed. The system must not remain stale after the status-change operation completes.

Billing versus abuse protection

Maintain two separate concepts:

Protective ingress accounting

  • Counts all bytes, requests, and events, including discarded events.
  • Prevents unlimited free parsing/cache traffic.
  • Produces 429/413/503, not customer usage charges.

Billable usage accounting

  • Applies only after discard classification.
  • Counts only successfully persisted, non-discarded events.
  • Never counts discarded, invalid, duplicate, or failed events.
  • Uses atomic reservation/commit/release semantics across API instances.

Add explicit tests proving a customer at billable quota can still submit an event belonging to an already-discarded stack, while unique/new or active-stack events are blocked.

Modern .NET performance requirements

Use current .NET 10 platform features where they materially reduce work:

  • Minimal API route handlers.
  • HttpRequest.BodyReader and PipeReader-based JSON streaming.
  • JsonSerializer.DeserializeAsyncEnumerable with top-level values.
  • Source-generated JsonSerializerContext and JsonTypeInfo.
  • Typed, sealed request/response contracts.
  • CancellationToken propagation.
  • ArrayPool or MemoryPool only for measured hot buffers with strict ownership.
  • Span/ReadOnlySpan parsing within synchronous parser boundaries.
  • Frozen collections for static lookup tables where appropriate.
  • ValueTask only for hot operations that frequently complete synchronously.
  • Bounded concurrency and request-timeout policies.
  • Built-in request decompression where it preserves streaming and all decompressed-size limits remain enforced.
  • System.IO.Hashing may be benchmarked for transient internal keys, but persisted stack signatures must remain stable and collision-safe.

Hot-path rules:

  • No request-sized byte arrays, strings, MemoryStreams, JsonDocuments, or JsonNode trees.
  • No per-stage contexts.Where(...).ToList() allocation.
  • No per-event repository/cache call when a microbatch operation is possible.
  • No unbounded Channel, Task collection, queue, dictionary, or parser buffer.
  • Avoid LINQ in measured inner loops where it creates enumerators or collections.
  • Pre-size bounded collections from configured microbatch limits.
  • Avoid exception-driven normal control flow.
  • Aggregate logs and metrics per request/microbatch; do not log each discarded event.
  • Preserve safe encoding; do not use UnsafeRelaxedJsonEscaping.

Step-by-step implementation plan

Phase 0: Record baseline and freeze semantics

  • Document current V2 end-to-end flow and ownership boundaries.
  • Capture current behavior for authentication, default project resolution, suspension, retention, duplicate reference ids, manual stacking, fixed/regressed stacks, discarded stacks, sessions, bots, usage, notifications, and retries.
  • Capture representative real payload shapes without retaining sensitive production data.
  • Establish current throughput, p50/p95/p99 time-to-durable-event, CPU, allocated bytes/event, GC counts, object-storage operations, queue operations, and Elasticsearch requests.
  • Establish discarded-event cost separately.
  • Define benchmark hardware, Elasticsearch topology, compression, client concurrency, event-size distribution, and stack cardinality so results are repeatable.
  • Add architecture decision records for inline acknowledgement, bounded segments, idempotency, quota settlement, and side-effect outbox semantics.

Phase 1: Add benchmark and load-test harnesses

  • Add a microbenchmark project for serialization, fingerprinting, raw stack parsing, route grouping, and materialization.
  • Add an end-to-end load harness that can stream NDJSON with controllable event size, signature reuse, discard percentage, compression, concurrency, and connection failure.
  • Include scenarios for all-new stacks, one hot stack, many active stacks, all discarded, mixed discarded/active, malformed events, duplicate replays, slow clients, and slow/unavailable Elasticsearch.
  • Capture allocation profiles and identify LOH allocations.
  • Make benchmark output easy to compare in PR evidence; do not make noisy timing thresholds ordinary CI blockers.

Phase 2: Define V3 contracts and generated JSON metadata

  • Add small V3 request/response models in an ingestion-specific namespace.
  • Require a stable client event id and document idempotency rules.
  • Define field, collection, nesting, per-event, and segment limits.
  • Add the source-generated JsonSerializerContext.
  • Verify snake_case, nullable annotations, unknown-field policy, safe Unicode encoding, and extension-data limits.
  • Add serialization contract tests for one event, multiple top-level values, Unicode, malformed JSON, oversize fields, unknown fields, and unsupported top-level arrays.
  • Add a V3 HTTP sample under tests/http.

Phase 3: Map the Minimal API endpoint

  • Add EventIngestionV3Endpoints with MapGroup/MapPost mappings.
  • Require AuthorizationRoles.ClientPolicy.
  • Resolve explicit/default project once and reject credential mismatches.
  • Apply endpoint-specific concurrency, timeout, body, and content-encoding policies.
  • Ensure OverageMiddleware does not reject V3 chunked requests or perform premature billable checks.
  • Stream from BodyReader into bounded microbatches.
  • Map terminal results and ProblemDetails without MVC controller/filter dependencies.
  • Register the endpoint from Startup endpoint mapping.
  • Add a separate V3 OpenAPI document/baseline without changing the V2 document.
  • Add integration tests proving the endpoint is a Minimal API route and works without Content-Length.

Phase 4: Extract canonical stack fingerprinting

  • Move signature construction out of ErrorPlugin/AssignToStackAction into IStackFingerprintService.
  • Preserve existing signature fixtures and grouping behavior.
  • Implement raw stack-trace fingerprint parsers for supported formats.
  • Add fallback normalization/hash behavior.
  • Support manual stacking and project grouping settings.
  • Add parity tests showing equivalent V2 structured errors and V3 raw traces resolve to the same signature where expected.
  • Benchmark allocation-free fingerprinting and collision behavior.

Phase 5: Add bulk stack-route resolution and cache lifecycle

  • Define StackRoute and cache-key versioning.
  • Add a projected bulk repository method for signature misses.
  • Add microbatch-local deduplication.
  • Add distributed bulk cache get/set behavior.
  • Add short negative-cache behavior with immediate replacement on stack creation.
  • Update stack status-change, stack creation, deletion, reset, and migration paths to maintain route-cache correctness.
  • Add scale-out tests with two service providers/cache consumers.
  • Add metrics for local dedupe, distributed hit, negative hit, repository miss, and route-resolution latency.

Phase 6: Implement discarded-event early termination

  • Branch on StackStatus.Discarded immediately after route resolution.
  • Release raw/pooled working state.
  • Aggregate discarded counts per microbatch.
  • Verify no quota reservation occurs.
  • Verify no PersistentEvent/Error/StackFrame graph is materialized.
  • Verify no event, stack-stat, notification, webhook, archive, or outbox write occurs.
  • Verify all-discarded requests succeed when Elasticsearch event indexing is unavailable, provided authoritative route state can still be resolved safely.
  • Add status-race and mark/unmark regression tests.

Phase 7: Add scale-out-safe quota reservation and settlement

  • Define reservation, commit, and release operations in IIngestionQuotaService.
  • Implement atomic distributed accounting through Foundatio abstractions.
  • Preserve current hourly/monthly usage and notification semantics.
  • Apply quota only to survivors.
  • Handle partial remaining quota deterministically.
  • Release reservations on validation, cancellation, timeout, and persistence failure.
  • Commit only persisted, nonduplicate events.
  • Add concurrent multi-instance tests proving the plan limit cannot be over-reserved materially.
  • Keep protective rate limits separate and test that discarded traffic is still bounded operationally.

Phase 8: Implement full server-side error materialization

  • Build structured Error and StackFrame data from raw V3 stack traces.
  • Preserve exception type, message, nested/caused-by information where present, file, line, column, method, module, and native/async indicators where the source format provides them.
  • Apply target selection/grouping consistently.
  • Apply truncation and privacy rules before persistence.
  • Avoid duplicate parsing by retaining compact pooled parse state if benchmarks justify it.
  • Add parser fixtures for every officially supported client/runtime.
  • Add fuzz/property tests for malformed and adversarial stack traces.
  • Verify parser failure never crashes or stalls the stream; use the documented fallback.

Phase 9: Implement direct batch persistence and idempotency

  • Create IEventBatchWriter.
  • Resolve/create true new stacks with single-flight locking.
  • Replace per-stack sequential work with bulk/project-grouped operations.
  • Derive deterministic event storage identity from project plus client id.
  • Use create-only semantics and treat existing ids as duplicate success.
  • Bulk persist events with bounded retry for clearly transient failures.
  • Classify partial bulk results.
  • Do not acknowledge ambiguous failures.
  • Add retry/replay tests where the response is lost after persistence.
  • Add cancellation tests at every await boundary.
  • Confirm one slow request cannot monopolize all persistence concurrency.

Phase 10: Move side effects behind durable intent

  • Inventory every action after event persistence.
  • Classify each as required before acknowledgement, derived/repairable, or user-visible side effect.
  • Persist deterministic outbox/repair intent for asynchronous work.
  • Batch stack-usage updates where semantics permit.
  • Move notifications, webhooks, archival, and nonessential enrichment out of the request.
  • Preserve ordering/deduplication guarantees for user-visible effects.
  • Add reconciliation for partial event/outbox bulk outcomes.
  • Add tests proving client replay cannot duplicate notifications/webhooks.

Phase 11: Add observability and resilience

  • Add activities/spans for deserialize, fingerprint, route lookup, discard, quota, materialize, stack resolve/create, event bulk write, outbox write, and settlement.
  • Add counters for received, persisted, discarded, duplicate, blocked, invalid, parser fallback, retry, and failure.
  • Add histograms for compressed/decompressed/event/microbatch sizes and every stage duration.
  • Add gauges for active streams, in-flight events, rented buffer bytes, persistence concurrency, and circuit state.
  • Add resilience policies through IResiliencePolicyProvider.
  • Fail fast when durable storage is unavailable instead of holding streams indefinitely.
  • Verify logs contain identifiers and aggregate counts but not API keys, raw payloads, or stack traces.
  • Add dashboards/alerts for discard ratio, quota anomalies, route-cache miss spikes, parser fallback, p99 latency, and 429/503 rates.

Phase 12: Adapt V2 to shared improvements without taxing V3

  • Keep V2 routes, controller contract, 202 response, payload formats, and existing client behavior unchanged.
  • Adapt EventPostsJob to call the shared fingerprint, route resolution, materialization, quota, and batch writer services where this reduces work.
  • Do not make V3 instantiate V2 PersistentEvent graphs, plugin contexts, queue models, or compatibility converters.
  • Retain the V2 queue until a separate compatibility/removal decision is made.
  • Run serialization-audit coverage for V2 request/storage compatibility.
  • Compare V2 performance before/after and retain only measurable improvements.

Phase 13: Rollout

  • Add V3 enable/disable configuration and project/organization allowlisting.
  • Add explicit settings for microbatch count/bytes, segment limits, concurrency, timeout, decompressed size, and retry policy.
  • Start with local/integration environments.
  • Dogfood through the internal Exceptionless project.
  • Roll out to selected projects and compare V2/V3 stack signatures, discard decisions, usage, and stored event shape.
  • Increase traffic gradually while watching Elasticsearch saturation, cache latency, GC, 429/503, and cost per million events.
  • Publish client implementation guidance and at least one reference client.
  • Document rollback: disable V3 and have clients return to V2; no V2 contract change is required.

Test matrix

Contract and transport

  • Single NDJSON event.
  • Multiple top-level events split at every possible network-buffer boundary.
  • Chunked request without Content-Length.
  • gzip and Brotli streaming requests.
  • Unsupported encoding/media type.
  • Truncated JSON at every byte position.
  • Invalid UTF-8.
  • Deeply nested JSON.
  • Oversize individual field/event/segment.
  • Client cancellation during read and during persistence.
  • Slow-loris input and request timeout.
  • Empty stream.

Stack and discard behavior

  • Existing active stack.
  • Existing discarded stack.
  • Mixed active/discarded signatures.
  • Repeated discarded signature in one microbatch.
  • Mark active stack discarded while requests are in flight.
  • Reopen discarded stack.
  • New stack with a formerly negative cache entry.
  • Manual stacking.
  • Structured V2 versus raw V3 signature parity.
  • Unsupported raw stack fallback.
  • Hash/signature collision safeguards.

Billing

  • All persisted.
  • All discarded.
  • All duplicate.
  • Mixed persisted/discarded/duplicate/invalid.
  • At quota with a discarded event.
  • At quota with a new or active event.
  • Partial remaining quota.
  • Concurrent reservations from multiple API instances.
  • Persistence failure releases reservation.
  • Lost response/replay does not double charge.

Persistence and side effects

  • New-stack race across instances.
  • Bulk event partial failure.
  • Elasticsearch timeout/unavailability.
  • Cache timeout/unavailability.
  • Response lost after event persistence.
  • Duplicate replay.
  • Outbox partial failure and reconciliation.
  • Notifications/webhooks emitted once.
  • Stack usage/stat repair.

Performance and scale

  • Long stream demonstrates bounded memory.
  • 0%, 10%, 50%, 90%, and 100% discarded traffic.
  • One hot signature and high-cardinality signatures.
  • Small, median, large, and maximum events.
  • Compressed and uncompressed traffic.
  • 1, 2, 4, and 8 API instances.
  • Increasing client concurrency until Elasticsearch saturation.
  • Slow Elasticsearch produces backpressure/429/503 without unbounded memory.

Performance gates

Record exact baselines in Phase 0 and attach benchmark evidence to implementation PRs. Initial gates:

  • No request-sized buffering or giant top-level array materialization.
  • No LOH allocation caused by normal request framing or microbatch list growth.
  • Working memory remains bounded by configured in-flight requests and microbatch limits during a long stream.
  • Discarded events perform zero event-index writes and zero full Error/StackFrame materialization.
  • One route lookup per unique signature per microbatch at most.
  • Normal V3 ingestion performs no object-storage payload write/read and no primary ingestion queue operation.
  • Horizontal scale from one to four API instances shows near-linear API throughput until the shared datastore becomes the measured bottleneck; target at least 80% scaling efficiency.
  • V3 reduces allocated bytes/event, CPU/event, and time-to-durable-event versus V2. Set numeric reduction targets after the repeatable V2 baseline is captured rather than choosing ungrounded numbers.
  • Cost per million persisted events and per million discarded events is measured and reported.
  • No correctness regression is accepted in exchange for a microbenchmark improvement.

Acceptance criteria

  • V3 event ingestion is implemented entirely with Minimal APIs.
  • Clients stream individual JSON objects and never need to build a JSON array.
  • The server parses raw stack traces into structured errors.
  • Request processing is bounded and allocation-conscious.
  • The normal V3 path processes inline without the EventPost object-storage/queue handoff.
  • A success response means the primary event data is durable or the event reached another documented terminal success such as discarded/duplicate.
  • Stable event ids make complete-segment replay idempotent.
  • Discarded stacks are identified before billable quota reservation and full materialization.
  • Discarded events are not charged.
  • Discarded traffic remains subject to protective ingress limits.
  • Marking and reopening discarded stacks behaves correctly across API instances.
  • Active/new events are bulk persisted; secondary effects are durable and asynchronous.
  • Elasticsearch failure produces bounded backpressure and retryable failures, not unbounded memory or an ambiguous server queue fallback.
  • V2 remains contract-compatible.
  • V3 OpenAPI, HTTP samples, client guidance, unit tests, integration tests, failure tests, and performance evidence are included.
  • Rollout can be enabled gradually and disabled without changing V2.

Follow-up decision trigger

If the product must acknowledge events while Elasticsearch is unavailable, create a separate design issue for an explicit durable buffered mode. Do not silently reintroduce the current queue into the inline V3 path.

Activity

  1. ejsmith commented on Jul 13, 2026

    @ejsmith
    MemberAuthor

    Implementation progress

    Implementation has started on branch feature/v3-event-ingestion, created from current origin/main at f02e0bf3b.

    Current slice:

    • Audit the V2 ingestion, serialization, quota, stack-routing, persistence, and side-effect boundaries against this issue.
    • Establish focused benchmark/load-test scaffolding and freeze V2 compatibility fixtures.
    • Implement the V3 source-generated contracts and Minimal API streaming boundary as the first reviewable commit.

    The implementation will follow the ordered phases in this issue. I will post checkpoints after each verified commit/slice, including tests and benchmark evidence.

  2. ejsmith commented on Jul 13, 2026

    @ejsmith
    MemberAuthor

    Phase 0 checkpoint

    Committed locally on feature/v3-event-ingestion:

    • 74e254c84 — Document V3 ingestion architecture

    Added docs/event-ingestion-v3-architecture.md covering:

    • The current V2 object-storage/queue/worker path and compatibility boundary.
    • Inline acknowledgement, bounded stream segments, and deterministic idempotency decisions.
    • The discarded-stack early-termination and distributed route-cache invariants.
    • Scale-out-safe billable reservation/settlement semantics.
    • Side-effect durability, baseline methodology, performance gates, and completion evidence.

    Next: Phase 1 benchmark harness and Phase 2 source-generated V3 contracts.

  3. ejsmith commented on Jul 13, 2026

    @ejsmith
    MemberAuthor

    Phases 1–2 checkpoint

    Committed locally:

    • 0f5bb6837 — Add V3 ingestion contracts and benchmarks

    Implemented:

    • A BenchmarkDotNet 0.15.8 project in the solution with array-versus-top-level-value ingestion benchmarks for 1, 100, and 1000 events.
    • Compact V3 event/user/request/environment contracts with explicit limits.
    • Source-generated System.Text.Json metadata using snake_case, safe encoding, nullable enforcement, and bounded JSON depth.
    • Focused tests for multiple top-level values, required properties, additive unknown fields, top-level-array rejection, and safe encoding.

    Verified:

    • dotnet build benchmarks/Exceptionless.Benchmarks/Exceptionless.Benchmarks.csproj -c Release --no-restore -v:minimal
    • BenchmarkDotNet dry job: 6/6 benchmark cases executed.
    • Focused V3 serialization tests: 5/5 passed.
    • NuGet vulnerability audit: no known vulnerable packages in the benchmark project.

    Next: Phase 3 Minimal API routes, chunked NDJSON handling, request limits, authentication/project resolution, and integration coverage.

  4. ejsmith commented on Jul 13, 2026

    @ejsmith
    MemberAuthor

    Implementation is complete on feature/v3-event-ingestion and the branch is pushed.

    Completed:

    • Added controller-free Minimal API routes at POST /api/v3/events and POST /api/v3/projects/{projectId}/events.
    • Streams top-level JSON values directly from HttpRequest.BodyReader with source-generated System.Text.Json metadata; no request-sized array or buffering.
    • Added bounded count/byte microbatches, adaptive concurrency limits, request timeouts, gzip/Brotli decompression, compressed/decompressed/event/request limits, cancellation, and ProblemDetails responses.
    • Simplified error clients to send exception_type plus raw stack_trace; server parsing covers .NET, Java, JavaScript, Python, caused-by chains, and a normalized fallback.
    • Added manual stacking support and V2/V3 fingerprint parity coverage.
    • Added distributed bulk route lookup with positive/negative caching and tested create/discard/reopen/delete propagation across independent consumers.
    • Discarded-stack detection happens before quota reservation, event materialization, and persistence. Discarded events are tracked but not charged.
    • Added atomic distributed quota reservations with deterministic partial admission and release/commit recovery.
    • Persists admitted events inline with deterministic IDs, bulk writes, reconciliation on ambiguous failures, and a durable async side-effect work item.
    • Added per-event idempotency markers for stack statistics, notifications/webhooks, and post-save plugins; replay tests prove no second event write or queued work item.
    • Added isolated V3 OpenAPI output, HTTP samples, observability, disabled-by-default rollout/allowlist settings, BenchmarkDotNet suites, and a repeatable streaming load harness.
    • Preserved the V2 endpoint/pipeline and bypassed V2 overage middleware only for V3.
    • Pinned Microsoft.OpenApi 2.7.5 to remediate CVE-2026-49451 discovered during the final dependency audit.

    Verification:

    • dotnet build Exceptionless.slnx --no-restore --disable-build-servers -m:1: 0 warnings, 0 errors.
    • Full backend suite: 2,350 passed, 0 failed, 2 intentional performance fixtures skipped.
    • Focused V3 persistence/replay/outbox suite: 26 passed.
    • BenchmarkDotNet dry run: all 10 ingestion benchmarks executed successfully.
    • NuGet vulnerability audit: no vulnerable packages across all eight projects.
    • git diff --check: clean.

    Commits:

    • 74e254c84 Document V3 ingestion architecture
    • 0f5bb6837 Add V3 ingestion contracts and benchmarks
    • bfa50dba1 Implement V3 ingestion processing pipeline
    • f45909066 Expose streaming V3 ingestion Minimal API
    • 78ac6df06 Add V3 ingestion load and rollout tooling

    Opening a draft PR next for review and staged performance validation.

  5. ejsmith commented on Jul 13, 2026

    @ejsmith
    MemberAuthor

    Draft PR opened for review and staged load validation: #2370

  6. ejsmith commented on Jul 13, 2026

    @ejsmith
    MemberAuthor

    CI is green on PR #2370: API build/tests/coverage, client lint/check/build/unit/integration, Aspire/Playwright E2E, Docker builds, Website, and versioning all completed successfully.

  7. ejsmith commented on Oct 7, 2026

    @ejsmith
    MemberAuthor

    Revised plan: streaming ingestion on /api/v2/events

    Review of #2370 and research into comparable products (Sentry, Rollbar, Bugsnag, Honeybadger, Raygun, Airbrake, Datadog, OTLP, Elastic APM, Splunk HEC, Segment, PostHog, Honeycomb, Application Insights, Loki, Seq) changed the direction:

    • Every one of them acknowledges after a cheap durable (or in-memory) enqueue. None waits for full processing. Inline processing would tie every client app's requests to Elasticsearch and pipeline health, and give up the queue as our load leveler.
    • No new endpoint version is needed. New encodings of the same events are added to an existing endpoint through Content-Type (Datadog, OTLP, Seq). New endpoints appear only when the wire format changes incompatibly.
    • Streaming means bounded requests, not long-lived uploads. OpenTelemetry rejected streaming, and Elastic APM agents rotate every 10s/750 KB. Proxies, load balancer timeouts, and Kestrel's minimum data rate all work against long uploads.

    #2370's inline V3 endpoint is closed. Its stream reader, concurrency limiter, and load harness are carried into the PRs below.

    Goals

    • Clients: SDKs write events as they occur, on an open NDJSON stream or one per request. They never build arrays or block the host app, and requests always return quickly.
    • Server: never hold a whole post in memory, at accept or in processing. Acknowledge after a durable enqueue and process asynchronously. Cut per-post storage and Redis cost.
    • Compatibility: existing SDKs and payloads keep working and benefit automatically.

    Wire contract

    All of this is added to POST /api/v2/events and /api/v2/projects/{id}/events; nothing existing changes.

    • Body: V2 event JSON as a single object, a JSON array, or application/x-ndjson (new). text/plain and the v1 API keep the legacy path.
    • Encoding and size: chunked or Content-Length, optional gzip or br, UTF-8 on the new path.
      • Per event: about 512 KB.
      • Per request: about 10 MB compressed and 50 MB decompressed, with a 60s maximum duration.
    • Event id: optional. When present, resends are deduplicated for 24 hours.
    Status Meaning
    202 Every event in the request is durably queued. Body (additive): {received, accepted, invalid, errors[{index, message}]}, where errors cover framing only.
    400 The body can't be split into events. Events before the error are queued.
    401/403/404 Stop.
    402 Over the plan limit; pause.
    413 Too large.
    429/503 Resend after Retry-After, with backoff and jitter.

    Client protocol:

    • Write one NDJSON line per event on an open request.
    • End the request at 30s or 1 MB, then reuse the connection for the next one.
    • Keep unacknowledged events for resending; event ids make resends safe.
    • Clients that can't stream (browsers, minimal senders) POST each event, or a few, with keep-alive.
    • Always send in the background:
      • A bounded queue that drops the oldest events when full.
      • A 10s timeout.
      • About a 2s flush deadline on shutdown.

    Server design

    Accept path. Frames events from the request stream and never deserializes the whole post.

    • Run cheap checks before reading the body: auth, whether submission is enabled, suspension, and events left.
    • Split the stream into events by JSON structure: check that each is an object, enforce the per-event cap, and skip oversized events without retaining them.
    • Group events into microbatches of 100 events, 256 KB, or 1s, whichever comes first.
    • Put each microbatch's compressed NDJSON directly into the queue entry. Fall back to blob storage for large microbatches or a high backlog.
    • Return 503 when the enqueue fails. Today that case silently returns 202.
    • For streams, turn off Kestrel's minimum body data rate and limit open streams per organization and globally.

    Job (EventPostsJob). Inline entries need no blob read. Blob payloads are decompressed and parsed as streams with limits, which closes the decompression-bomb risk. Failed events are requeued inline instead of being re-uploaded. Leftover dead-letter payloads are cleaned up.

    Pipeline:

    • Event-id idempotency: an id is pending while processing, stored after SaveEventAction succeeds, and released on failure. The same fix applies to reference-id dedupe, which today marks events before saving them.
    • The plan limit is applied after stack lookup, so discarded-stack events are never counted as blocked.
    • Raw @simple_error stack traces are parsed on the server into structured @error values and stacked by the normal error algorithm, so there is one stacking algorithm. Existing simple-error stacks may regroup once, and that is acceptable.

    Decisions

    1. Ingest requests are no longer counted by the per-request throttle (3,500 per organization per 15 minutes). Ingestion is bounded by the plan quota, per-organization stream and processing limits, and an events/s limit per organization.
    2. Server-side parsing replaces the simple-error signature; there is one algorithm.
    3. JSON arrays move to the new accept path behind a flag once NDJSON is proven.
    4. SDKs enable NDJSON only when the server advertises support through a capability header. Older self-hosted servers return 202 and then drop NDJSON, so SDKs can't fall back after the fact.
    5. Add V3 streaming event ingestion endpoint #2370 is closed; its branch is kept for reference.

    Work breakdown

    # PR Scope Done when
    0 Baseline Load harness on main; measure V2 with single-event posts, 50-event arrays, and Elasticsearch slow or down Accept p50/p99, CPU and allocations per 1k events, and blob/Redis ops per post recorded
    1 Correctness and signals 503 + Retry-After on failed enqueue and disabled submission; Retry-After on throttle 429; streaming decompression with a limit in the job; dead-letter payload cleanup No 202 without a durable enqueue; decompression-bomb test passes
    2 Streaming accept + NDJSON Boundary reader at accept, microbatching, inline queue entries with blob fallback, limits and stream limiter, ingest throttle change, capability header Chunked and slow streams, rotation, and cut streams tested; array and NDJSON parity; accept p99 and blob ops beat the baseline
    3 Arrays on the new path (flagged) Framed JSON arrays; streamed blob payloads in the job; per-event retry without re-upload Existing SDK payloads produce identical events, stacks, and usage on both paths
    4 Event-id idempotency Pending, stored, and released lifecycle; fix reference-id dedupe ordering Retry after failure is processed, retry after success is deduplicated, in-flight duplicates are not lost
    5 Discard before quota Plan limit after stack lookup, without creating stacks for blocked events Discarded events never counted as blocked
    6 Server-side stack parsing @simple_error to @error (.NET, Java, JS, Python, fallback) with golden fixtures from real SDK output Fixture parity with SDK-produced structured errors
    7 Protocol and conformance Client protocol doc, status matrix, JSON Schema, golden fixtures, reference senders (curl, Python, Go) A sender built from the docs passes the fixtures
    C1 .NET SDK Streaming transport with rotation, keep-alive, 10s timeout, ids, fix for the broken batch shrink after 413, async shutdown flush with a deadline, capability check Never blocks the app; bounded memory during outages
    C2 JS SDK Node streaming; browser fetch with keepalive/sendBeacon; timeouts; fix the queue stall on a hung request and the one-batch-per-tick throughput limit; SIGTERM flush; capability check Same as C1
    I1 Infra API App Gateway backend timeout of at least 90s; confirm streamed bodies aren't buffered A 30s rotating stream works in staging

    Order: 0 → 1 → 2, then 3, 4, and 5 in parallel, then 6. Protocol docs (7) and SDK work (C1, C2) follow 2. I1 must land before 2 is enabled in production.

    Rollout

    • Feature flags: NDJSON accept, arrays on the new path, inline payload thresholds, throttle mode.
    • Metrics:
      • Accept latency.
      • Events per request.
      • Queue depth and age.
      • Redis memory.
      • Blob operations per second.
      • Job memory and CPU.
      • Dedupe hits.
      • Framing errors.
    • Stages: internal project, then a percentage of projects, then everyone. Every flag can be turned back off.
    • Chaos checks:
      • Elasticsearch down: clients keep getting 202 and the backlog drains.
      • Redis down: 503 with Retry-After.
      • Cut streams: no events lost when ids are used.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions