Skip to content

bug(ingest): dedupe marks the event id before the NATS publish — a failed publish + client retry permanently drops the event #384

Description

@taitelee

Summary

processRecord durably records the dedupe id before publishing to NATS, so any publish failure followed by the sanctioned client retry silently loses the event forever.

Detail

internal/api/ingest.go:

  • :419 h.Dedup.CheckAndMark(ctx, eventID) — durably writes the id to Pebble (internal/dedupe/embedded.go:44, pebble.Sync, survives restart).
  • :450 h.Publisher.Publish(...) runs after. On "maximum bytes exceeded" it returns 503 + Retry-After: 30 (the documented backpressure contract, design decision ci: bump actions/checkout from 4 to 6 #4); on marshal/other failure a 500.

When the client retries as instructed, the retry hits dup == true at :424 and is skipped. The event is never in NATS, never in ClickHouse, never in the DLQ. There is no unmark/rollback. Same for a mid-batch client abort after some records were marked.

Impact

Silent, permanent data loss triggered by exactly the documented "NATS full → retry" path.

Fix direction

Mark only after a successful publish (mark-after-publish), or make the mark provisional and confirm on publish success / roll back on failure.

Found in a repo-wide audit; independently confirmed by two auditors. Distinct from #220/#221/#222/#370.

Activity

  1. added
    bugSomething isn't working
    area/ingestIngest pipeline (Bento, batching, DLQ)
    area/dedupeDeduplication (Pebble, ScyllaDB)
    on Jul 8, 2026
  2. coderabbitai commented on Jul 8, 2026

    @coderabbitai
    🔗 Related PRs

    #122 - fix(ingest): break infinite-Nak loop on permanent delete errors [merged]
    #182 - refactor: full api --> ingest --> clickhouse --> dlq refactor [closed]
    #364 - feat(dedupe): observe or reject ingest rows missing the id_field [merged]


    🧪 Issue enrichment is currently in open beta.

    You can configure auto-planning by selecting labels in the issue_enrichment configuration.

    To disable automatic issue enrichment, add the following to your .coderabbit.yaml:

    issue_enrichment:
      auto_enrich:
        enabled: false

    💬 Have feedback or questions? Drop into our discord!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/dedupeDeduplication (Pebble, ScyllaDB)area/ingestIngest pipeline (Bento, batching, DLQ)breaking-changeBreaking change to public API, CLI, or configbugSomething isn't working

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions