Skip to content

Background worker has no single-runner guarantee — double-runs if scaled >1 instance #271

Description

@mforce

Coding brief — read first. Start with epic #530 (the canonical decision record); this is slice S1 of that epic. This body is complete and current as of 2026-08-16, folding every review correction — no comments to cross-reference. Repo conventions in AGENTS.md apply. Definition of done = the acceptance criteria at the end. Originally filed under the #244 deployment-readiness audit; now a Phase 1.6 scale-out blocker.

Problem

The background worker has no single-runner guarantee, so it silently double-runs if the app is ever scaled past one instance. Every API instance registers DurableJobWorker as a hosted service; the poll selects Pending jobs with no row claim — no FOR UPDATE SKIP LOCKED, no lease (DurableJobWorker.cs) — and the three recurring sweeps (DailyEntryLockSweep, RefreshTokenPurgeSweep, IdempotencyRecordPurgeSweep) ride the same poll and run unconditionally per instance. With one instance this is fine; two instances each process the same jobs and both sweep.

The contract is "at most one active leader", NOT "exactly once"

This is the correction from the epic #530 review, and it is load-bearing — do not write "exactly once" into the finish line. A Postgres advisory lock provides mutual exclusion only: a session lock is released when the session dies, so a leader can lose the lock while its work is still running, and a process can crash after an external effect but before marking a job complete. Exactly-once is therefore unachievable at this layer. The achievable, correct contract:

  • at most one active leader under ordinary two-replica operation;
  • durable jobs are at-least-once; handlers are idempotent (this is the property that makes at-least-once safe);
  • recurring schedules have durable next_run/claim state, not "whoever holds the lock right now may run";
  • jobs claimed with leases/fencing or FOR UPDATE SKIP LOCKED, with crash recovery.

Scope

Move the worker and the three sweeps behind a Postgres advisory-lock lease (the lease is deliberately Postgres, not one of #543's Redis ports — a Redis lock's failure mode is precisely the double-execution the lease prevents; epic decision 13). Add crash recovery so a dead leader's lease is reacquirable.

Note: the two purge sweeps are already idempotent deletes; DailyEntryLockSweep reads-then-writes under an optimistic Version token, so a losing replica currently throws a concurrency exception its per-account catch logs as a DB fault. The lease removes the double-run rather than relying on that catch.

Sequencing

Slice S1 of epic #530, after #543 (S0). The lease mechanism itself does not consume #543's ports, but the scale-out slices land in order and the epic gates a second replica on all four (#271, #338, #544, #545) closing.

Acceptance criteria

Verify

dotnet test Cluckwork.sln

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:apiAPI/endpoint layerbugSomething isn't workingepic-1.6Phase 1.6 — Multi-farm tenancysliceThin vertical work item

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions