Skip to content

Distributed IP-keyed auth rate limiters with in-process fallback #544

Description

@mforce

Coding brief — read first. Start with epic #530 (the canonical decision record). This body is complete and current as of 2026-08-16 — no comments you must cross-reference. Repo conventions in AGENTS.md apply automatically. Definition of done = the Verify section at the end.

Slice S3 of epic #530 (Phase 1.6 — Multi-farm tenancy). Depends on S0. One of the four blockers that must close before a second replica is permitted.

Problem

AddRateLimiter's partitions live inside each process, so N replicas allow roughly N times the intended attempts per IP before lockout. This covers login, refresh and the client-error report endpoint alike. It is listed in AGENTS.md as a scaling blocker and has never had an issue of its own.

Scope

Move the IP-keyed auth limiters onto the increment-plus-expiry port from S0.

Keying stays IP, not account — these endpoints are pre-authentication, so there is no account yet, and the farm code supplied on a login attempt is attacker-controlled and must never become a limiter key or a metric label.

Failure policy

When the shared store is unreachable, fall back to today's in-process limiter and raise an alarm. Bounded degradation (N replicas × the intended budget) rather than either extreme: fail-open is unlimited, and because an attempt against an unknown user still pays the dummy PBKDF2 cost, unlimited means CPU exhaustion, not merely unlimited guesses. Fail-closed would turn a store blip into nobody being able to log in.

The alarm is the load-bearing half. A limiter that has silently been in fallback for three weeks is the worst of both.

Interaction with the unknown-farm response (#530 decision 6)

/auth/login returns a distinct response for an unknown farm code, which makes it a farm-enumeration endpoint by design. That branch is cheaper than the password branch — it skips PBKDF2 — so it is the one an attacker would flood.

Therefore: the limiter is applied before the account lookup, not after, and the attempted slug is never used as a metric label or an unbounded log dimension.

Tests

  • Two instances sharing one store enforce one combined budget, not two.
  • Store unreachable → falls back to in-process, alarm raised, requests still served.
  • Store recovers → the shared budget resumes without a restart.
  • The unknown-farm branch is limited before any account lookup.
  • No metric or log dimension carries an attacker-supplied slug.

Verify

dotnet test Cluckwork.sln

Plus a two-instance run asserting the combined budget.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:apiAPI/endpoint layerepic-1.6Phase 1.6 — Multi-farm tenancysliceThin vertical work item

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions