You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
Build an OAuth 2.1 authorization server (OpenIddict) so MCP clients can authenticate #788
The design is finished and merged. Do not redesign it. Read docs/plans/788-mcp-oauth/
first, all five files, in order. It records what was decided, what was rejected and why, and four
claims that were checked against source. Several conclusions are counter-intuitive and came from
verification rather than reasoning. Re-deriving them from scratch will cost a day and may reach the
wrong answer, as it did twice during the design.
The slice that matters is #796. It produces nothing visible, and an error there is silent and
security-relevant. Adversarial review found two of its requirements after the design looked
complete: stale role claims, and revocation not being a property of the token format. Read its
comments, not just its body.
Three traps, all documented, all easy to hit:
Adding the 4 OpenIddict tables turns TenantBypassDiscoveryTests.DiscoveredSurface_Floor red.
That is correct behaviour, because it asserts exact set equality. Add the entities with a stated
reason. Do not relax it to a subset check, which would quietly disable a tenancy guard.
Roles are not read live. Nothing reloads them from the database. Freshness comes from
revocation. An OAuth token with stale role claims keeps a demoted user's authority.
Reference tokens do not give instant revocation by themselves. OpenIddict does not check
authorization-grant status by default.
Important
Amended. This issue's question is answered. The body below stays as written, for history. It
poses a question and leans toward the cheaper options. The decision went the other way, to build an
OAuth 2.1 authorization server with OpenIddict. See the decision comment
for what was chosen, what was designed and rejected (a full Personal Access Token flow), the verified
integration cost, and the one guard test that will break. This issue is now the OAuth implementation.
Slices: engine first, screens after
Dependencies set the order. Risky plumbing lands first, where it is cheapest to fix. Slices 1-3 are
invisible to users, by design. The sizes total four L and two M.
Wave 1, foundation. Nothing else starts until this lands.
The MCP SDK's discovery metadata (McpAuthenticationHandler) and scope enforcement on the MCP tools are not part of this issue. They live in epic #789, because they belong to MCP rather than to the OAuth server.
#770's design work raised this issue (see docs/plans/770-mcp-server/01-design.md, "Open questions and risks"). This open question decides whether MCP support is usable in the real world, as distinct from whether it is correct.
The mismatch
#770's design authenticates an MCP session with an ordinary Cluckwork access token. It is the same bearer the SPA carries, from POST /api/v1/auth/login with a farm code. That is the smallest correct first step. It reuses the whole existing pipeline (tenant resolution, credential epoch #364, authorization, must-change-password) with no MCP-specific identity and no new config key.
It works. It may not be usable:
Cluckwork's access token is short-lived by design, and it rotates through a refresh token in an HttpOnly cookie. The cookie exists to keep the refresh token out of reach of JavaScript in a browser.
An MCP client is not a browser. It has no cookie jar semantics to rely on, and the MCP ecosystem expects OAuth 2.1 with discovery. A client calls the server, the server points it at an authorization server, and the client obtains and refreshes its own token.
So a user configuring a Cluckwork MCP server would paste a bearer that stops working shortly afterwards, with no in-protocol way to renew it.
Options, roughly in increasing cost
Ship phase 1 as-is, documented as requiring a caller that can inject a fresh bearer (a script, a gateway, a developer testing). This is honest, cheap, and possibly fine for the first slice, but it is not "point your assistant at your farm".
Add the SDK's McpAuthenticationHandler and serve protected-resource metadata at /.well-known/oauth-protected-resource/.... The SDK has first-class support for this (ModelContextProtocol.AspNetCore.Authentication). This tells a client where to authenticate. By itself, it does not make Cluckwork an authorization server.
Become an OAuth 2.1 authorization server for MCP clients. This is the largest option, and probably not phase 1.
Why this needs a decision before the tool work, not after
The design deliberately keeps this choice from invalidating it. The identity bridge reads the ClaimsPrincipal the pipeline produces, regardless of how that principal was authenticated. So the choice does not affect the tools or the guards.
But the phasing depends on the answer. If option 1 is acceptable for a first slice, MCP support can ship and be exercised. If it is not, then whatever closes this gap is a prerequisite for MCP being worth shipping at all. That work should be sized alongside MCP rather than discovered afterwards.
Not claimed
This issue claims no security defect. Option 1 is not insecure. A short-lived bearer over an authenticated endpoint is exactly what the SPA does. The problem is ergonomics and reach. The risk is finishing the tool work and finding that nobody can practically connect.
Note
Picking this up? Start here.
The design is finished and merged. Do not redesign it. Read
docs/plans/788-mcp-oauth/first, all five files, in order. It records what was decided, what was rejected and why, and four
claims that were checked against source. Several conclusions are counter-intuitive and came from
verification rather than reasoning. Re-deriving them from scratch will cost a day and may reach the
wrong answer, as it did twice during the design.
Build in this order: #794 (prerequisite) → #795 → #796 → #797 → #798 → #799. #800 can run in
parallel after #795.
The slice that matters is #796. It produces nothing visible, and an error there is silent and
security-relevant. Adversarial review found two of its requirements after the design looked
complete: stale role claims, and revocation not being a property of the token format. Read its
comments, not just its body.
Three traps, all documented, all easy to hit:
TenantBypassDiscoveryTests.DiscoveredSurface_Floorred.That is correct behaviour, because it asserts exact set equality. Add the entities with a stated
reason. Do not relax it to a subset check, which would quietly disable a tenancy guard.
revocation. An OAuth token with stale role claims keeps a demoted user's authority.
authorization-grant status by default.
Important
Amended. This issue's question is answered. The body below stays as written, for history. It
poses a question and leans toward the cheaper options. The decision went the other way, to build an
OAuth 2.1 authorization server with OpenIddict. See the decision comment
for what was chosen, what was designed and rejected (a full Personal Access Token flow), the verified
integration cost, and the one guard test that will break. This issue is now the OAuth implementation.
Slices: engine first, screens after
Dependencies set the order. Risky plumbing lands first, where it is cheapest to fix. Slices 1-3 are
invisible to users, by design. The sizes total four L and two M.
Wave 1, foundation. Nothing else starts until this lands.
Its prerequisite is Data Protection has no persisted key ring, so Identity's token providers break across restarts and replicas #794 (Data Protection key ring).
Wave 2. These two can run in parallel, and both need #795.
must-change-password, flock scoping, rate limits and revocation. This slice carries the
security work, and it produces nothing visible, so it is the most likely slice to be under-scoped.
unapproved apps expire.
Wave 3, the screens.
open-redirect guard on the hop from login to consent. Needs OAuth slice 1: stand up OpenIddict — tables, migration, token issuance #795, OAuth slice 3: dynamic client registration with rate limiting and unapproved-app expiry #797.
Owner's farm-wide view. Needs OAuth slice 1: stand up OpenIddict — tables, migration, token issuance #795, OAuth slice 4: consent screen, with step-up and return-URL validation #798.
In parallel with waves 2-3.
Needs OAuth slice 1: stand up OpenIddict — tables, migration, token issuance #795 only.
The MCP SDK's discovery metadata (
McpAuthenticationHandler) and scope enforcement on the MCP tools are not part of this issue. They live in epic #789, because they belong to MCP rather than to the OAuth server.#770's design work raised this issue (see
docs/plans/770-mcp-server/01-design.md, "Open questions and risks"). This open question decides whether MCP support is usable in the real world, as distinct from whether it is correct.The mismatch
#770's design authenticates an MCP session with an ordinary Cluckwork access token. It is the same bearer the SPA carries, from
POST /api/v1/auth/loginwith a farm code. That is the smallest correct first step. It reuses the whole existing pipeline (tenant resolution, credential epoch #364, authorization, must-change-password) with no MCP-specific identity and no new config key.It works. It may not be usable:
Options, roughly in increasing cost
McpAuthenticationHandlerand serve protected-resource metadata at/.well-known/oauth-protected-resource/.... The SDK has first-class support for this (ModelContextProtocol.AspNetCore.Authentication). This tells a client where to authenticate. By itself, it does not make Cluckwork an authorization server.Why this needs a decision before the tool work, not after
The design deliberately keeps this choice from invalidating it. The identity bridge reads the
ClaimsPrincipalthe pipeline produces, regardless of how that principal was authenticated. So the choice does not affect the tools or the guards.But the phasing depends on the answer. If option 1 is acceptable for a first slice, MCP support can ship and be exercised. If it is not, then whatever closes this gap is a prerequisite for MCP being worth shipping at all. That work should be sized alongside MCP rather than discovered afterwards.
Not claimed
This issue claims no security defect. Option 1 is not insecure. A short-lived bearer over an authenticated endpoint is exactly what the SPA does. The problem is ergonomics and reach. The risk is finishing the tool work and finding that nobody can practically connect.