The adversarial review of the MCP design (#770, docs/plans/770-mcp-server/05-adversarial-review.md) raised this and flagged it as unexamined. No design document, guard or issue covers it. It is a prerequisite in spirit for #789 (EPIC: MCP server support), though not a hard blocker on any single slice.
The problem
Every MCP tool result becomes prompt content in a model's context window. Some of that content is free text that a farm's own users typed. A model reading it cannot reliably distinguish "data I was asked to summarize" from "instructions addressed to me".
Concrete fields already in the phase-1 tool surface:
| Field |
Source |
Customer.Name, Phone, Email, Address, Note |
src/Cluckwork.Domain/Sales/Customer.cs:16-20 |
SalesOrder.DiscountReasonNote |
src/Cluckwork.Domain/Sales/SalesOrder.cs:28 |
Flock.Name, Breed |
src/Cluckwork.Domain/Flocks/Flock.cs:12-13 |
Customer.Note and DiscountReasonNote are genuinely free-form.
Why this is more than the generic prompt-injection worry
There is a privilege gradient, and it runs the wrong way.
- Writing a customer record, including
Note, requires AuthPolicies.SalesFlow (CustomerEndpoints.cs:18-39). SalesFlow is everyone but ReadOnly, including a plain Worker.
- The MCP receivables and orders tools are gated at
AdminOnly / SalesAccess.
So a Worker can write free text that later lands in an Owner's model context. If that text reads as an instruction and the assistant acts on it, the action executes with the Owner's privileges, not the Worker's. A low-privilege write escalates through a high-privilege read.
A second path needs no insider at all. A customer supplies their own name or address, a clerk enters it verbatim, and it reaches the same place.
What bounds the damage today
This is worth stating so the risk is not over-rated:
- Every tool call is independently authorized against the caller's token. An injected instruction cannot make the model call a tool the caller's role cannot reach.
AddAuthorizationFilters() removes it from tools/list and refuses the call.
- The phase-1 surface has exactly one write (record daily entry). The reachable harm is therefore narrow today.
- Most MCP clients require user confirmation for writes.
The caller's own RBAC tier bounds the blast radius. Within that tier nothing bounds it, and the tier of someone running an assistant over the books is usually Owner or Manager.
Why decide it now rather than after
The cheap mitigations are shape decisions, and shape is expensive to change once tools exist and clients depend on the schema:
- Whether tool results are structured JSON with typed fields rather than prose sentences.
- Whether free-text fields are returned at all by default, or only on explicit request, or truncated.
- Whether user-supplied text is delimited or fenced, so a model has a syntactic cue about provenance.
- Whether the write tool's arguments may ever be derived from text a read tool returned, or must be caller-supplied.
- Whether
McpServerTool's ReadOnly/destructive annotations are set correctly, so clients can apply their own confirmation policy.
Retrofitting any of these means changing published tool schemas.
What this issue is NOT asking for
This issue does not ask for a solution to prompt injection. Nobody has solved that industry-wide, and Cluckwork will not solve it. It asks for a recorded decision about presentation and boundaries, at the strength the argument supports, so the tool surface is designed with the hazard in view rather than discovering it after a schema ships.
An acceptable outcome is an explicit accepted-risk decision record saying the bounded blast radius is tolerable for phase 1, and why. Per AGENTS.md, an accepted-risk record with No incident is a legitimate result, and it is load-bearing.
Suggested first step
Enumerate every free-text field reachable from the phase-1 surface. The table above is a start, not a complete walk. Derive the full list from the model, per the repo's "walk everything, exclude deliberately" rule. Then decide presentation per field.
The adversarial review of the MCP design (#770,
docs/plans/770-mcp-server/05-adversarial-review.md) raised this and flagged it as unexamined. No design document, guard or issue covers it. It is a prerequisite in spirit for #789 (EPIC: MCP server support), though not a hard blocker on any single slice.The problem
Every MCP tool result becomes prompt content in a model's context window. Some of that content is free text that a farm's own users typed. A model reading it cannot reliably distinguish "data I was asked to summarize" from "instructions addressed to me".
Concrete fields already in the phase-1 tool surface:
Customer.Name,Phone,Email,Address,Notesrc/Cluckwork.Domain/Sales/Customer.cs:16-20SalesOrder.DiscountReasonNotesrc/Cluckwork.Domain/Sales/SalesOrder.cs:28Flock.Name,Breedsrc/Cluckwork.Domain/Flocks/Flock.cs:12-13Customer.NoteandDiscountReasonNoteare genuinely free-form.Why this is more than the generic prompt-injection worry
There is a privilege gradient, and it runs the wrong way.
Note, requiresAuthPolicies.SalesFlow(CustomerEndpoints.cs:18-39).SalesFlowis everyone but ReadOnly, including a plain Worker.AdminOnly/SalesAccess.So a Worker can write free text that later lands in an Owner's model context. If that text reads as an instruction and the assistant acts on it, the action executes with the Owner's privileges, not the Worker's. A low-privilege write escalates through a high-privilege read.
A second path needs no insider at all. A customer supplies their own name or address, a clerk enters it verbatim, and it reaches the same place.
What bounds the damage today
This is worth stating so the risk is not over-rated:
AddAuthorizationFilters()removes it fromtools/listand refuses the call.The caller's own RBAC tier bounds the blast radius. Within that tier nothing bounds it, and the tier of someone running an assistant over the books is usually Owner or Manager.
Why decide it now rather than after
The cheap mitigations are shape decisions, and shape is expensive to change once tools exist and clients depend on the schema:
McpServerTool'sReadOnly/destructive annotations are set correctly, so clients can apply their own confirmation policy.Retrofitting any of these means changing published tool schemas.
What this issue is NOT asking for
This issue does not ask for a solution to prompt injection. Nobody has solved that industry-wide, and Cluckwork will not solve it. It asks for a recorded decision about presentation and boundaries, at the strength the argument supports, so the tool surface is designed with the hazard in view rather than discovering it after a schema ships.
An acceptable outcome is an explicit accepted-risk decision record saying the bounded blast radius is tolerable for phase 1, and why. Per
AGENTS.md, an accepted-risk record withNo incidentis a legitimate result, and it is load-bearing.Suggested first step
Enumerate every free-text field reachable from the phase-1 surface. The table above is a start, not a complete walk. Derive the full list from the model, per the repo's "walk everything, exclude deliberately" rule. Then decide presentation per field.