Problem / motivation
I run six agents on the floor most days. I give them a few standing principles, such as "never merge a pull request" and "do not push without my approval". Two things go wrong:
The principles fade because they live in a prompt or a memory file. After a compaction, a restart, or a long session, an agent can lose them or treat them as old context. Some principles need a backstop, not just a reminder. For a push or a merge, "the agent usually remembers" is not enough. When an agent forgets, the cost is real and hard to undo. Today my only backstops are Claude Code's own permission rules or a hand-written hook. Both live outside Munder Difflin. They are per engine, invisible on the floor, and they answer with a terminal prompt in an agent I am not watching.
Proposed solution
An opt-in guardrail, off by default. With no rules file, nothing changes: nothing is enforced and nothing is logged.
When it is turned on:
One Rules screen in Settings. A rule is a plain sentence ("Never merge a pull request"). Each rule is put into every agent's context, and put back after a compaction or a resume.
An optional backstop per rule. Before a tool call runs, the hook server checks it against the rule. The backstop can log only, ask, or block. Bash commands are parsed first (pipes, &&, sh -c, xargs, find -exec and so on), so a rule matches what actually runs, not just the first word.
Approve on a card. For a push, "ask" puts a card in ASK ME instead of a terminal prompt. The card shows the agent, the commit, the branch and where it goes. Approve lets that one exact push through, once, within an hour. Deny refuses it, and the agent is told why.
A hook token. Each agent the app starts gets its own token, so a process that claims to be an agent cannot raise cards or use approvals in that agent's name.
"Turn on guardrail" adds three starter rules: never merge a PR, never push to a protected branch, and push only with my approval. Rules are stored in one JSON file, with a backup on every save.
I have been running this on my own floor for about three weeks, and it went through several rounds of review there.
How I am thinking of landing this into PR(s): I currently have one big change and I plan to break it up into multiple small PRs off main. Each PR comes with tests (npm run typecheck and npm run test:focused) and a before and an after.
PR What Size (code, tests extra) Behavior change
1 Hook token: the hook server knows its own agents about 30 lines none visible
2 Command parser and policy engine, log only about 2,300 lines none until a rule is set to ask or block
3 Rules file and Rules screen, rules kept in context about 1,700 lines new Settings screen
4 Approvals: the card, one-time grants for a push about 700 lines new ASK ME card
The test suite is large (several thousand lines in all), mostly cases for the command parser. I am happy to trim it if you would rather review fewer.
What I am asking you:
- Do you want this in Munder Difflin at all, or would you rather it stayed a plugin or a fork?
- Naming: "guardrail", "rules", "backstop", "approvals". Do any of these clash with words you already use or plan to use?
- Is the PR order right for you, or would you rather review it another way? For example: the screen first in log-only mode, or the engine and the screen together.
- Anything you would want different before I start: where the file lives, the default rules, or the ASK ME placement.
I will not open PR 1 until you have said yes here.
Screenshot, mockup or recording
I will add these today
Alternatives considered
- Claude Code permission rules or my own hook script. They work for one engine. They are not visible on the floor, and an ask is a terminal prompt.
- Principles in the agent prompt only. This is what I had before. It works until a compaction or a long session; then it is advice, not a check.
- A separate tool outside the app. It would need its own copy of the agent list, the hook socket and ASK ME, all of which Munder Difflin already has.
Area
The hive (memory, routing, GOD agent)
Pre-flight
Problem / motivation
I run six agents on the floor most days. I give them a few standing principles, such as "never merge a pull request" and "do not push without my approval". Two things go wrong:
The principles fade because they live in a prompt or a memory file. After a compaction, a restart, or a long session, an agent can lose them or treat them as old context. Some principles need a backstop, not just a reminder. For a push or a merge, "the agent usually remembers" is not enough. When an agent forgets, the cost is real and hard to undo. Today my only backstops are Claude Code's own permission rules or a hand-written hook. Both live outside Munder Difflin. They are per engine, invisible on the floor, and they answer with a terminal prompt in an agent I am not watching.
Proposed solution
An opt-in guardrail, off by default. With no rules file, nothing changes: nothing is enforced and nothing is logged.
When it is turned on:
One Rules screen in Settings. A rule is a plain sentence ("Never merge a pull request"). Each rule is put into every agent's context, and put back after a compaction or a resume.
An optional backstop per rule. Before a tool call runs, the hook server checks it against the rule. The backstop can log only, ask, or block. Bash commands are parsed first (pipes, &&, sh -c, xargs, find -exec and so on), so a rule matches what actually runs, not just the first word.
Approve on a card. For a push, "ask" puts a card in ASK ME instead of a terminal prompt. The card shows the agent, the commit, the branch and where it goes. Approve lets that one exact push through, once, within an hour. Deny refuses it, and the agent is told why.
A hook token. Each agent the app starts gets its own token, so a process that claims to be an agent cannot raise cards or use approvals in that agent's name.
"Turn on guardrail" adds three starter rules: never merge a PR, never push to a protected branch, and push only with my approval. Rules are stored in one JSON file, with a backup on every save.
I have been running this on my own floor for about three weeks, and it went through several rounds of review there.
How I am thinking of landing this into PR(s): I currently have one big change and I plan to break it up into multiple small PRs off main. Each PR comes with tests (npm run typecheck and npm run test:focused) and a before and an after.
PR What Size (code, tests extra) Behavior change
1 Hook token: the hook server knows its own agents about 30 lines none visible
2 Command parser and policy engine, log only about 2,300 lines none until a rule is set to ask or block
3 Rules file and Rules screen, rules kept in context about 1,700 lines new Settings screen
4 Approvals: the card, one-time grants for a push about 700 lines new ASK ME card
The test suite is large (several thousand lines in all), mostly cases for the command parser. I am happy to trim it if you would rather review fewer.
What I am asking you:
I will not open PR 1 until you have said yes here.
Screenshot, mockup or recording
I will add these today
Alternatives considered
Area
The hive (memory, routing, GOD agent)
Pre-flight