Repository navigation
refactor(v4-removal)!: replace the v4 policy language with ACS v5 across the Python runtime - #3444
Merged
Merged
Conversation
PR Review Summary
Verdict: AI review comments are untrusted advisory output. The summary reports workflow-generated completion status only, not model-authored pass/fail claims. |
Dependency ReviewThe following issues were found:
License Issuesagent-governance-python/agent-governance-toolkit-core/pyproject.toml
OpenSSF Scorecard
Scanned Files
|
📦 Dependency diff (SBOM)Comparing main → liamcrumm/v4-removal-python-core. ✅ No dependency changes detected. |
liamcrumm
force-pushed
the
liamcrumm/v4-removal-python-core
branch
from
July 27, 2026 19:50
f0845b6 to
123154f
Compare
liamcrumm
force-pushed
the
liamcrumm/v4-removal-python-core
branch
from
July 27, 2026 20:17
123154f to
4e9276b
Compare
liamcrumm
force-pushed
the
liamcrumm/v4-removal-python-core
branch
from
July 27, 2026 21:57
4e9276b to
f3e6563
Compare
Both found by re-reviewing the census itself rather than the code it guards. A raise inside the guarded branch's own except handler was counted as refusing. That raise fires when the write THROWS, which is the class census 3 covers; it says nothing about the branch never being entered. Counting it exempted the repo's most common shape, a guarded write inside a try, which is the likeliest way a skipped write gets back in. Only a raise reachable when the guard is false counts now. Restoring main's shape at langchain's tool-post site is caught, where before it passed. The fallback clause stays scoped to the whole block, deliberately, and now says why. llamaindex and semantic_kernel chain several target guards and end the block with `return <result>.transformed_value`, so a guard falling through reaches that fallback rather than carrying on with the original. Scoping it to the branch flagged all four as drops. The two fixes interact and only one of them is right. agent-os 2921, agt-policies 351, agent-sandbox 499, ratchet 1415, links 0 new broken, censuses 42 apply / 24 refuse / 0 naive, 26 guards reachable, 20 raise and 5 flag-guarded rewrites, 0 swallowing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com>
MohammadHaroonAbuomar
approved these changes
Jul 30, 2026
Merged
18 of 27 tasks
MohammadHaroonAbuomar
added a commit
that referenced
this pull request
Jul 30, 2026
…c and autogen adapters (#3519) * fix(agent-os): enforce output verdicts in anthropic and autogen adapters Four adapter call sites evaluated the output intervention point via kernel.post_execute(ctx, ...) and discarded the returned (allowed, reason) tuple, so an output deny still disclosed the model response to the caller and an output transform was silently dropped. A probe against main at the #3444 stack merge confirmed all four: - anthropic_adapter.py _GovernedMessages.create (deprecated wrap path) - anthropic_adapter.py GovernanceMessageHook.create - autogen_adapter.py governed_initiate_chat - autogen_adapter.py governed_receive Each site now consumes the verdict the way the sibling adapters (bedrock, mistral, gemini, openai) already do: a deny raises PolicyViolationError with the public reason before the response is returned, and a transform — which the two-value post_execute contract reports as refused (transform_not_applicable) — blocks as well instead of forwarding the un-rewritten output. Regression tests drive the four fixed surfaces with a per-point stub runtime and assert the secret payload never reaches the caller, on deny and on an unappliable transform, plus allowed-path parity. test_transform_never_dropped.py gains a fifth census that AST-walks src/ and fails on any evaluate_*/pre_execute/post_execute call whose result is discarded (bare statement or assignment to _), the class this bug belongs to. The tree passes it after the fixes. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * fix(agent-os): give reason-less output denies a public message Copilot review on #3519: post_execute may return (False, None); the four new raise sites would surface PolicyViolationError("... None"). Default to 'denied by output policy' at all four sites and pin the behavior with one regression test per adapter. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * chore: add the census constant DISCARDABLE to the spell dictionary Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * chore: use US spelling in the census self-test name Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * test(agt-policies): bind output in the anthropic scenario manifest The adapter now enforces the output verdict (this PR), and the shared scenario manifest bound only input and pre_tool_call, so the uniform deny-on-unbound-point rule blocked the allow/transform/escalate scenarios. Bind output like the gemini scenarios do and script the second decision each flow now consumes. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> --------- Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com>
3 tasks done
MohammadHaroonAbuomar
added a commit
that referenced
this pull request
Jul 30, 2026
…ntmesh integrations (#3451) * refactor(v4-removal)!: replace the v4 policy language with ACS v5 across the Python runtime Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix(agent-os): remove dead branches and duplicate assignments in the adapters Review findings on the native adapter rewrite, all on paths the rewrite carried across rather than introduced: openai_adapter.health_check branched on has_clients to return "healthy" either way. Collapsed to match the sibling adapters. openai_adapter tool mediation had an "escalate" branch after the not-allowed guard. escalate is not a permitting verdict, here or on main, so the guard always raised first and the branch never ran. The docstring already documents the surviving behaviour: an escalate the resolver refuses surfaces as a deny. agentshield_adapter called _get_or_create_context() twice in a row at three sites. ToolCallInterceptor.intercept used a bare "..." body; the sibling AdapterRuntime protocol in this package uses a docstring. test_policy_replay_metadata imported agt.policies both ways. It now uses the module form only, which is the one patch.dict(vars(policies)) needs. The maf_adapter PolicyViolationError re-export stays. It looks unused, but test_adapter_exception_identity asserts every adapter module resolves the canonical error type, so a comment now records why it is there. agent-os 2683 passed, agent-compliance 457 passed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * ci: fail policy validation on a manifest that omits its ACS version The manifest scan keyed only on the agent_control_specification_version substring, so a manifest that omits that field was not linted at all. For a policy gate, skipping silently is the wrong direction: the file most likely to be malformed is the one missing its version. Files declaring intervention_points without a version are now reported. intervention_points is the ACS-specific tell; keying on policies or agents instead flags 22 unrelated config files in this repo that happen to use those names, which is why the check is narrow. Verified against the tree at every layer of this stack: no new findings, and a probe manifest with intervention_points and no version is caught. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * ci: let BREAKING_CHANGES.md name the v4 symbols it removes The ratchet counts any mention of a v4 name as a use of it. A breaking-changes record has to name each removed symbol so a reader can find its replacement, so the file is allowlisted. Naming a symbol in order to say it is gone is not a use of it. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: stop treating a transform verdict as plain permission A transform verdict permits the call, but it permits a rewritten one: the verdict carries the replacement the caller is expected to apply. Several gates read only whether the verdict permits, so they ran the original value while the policy believed it had been rewritten. A policy that redacts a credential did not redact it, and nothing reported a failure. Three groups of sites, three different resolutions: The five sandbox providers now refuse a transform outright at the pre_tool_call gate. That gate gets code it is about to execute and cannot rewrite it, so refusing is the only honest answer. This gate is new in this change, so the gap came in with it rather than predating it. The four adapters that called post_execute and dropped the return value now raise PolicyViolationError. main evaluated no output policy at all here, so these call sites only became load-bearing once this change added it. base.py returns (False, "transform_not_applicable") from the tuple helpers instead of silently discarding the replacement, and NativeAdapterResult grows permits_unchanged for callers that cannot apply one. allowed stays as it was; a site that can apply a transform should keep using it. Tests drive the real execute_code gate on all five providers with a real PolicyEvaluation. The five transform cases fail without the guards and the deny and allow cases pass either way, so they discriminate. agent-sandbox 499 passed, agent-os 2827 passed, ratchet 1415, lint clean on every changed file. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * docs: record this PR's removed exports where the removal happens The stack documented every Python removal in the last PR, six merges after the removals land. Anyone on main between the two saw 75 public names disappear with nothing to read. The list is computed from the export diff against the merge base rather than assembled by hand. That matters: the hand-written version I started from claimed 51 names, and of those 31 were still exported at this point in the stack because their removal happens in a later PR. The computed set is 75, and every one of them is gone here. The budget-counting change is recorded with them, since a deployed policy now permits one more tool call than it used to. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: an unbound output point is not a denial Making the adapters honour the output verdict broke every manifest that does not bind an output intervention point. The engine reports an unconfigured point as runtime_error:intervention_point_unknown, which is the right answer to a request naming an unknown point. The adapter is not answering a request though. It evaluates output after every call whether or not the host asked for output governance, so reading that error as a denial blocked every response under a manifest binding only input or only tool calls. Seven adapter scenarios caught it, and the previous code hid it by discarding the verdict. post_execute now treats that one reason as "no output policy configured", permits, and records completion. Completion mattered: it was already gated on allowed, so budgets drifted on those manifests before this change too. Every other reason still denies, including the other runtime errors. Tests pin both directions, since the risk in relaxing this is that output enforcement quietly stops working. Only the unconfigured case fails without the fix; the six covering deny, transform, allow, and the other runtime errors pass either way. agt-policies 350, agent-os 2834, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: apply the unbound-point permit at the runtime, not the base class The previous commit put the carve-out in BaseIntegration.post_execute. Only six adapters reach output evaluation that way. Roughly eleven others call evaluate_output directly (langchain, semantic_kernel, smolagents, pydantic_ai, agentshield, google_adk, guardrails, crewai, llamaindex, openai, openai_agents_sdk) and still hard-failed, so one manifest behaved two ways inside one package. Reported by Haroon on #3444. The permit now lives on NativeAdapterResult behind permit_if_unconfigured, which NativeAdapterRuntime.evaluate_output sets. Every caller goes through that method, so there is one place to reason about instead of 26 call sites. point_not_configured is exposed so callers can tell an unconfigured point from a denial. base.py drops its copy and reads allowed like everything else. The flag is set for output only. input and pre_tool_call stay fail-closed: omitting them omits governance of an action about to happen, and the fix leans on pre_execute staying strict, so a test pins that. The skip warns once per runtime rather than staying silent, since an operator otherwise cannot see that output enforcement is off. BREAKING_CHANGES records the split, including that the bridge used to rewrite any unconfigured point to an allow. Tests build the real NativeAdapterResult rather than a stand-in, because a stand-in for the class carrying the permit would test the stand-in. Removing the flag from the chokepoint fails exactly the direct-adapter test and no other, so the evidence discriminates. agent-os 2839, agt-policies 349, agent-sandbox 499, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: permit every unconfigured post-hoc point, not just output Fixing output alone left the same bug at post_tool_call. Thirteen scenario manifests bind pre_tool_call and none bind post_tool_call, so openai_agents_sdk.py:521 raised on the configuration every scenario uses. Nothing caught it because that adapter has no scenario test. The line is not output versus the rest, it is whether the point gates an action or runs after one. input, pre_tool_call and pre_model_call still deny: the action has not happened, so refusing protects it. output, post_tool_call and post_model_call permit: the tool already ran and the model already answered, so refusing protects nothing and only breaks the caller. One helper carries this now, so a fourth post-hoc point cannot pick up the wrong behaviour by being written somewhere else. Each point warns once rather than once per call. Tests cover all three post-hoc points and pin that both pre-points still deny, since the split is the whole design. agent-os 2843, agt-policies 349, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: refuse an unapplied transform everywhere, and stop permitting unbound points Two findings from re-reviewing this PR at its own head. The unbound-point permit was wrong and is reverted. I had argued a post-hoc point cannot prevent anything because the tool already ran, so refusing an unconfigured one only breaks the caller. That is false. The SDK says it plainly in AgentControlBlocked: "a block prevents the result from propagating, it does not undo the side effect." A post_tool_call block stops tool output reaching the model and an output block stops it reaching the caller, so permitting those forwarded results no policy was consulted about, on the exfiltration boundary. It also contradicted this stack's own Rust guidance, which already says a manifest must bind every point the adapter evaluates. The gemini and mistral scenario manifests now bind output, which is what langchain and llamaindex already did and what made them pass. Thirteen gates still read bare `allowed`, which is true for a transform. They had nowhere to put the replacement, so they ran the original value while the policy believed it was rewritten. They now gate on `permits_unchanged` and refuse. maf keeps reading `allowed` because it applies the transform fifteen lines further down; a first pass changed it and its tests caught the mistake. to_policy_violation now says the integration cannot apply the transform rather than reporting the policy's own reason, which described the policy instead of the problem. post_execute no longer records completion for a transform it is about to report as refused, which was charging the budget and seeding the drift baseline from output the caller is told not to use. The unused to_legacy_tuple is deleted; it returned bare `allowed` and carried the same trap. The gate test is anchored on class and method and asserts the whole set of thirteen, so a fourteenth cannot be added by copying a naive one. agent-os 2854, agt-policies 349, agent-sandbox 499, ratchet 1415, links 0 new broken, no-stubs clean. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: refuse a transform in the MCP gateway and the trust root A re-review at this PR's own head found two more consumers that dropped a transform, both migrated to the native runtime by this PR and both outside integrations/, which is why the earlier sweep missed them: it was scoped to a directory rather than to the diff. MCPGateway._evaluate answers with a bool and a reason, so it has nowhere to put rewritten tool arguments; it permitted and forwarded the originals. TrustRoot.validate_action returns a TrustDecision that carries no replacement, so it downgraded a transform to a plain allow and charged the call. This is the final authority, so that is the worst place for it. Both now refuse, as the five sandbox providers already do. The gate test is replaced. The old one listed the known sites by name, so it could not catch a consumer written somewhere new, and it asserted on source shape, which a no-op gate would satisfy. The new one is a census: it walks agent_os, finds every function that evaluates a policy and branches on the verdict, and requires each to either read transformed_value or gate on permits_unchanged. Forwarders that hand the result back are excluded, since they judge nothing. It reports 42 applying and 24 refusing, and it fails on both new findings when their guards are removed, so it would have caught them. Two behavioural tests drive the gateway and the trust root with a real transform and assert they refuse rather than forward. The unconfigured-point test also now asserts that a refused transform records no completion, which it previously left open. agent-os 2903, agt-policies 349, agent-sandbox 499, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * test: make the transform census see through delegation and past docstrings The census had two holes, both found by re-review rather than by the census itself. It only looked at functions that both evaluated a policy and branched on the verdict. Four consumers split those across two functions, so neither half was classified and all four went unchecked: pre_execute with _tuple_for, guardrails _validate with _apply_bridge_result, and agentshield's two validate_* entries with _merge_bridge_verdict. All four are correct today, but _tuple_for is the base-class transform contract that an earlier commit in this PR had to fix, so leaving it uncovered by the test written to cover exactly that class was not defensible. The census now follows one level of same-file delegation and judges the pair together. It also matched raw source, so a docstring counted. That was live, not hypothetical: deleting _tuple_for's transform check left the suite green because its docstring mentions transformed_value. Matching now runs on executable code with docstrings stripped and comments gone, and the same deletion fails the census. guardrails now fails validation on a non-string transform instead of passing the original through, since it validates strings and has nowhere to put a payload of another shape. I tried the same guard in agentshield and it was wrong. Its tool-call path applies a dict replacement to params before merging, so a non-string payload there has already been applied. A scenario test caught it; the merge keeps its string check and now records why. agent-os 2907, agt-policies 349, agent-sandbox 499, ratchet 1415, links 0 new broken. Diff-wide consumer census: 42 apply, 24 refuse, 0 naive. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: remove the pydantic-ai tool wrapper this PR orphaned _wrap_single_tool was called from _wrap_tools_for on main. This PR removed that wiring but kept the function, and its body calls kernel.intercept_tool_call, which this PR also deleted. It would raise AttributeError if anything reached it. Nothing does: it has no caller and is not exported. Found by re-reviewing this PR's own deletions for survivors, not by a test, since dead code passes every suite. agent-os 2907, agt-policies 349, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: a replacement of the wrong shape no longer falls through Forty-three sites applied a transform only when the replacement matched the shape the surface takes, and did nothing when it did not, so the original value the policy meant to rewrite was forwarded. ACS lets a transform value be any JSON value, so a manifest can return a dict for a string target and reach this; it is the same silent drop as ignoring a verdict, one level down. I had deferred this as pre-existing, since main carries 30 of these and this PR only renamed the accessor. That was not a defensible reason. It is a fail-open in a governance library, every one of the sites is in a file this PR already rewrites, and no new file is added, so the 300-file display limit does not come into it either. Twenty-six sites already raised on a denial and now raise the same way when the replacement will not fit. Sixteen with other refusal styles fold applies_to into the deny check they already had, so an unusable replacement takes that site's own path rather than a duplicated one. agentshield's validate_tool_call folds it into the returned verdict. Two sites needed individual treatment. autogen's on_send takes a dict or a string and swallows a failed write, so applies_to could not express it; it tracks whether the rewrite happened and drops the message if not. agentshield's _merge_bridge_verdict is exempt with a comment: it runs after validate_tool_call has already written a dict replacement, so refusing there would reject a correct rewrite. A first attempt did exactly that and a scenario test caught it. A second census covers the class: every guarded application must have a path for the shape it cannot take. Removing one fold fails it. Still open and reported rather than changed: 25 sites swallow a failed rewrite with `except Exception: pass`, so a type-correct replacement that cannot be written also forwards the original. Unlike the shape mismatch that is not demonstrably reachable, and whether a failed best-effort write should abort the call is a behaviour decision across every adapter. agent-os 2923, agt-policies 349, agent-sandbox 499, ratchet 1415, links 0 new broken. Consumer census 42 apply, 24 refuse, 0 naive. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * test: fold the unconfigured-point tests into the transform census file Both modules cover the same thing: what an adapter must do with a verdict it cannot act on. Keeping them apart pushed this PR to 301 files, one past the point where GitHub stops rendering the diff, and the split bought nothing. No coverage is lost. The unconfigured-point classes move across whole and reuse the census file's stubs rather than carrying their own near-duplicates. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * test: move the framework SDK stubs into conftest They were a module three test files imported by hand. conftest is where pytest already looks, so the separate file bought an import line and a file. This does not change the file count, which I claimed in the first draft of this message before checking: deleting the module drops one file but conftest becomes modified, so the diff stays at 300. That is at GitHub's rendering limit rather than past it. Keeping the change because it is the right home for the stubs, not because it buys headroom. agent-os 2923 passed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: catch the transform drops hiding behind a third condition Two sites survived the previous pass and the census written to guard it, both for the same reason: their guard ANDs a third condition, and the fixing scripts and the census all keyed on a two-operand shape with the transform check first. maf's input middleware applied a redaction only when the replacement was a string AND a message existed, so a dict replacement forwarded the original text with call_next proceeding. langchain's wrap_tool_call applied only when the replacement and the tool call were both dicts, so anything else ran the original arguments. Both were reachable and neither denied. The census now matches an AND of any arity where any operand is the transform check and any is an isinstance call. Broadening it first made it fail on exactly these two and nothing else, which is what says it would have caught them. Both use the applied-flag form rather than applies_to, because both wrap the write in except Exception: pass. That closes the write-failure drop at these two sites as well as the shape drop. Verified after: 26 inserted guards still leave their apply body reachable, 16 applies_to folds still match their site's function, variable and type, census 42 apply / 24 refuse / 0 naive. agent-os 2923, agt-policies 349, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: audit and surface a refused transform the way a denial is Both refusal paths added in the previous commit were incomplete. maf raised MiddlewareTermination and nothing else, while the deny path beside it logs, sets a user-visible context.result and writes an audit record. A block that leaves no record is not acceptable in a governance runtime, so the refusal now does all three. langchain logged a refused transform as "Policy DENY" and printed the transform's own reason, so an operator would read "DENY: pii_redaction" for a call that was refused because the replacement could not be written. The line now distinguishes the two. The raised exception was already correct, since to_policy_violation branches on transform. A behavioural test drives the maf middleware with a dict replacement and asserts it refuses, does not call the agent, leaves the message intact and records the block. It fails on the pre-fix shape. It lives with the other maf scenarios rather than in agent-os, where driving the middleware would have needed a stub of everything around it. agent-os 2923, agt-policies 350, agent-sandbox 499, ratchet 1415, links 0 new broken. Census 42 apply, 24 refuse, 0 naive; 26 inserted guards reachable; 16 folds correctly placed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: a rewrite that fails to land is refused, not swallowed The last open finding. Twenty-four sites applied a transform inside a try whose handler only passed, so a write the framework object refused left the original value in place and execution carried on while the policy believed it had rewritten. The write targets objects the adapter does not own, and frozen pydantic models and immutable message contents do raise on assignment, so this was reachable rather than theoretical. I had reported it rather than fixed it, on the grounds that it needed a behaviour decision. The same reasoning was rejected for the shape mismatch and it is no better here. Eighteen sites whose function already raises to_policy_violation now raise it from the write failure. maf raises MiddlewareTermination, matching its deny path. crewai's before_tool and before_llm and autogen's on_send track whether the rewrite landed and take the refusal each already had, because their targets are searched for in a loop or written through two different shapes, so a raise would be in the wrong place. crewai's before_llm also refuses when no message takes the rewrite at all, which had the same effect as a failed write. A third census fails on any transform rewrite whose handler only passes, unless the function tracks whether it landed. It reads both spellings, since some sites hold the replacement in a local first. Reintroducing a swallow in gemini fails it. agent-os 2921, agt-policies 350, agent-sandbox 499, ratchet 1415, links 0 new broken. Census 42 apply, 24 refuse, 0 naive; 26 inserted guards reachable. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: a write target that takes nothing is refused, not passed over An eighth finding, and a fourth way the same value gets dropped. The three guards so far cover the verdict, the replacement's shape, and a write that raises. This one is a write that is never attempted: a guard on what is being written TO fails, no exception is raised, and the original value goes out. Two were demonstrated end to end. autogen's _apply_content buried the swallow in a helper that returned None, so on_send's rewritten flag could never go false and on_publish had no check at all; a message whose content rejects assignment forwarded the original with allowed=True. langchain's wrap_tool_call wrote only when the result had .content, and a tool returning a plain string has none, which the same function's own str(getattr(result, "content", result)) shows is expected. _apply_content now returns whether the write landed and both callers read it. The rest refuse on the branch that cannot take the replacement: langchain's model input loop, which could complete without writing anything, its response path, anthropic's and mistral's input loops on a non-dict message, mistral's tool call with no function to write to, google_adk's non-dict tool args, and smolagents' tool call carrying neither argument attribute. A fourth census covers the shape: an application guarded on the target with no refusal, no flag and no else. It found a ninth site in mistral that was not in the report, and reintroducing the anthropic guard fails it. agent-os 2921, agt-policies 350, agent-sandbox 499, ratchet 1415, links 0 new broken. Census 42 apply, 24 refuse, 0 naive; 26 inserted guards reachable; 20 raise and 5 flag-guarded rewrites, 0 swallowing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: redact and continue where the value is the return, not an attribute The previous commit over-corrected langchain. A tool returning a plain string has no .content to write to, and I made that a hard denial, which turns a redaction policy into a block for the most ordinary kind of tool. llamaindex and semantic_kernel already handle the same case by returning the replacement itself, which is right: the value IS the return, so handing back the redacted one applies the policy rather than refusing it. Both langchain sites now do the same. The comment I left there was also wrong; returning transformed_value hands back the redaction, not the original. Census 4 flagged the corrected form, since it only recognised a raise, an else or a landed-flag. Returning the replacement is a third correct handling and it now counts. Reintroducing the anthropic drop still fails the census, so it did not go slack. agent-os 2921, agt-policies 350, agent-sandbox 499, ratchet 1415, links 0 new broken, ratchet self-test and tests/ci 108, no-stubs clean. Census 42 apply, 24 refuse, 0 naive; 26 guards reachable; 20 raise and 5 flag-guarded, 0 swallowing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * test: make the verdict stubs read the same at every layer of the stack The stubs modelled the pre-ACS evaluation, which reads verdict, transform and reason off the evaluation itself. The ACS runtime reads them off a Verdict object instead, so the same tests failed at the last PR in the stack with 'str' object has no attribute 'transform', and the session there routes every intervention point through one evaluate() rather than a method per point. The stubs now carry both shapes: a verdict that is still a string but also answers .decision, .reason and .transform, and a session that accepts either spelling. One version of the file passes at both ends of the stack, so the tests do not have to be ported layer by layer as the runtime changes underneath them. agent-os 2921 at this layer and at the tip. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: spell-check on the words this PR introduces cspell runs on changed lines and flagged five. Three were mine to reword or rename: a comment said "unrewritten", and two locals in the census helper were funcname and fbody, now func_name and enclosing_src. conftest and orelse go in the dictionary instead. The first is pytest's own filename, which this PR imports from now that the framework stubs live there; the second is the attribute name on ast.If, which the census reads to tell an unguarded branch from one that refuses. agent-os 2921, ratchet 1415. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: refuse a transform langchain cannot write, instead of returning a bare string The previous commit had these two sites return post_result.transformed_value when the target had no .content, on the grounds that llamaindex and semantic_kernel do the same. That analogy was false and the change was wrong. LangChain types wrap_tool_call as returning ToolMessage | Command and wrap_model_call as ModelResponse | AIMessage | ExtendedModelResponse. A str is none of those. It also breaks the premise I wrote in the comment: a tool that returns a plain string is wrapped as a ToolMessage before middleware sees it, and ToolMessage has .content, so what actually arrives without one is a Command. Returning a str there makes ToolNode fold it into messages as a HumanMessage, losing the command's updates and its tool_call_id; on the model path the caller then reads .result off a str and raises AttributeError. The siblings differ because their return type is Any by contract. Both sites refuse again, which is protocol-safe, fail-closed, and the same answer the sandbox providers, the MCP gateway and the trust root already give when they cannot apply a replacement. Refusing also removes a budget bug the early return had introduced: it skipped self._ctx.call_count += 1, so a redacted tool call was never charged and max_tool_calls undercounted. A scenario test drives wrap_tool_call with a Command and a transform verdict and asserts it refuses. It fails on the previous revision, which returned str. Census 4 now reads the If node's own else branch rather than searching the unparsed source for "else:", which also matched an else on a nested for, try or inner if and could mark a dropped transform as handled. agent-os 2921, agt-policies 351, ratchet 1415, censuses 42 apply / 24 refuse / 0 naive, 20 raise and 5 flag-guarded rewrites, 0 swallowing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * test: close two holes in the target-guard census Both found by re-reviewing the census itself rather than the code it guards. A raise inside the guarded branch's own except handler was counted as refusing. That raise fires when the write THROWS, which is the class census 3 covers; it says nothing about the branch never being entered. Counting it exempted the repo's most common shape, a guarded write inside a try, which is the likeliest way a skipped write gets back in. Only a raise reachable when the guard is false counts now. Restoring main's shape at langchain's tool-post site is caught, where before it passed. The fallback clause stays scoped to the whole block, deliberately, and now says why. llamaindex and semantic_kernel chain several target guards and end the block with `return <result>.transformed_value`, so a guard falling through reaches that fallback rather than carrying on with the original. Scoping it to the branch flagged all four as drops. The two fixes interact and only one of them is right. agent-os 2921, agt-policies 351, agent-sandbox 499, ratchet 1415, links 0 new broken, censuses 42 apply / 24 refuse / 0 naive, 26 guards reachable, 20 raise and 5 flag-guarded rewrites, 0 swallowing. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix(integrations): stop transform verdicts from proceeding unapplied A transform verdict permits, so code that gates on is_allowed alone lets the run continue with the original payload while the policy believed it had been rewritten. A redaction policy silently does not redact. Neither surface here can apply a replacement: the lifecycle hooks observe a run, and a guardrail can only trip or not. Both now refuse a transform rather than proceeding. The guardrail fix covers its input and output sites, not just the one flagged, since the two shared the defect. Also in this package, the input guardrail passed a list of input items straight to evaluate_input, which declares body as str or dict. The list branch now encodes to a dict. Two packaging faults that would ship: adk_agentmesh/__init__.py is a deprecation shim with no re-exports, so every import in its own README raised ImportError. The four documented names are re-exported, and each one was imported to confirm it. agent-governance-toolkit-integrations declared the console script openshell-agentmesh = openshell_agentmesh.cli:main while this stack deletes cli.py, so the wheel would install a command that crashes on first run. The entry is removed; only a deprecation shim remains in that package. Tests cover every verdict on both surfaces. They skip where the openai-agents SDK is missing, as the package's existing tests already do, so the behaviour was additionally verified against stubbed SDK modules: allow and warn proceed, deny, escalate and transform refuse, and the guardrail body encodes to a dict. Ratchet 1180. Links: 0 new broken. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * refactor(v4-removal)!: move the examples onto ACS v5 manifests (#3452) * refactor(v4-removal)!: move the examples onto ACS v5 manifests Example policies and bundles move to ACS manifests, and the demos that existed only to show v4 syntax are deleted with their fixtures. The package-local examples, docs, tutorials, benchmarks and notebooks under agent-os follow the same move. docs/compliance/nist-rfi-2026-00206.md moves here too, because it links to the native A2A runtime example this PR adds and the link would dangle if the two landed separately. Also carries two review fixes: sync_atr_rules interpolated third-party ATR patterns into a Rego backtick raw-string literal. Raw strings have no escape mechanism, so a pattern carrying a backtick closed the literal and the rest parsed as policy source. The weekly sync compiles that corpus from npm unattended, so a poisoned upstream rule could flip a verdict on benign input. Patterns are now emitted as JSON-quoted strings and the validator rejects backticks. atr-import shares the emission path through compile_patterns. The demos README documented AgtRuntimeMiddleware, which is RuntimeGovernanceMiddleware, and advertised real LLM calls with Ollama setup instructions. The example it points at makes no model call, so that section described a mode that does not exist. Ratchet: 656. Links: 0 new broken. atr tests: 5 passed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * docs(rust): correct the dependency audit and document the unbound-point deny (#3445) The audit claimed nothing was pulled from crates.io that was not already in the graph. Cargo.lock gains 66 entries: 2 first-party and 64 third-party crates that were not in the workspace before. It now lists them and names the two clusters. The second cluster is the one that matters: ureq, rustls, ring and webpki-roots link an HTTP and TLS stack into agentmesh. They arrive because the ACS Rust SDK pins its core features to include the remote annotators and because Manifest::from_url exists, and a consumer cannot opt out since the feature set is fixed inside the SDK crate. Nothing in agentmesh calls a remote annotator or from_url, so no egress is added in practice, but for a governance library the linked capability belongs in the security record. Narrowing the SDK's features is a follow-up on that crate. MIGRATION_V5.md also now documents that the runtime denies an intervention point the manifest does not declare. Failing closed is right, but the adapter evaluates input on every execute() and pre_tool_call on every tool call, so a manifest binding only one denies the other path with no obvious cause. Links: 0 new broken. Signed-off-by: Liam Crumm <liamcrumm@gmail.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f * docs(v4-removal): rewrite the documentation for ACS v5 (#3446) * docs: use the real class names in the integration guides A mechanical rename left several component names that resolve to nothing: AgtRuntimeMiddleware and AgtRuntimeMW are RuntimeGovernanceMiddleware, and CapabilityGuardMW, AuditTrailMW and RogueDetectionMW are the corresponding *Middleware classes in maf_adapter. AgtRuntimeChecker never existed in any SDK. The Haystack integration ships TrustGate and AuditLogger, which is what the tables now list. PolicyEvaluation.* in the .NET tutorial was the rename applied where it does not belong: the .NET engine still defines PolicyAction and no PR in this stack touches dotnet source. Reverted, and the file is allowlisted in the ratchet so documenting that engine's own enum does not read as v4 usage. Ratchet: 94, unchanged. Links: 0 new broken. Ratchet tests: 35 passed. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * refactor(v4-removal)!: confine the v4 policy language to the migration tool (#3447) * docs: scope the breaking-changes record to what each PR removes The Python entry listed the whole stack's removals in one place, so it overlapped the record the first PR now carries and named 33 symbols twice. It also claimed agent_os.kernel keeps a PolicyRule; the surviving one is re-exported from agent_control_plane. What is left here is what could only go once the rule model went: manifest resolution, the evaluators and decision types that read the rule model, the framework integrations' local policy surfaces, and the unreleased cedarling backend. Names that survive under the same spelling in an unrelated module are called out, since agent_os still exports a PolicyDecision and a PolicyRule that have nothing to do with the v4 language. The langflow README documented GovernanceComponent, which this stack deletes. Its section, the YAML policy that only configured it, its table row, and its node in the flow diagram are gone, matching what the haystack and openshell READMEs already do. ratchet 0, links 0 new broken. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * refactor!: delete agt.policies and evaluate through ACS directly (#3449) * fix(acs): drop the dead manifest-timeout branch and correct the docs it claimed _resolve_approval_timeout read approval.timeout_seconds off getattr(control, "manifest", None). AgentControl never stores that attribute, so the branch was dead in production and only ran in tests because the stub defined one. BREAKING_CHANGES then documented behaviour the code could not deliver. The core treats the approval section as opaque host configuration (SPECIFICATION §17.1) and does not surface it across the native boundary, so honouring it needs an SDK accessor that does not exist. The branch is removed, the docstring and BREAKING_CHANGES say a caller must pass the value, and the 300 second default that closes the original fail-open stays. The tests that covered the branch went with it. A replacement drives a resolver that never returns through a real HostSession and asserts it denies in about a second, which the removed tests could not have caught. Also repairs what earlier mechanical renames left behind: AgentControlMiddleware and AgentControlMW in the Python and MAF docs are RuntimeGovernanceMiddleware. The .NET class of that name is real but belongs to the ACS dotnet SDK, so the name resolved while meaning the wrong thing. AgentControlChecker and HaystackAdapter exist nowhere; the Haystack package ships TrustGate and AuditLogger. NoOpAgentControl and the compat.py path beside it were both deleted by this stack. Ten `the ACS manifest` substitutions where a symbol name used to be, across AGENTS.md, three compliance docs and openclaw-sidecar.md. Fourteen compliance citations pointing at agt/policies and agent_os/policies files this stack deletes, repointed at the ACS SDK. The agt-policies README still described the deleted wrapper, including two submodules that never existed and a snippet calling AgentControl.evaluate. Rewritten for what the package now is, a migration tool, and every command and symbol in it was run. ACS host: 25 passed. agent-os: 2683 passed. Ratchet 0. Links: 0 new broken. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * docs: cite the live Haystack integration, not the deprecated package The framework tables in docs/FAQ.md and agent-mesh/README.md pointed at the deprecated haystack_agentmesh package and its TrustGate/AuditLogger components. The maintained module is agentmesh.integrations.haystack, whose exports are TrustGateComponent, TrustAgentComponent, TrustedPipeline, and PipelineAuditEntry. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * test: delete the dead __init_result__ stub from _RecordingControl The method is not a dunder Python calls, nothing in the suite invokes it, and its body only raises NotImplementedError. Verified by AST that _RecordingControl keeps __init__ and evaluate_intervention_point; the PyO3 _native extension is not built in this environment, so the module was checked with py_compile rather than by running pytest. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * docs: describe the post-removal world with AgentControl, not AgtRuntime Two sentences in BREAKING_CHANGES.md used AgtRuntime for what code takes after the removal, contradicting the file's own removed-symbols table, which maps AgtRuntime(path) to AgentControl.from_path. The mentions that remain describe the removed v4 surface and stay as they are. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * docs(openai-agents-trust): note the hooks path carries no tool arguments The openai-agents lifecycle hook signature (on_tool_start/on_tool_end) does not include the tool-call arguments, so GovernanceHooks evaluates pre_tool_call and post_tool_call with args={}. Policy rules conditioned on tool arguments therefore never match on this path. Say so at both call sites and in the README, pointing at the guardrail integration for arg-level enforcement. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * fix(openshell): stop pointing the deprecation at a package that ships this shim The DeprecationWarning and README told users to install agent-governance-toolkit-integrations[openshell], but that extra force-includes this same package, whose skill (GovernanceSkill, ShellPolicyViolation, governed_shell) this stack removes - the pointer was circular. Say what is true: the skill is removed in the v5 ACS migration with no OpenShell-specific replacement; build an AgentControl from an ACS manifest and evaluate intervention points in the host. Packaging and force-includes are unchanged. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * docs(examples): remove the openshell-governed demo The v4-removal stack deletes openshell_agentmesh.skill (GovernanceSkill), which was this demo's only entry point. The example could no longer run; remove it with the rest of the retired openshell surface. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * docs: retire the openshell integration page The v4-removal stack deletes the openshell governance skill this page documents (its quickstart imports the removed GovernanceSkill). Remove the page and the openclaw-sidecar cross-link along with the example removed in the previous commit. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> --------- Signed-off-by: Liam Crumm <liamcrumm@gmail.com> Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f --------- Signed-off-by: Liam Crumm <liamcrumm@gmail.com> Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f --------- Signed-off-by: Liam Crumm <liamcrumm@gmail.com> Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f --------- Signed-off-by: Liam Crumm <liamcrumm@gmail.com> Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f * fix(ratchet): restore enhanced scanner, dedupe BREAKING_CHANGES entry, add _migrate_resolution prefix, update baseline post-merge Signed-off-by: GitHub <noreply@github.com> * fix: restore what the conflict resolution dropped, and address review The automatic resolution of the merge with main deleted 158 files that both main and this branch still had. They were not deliberate removals: examples, benchmarks, docs, and the ``agt.policies`` modules ``manifest``, ``session`` and ``_re2``, whose absence broke every agent-os test that loads a manifest. All are restored from the pre-merge branch tip. The resolver had also taken the downstream package's ``agt/__init__.py``, which assumes ``agt.policies`` is gone; the version here exports both surfaces and resolves ``cli`` lazily, because importing the migrator eagerly pulls in the native SDK and made ``from agt.policies import AdapterRuntimeSession`` fail on a host without it. Three moves the resolver got half-right are completed rather than reverted: the resolution layer now lives only under ``agt.cli._migrate_resolution``, so its last two importers point there, and the scenario harness that reached into it from outside the CLI is gone along with its only consumer, which is what the migrator-boundary test asks for. The ratchet baseline was recomputed on the corrected tree; it had been written against the tree missing those files. Review feedback: the ADK README's quickstart named a class the package does not export, so it raised ImportError on the first line a reader runs; it names ``ADKPolicyEvaluator`` now and the snippet was executed. The lifecycle hooks in openai-agents-trust called ``audit_record()`` on a result type that has no such method, and handed the audit log a verdict object where it declares a string and then hashes it, so every governed tool event raised before it could decide anything; both writers now share one JSON-safe record builder. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix: rebuild the merge on the stack's end state and green the CI gates The automatic resolution merged #3519 — a fix branched from main — into the stack and took main's side wherever they overlapped. That resurrected ``agt.policies``, which this stack deletes, leaving a tree that was neither state: the wheel gate failed because the migration wheel shipped a second policy layer, and the resolver package was half-present. The tree is rebuilt as the stack's end state plus main's independent changes, which is what the merge was supposed to produce. #3519's behavioural fix is re-applied on top rather than dropped: the four adapter sites that evaluated the output intervention point and discarded the verdict now raise on a deny, so a blocked response is not disclosed, and the anthropic scenario manifest binds the output point those tests now exercise. The remaining gate failures were consequences of the same move. The crypto allowlist still named ``manifest_resolution/build.py`` after the stack moved it under the private migrator, so its long-standing SHA-256 exemption stopped matching. Two integration packages lost their license headers. The spell checker met crate names from the Rust dependency audit and two British spellings. Two gitleaks findings are documentation placeholders, fingerprinted the way this repository already handles that file. Review feedback: the ADK README named a class the package does not export, so its quickstart raised ImportError on the first line; it names ``ADKPolicyEvaluator`` now and the snippet was executed. The openai-agents lifecycle hooks and guardrails called ``audit_record()`` on a result type that has no such method and handed the audit log a verdict object where it declares a string and then hashes it, so every governed event raised before deciding; both writers share one JSON-safe record builder. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix(ci): repair dependency pins, ACS SDK installs, and stale doc links Fixes the CI failures left by the automatic merge-conflict resolution on this branch. Six distinct root causes, not one: - agent-governance-toolkit-core kept the stack's stale `agentrust-trace>=0.2.0,<0.3.0` instead of main's dependabot bump to `>=0.5.1,<0.6.0`, making agent-mesh and docker-compose-test unresolvable (ResolutionImpossible) and breaking the TRACE v0.2 sinks. - agent-compliance, agent-marketplace and agent-sandbox now import `agent_control_specification` at module scope, but the CI gate that builds the native ACS SDK only covered agent-os and agt-policies. - The Policy Validation workflow calls the ACS-backed manifest linter without ever building the ACS SDK. - test-integrations resolved the published ACS 0.3.1b0 wheel, which predates `HostSession`; build the vendored SDK first so the resolver sees the requirement already satisfied by this commit's source. - test_hooks_verdicts asserted a `HostSession.evaluate_input(body=...)` API that does not exist. The real intervention point is `HostSession.input(body: JsonValue)`, which accepts a list, so the test now covers what actually matters: every non-dict item is stringified so the body stays JSON-serializable. - Three i18n READMEs still linked `docs/integrations/openshell.md`, which the stack deletes, and the dcbadge ignore pattern pinned the previous Discord invite code. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f Signed-off-by: Liam Crumm <liamcrumm@gmail.com> * fix(merge): keep the post-base Discord invite fix and Rust dependency pins Two main-side deltas from after the merge-base were reverted by the stack-end-state rebuild and must survive the merge: - #3481: README.md and README.ko.md returned to the expired 7aVPCcVh Discord invite; restore the non-expiring TxMRqY3pFr link - Dependabot: re-lock the base64 0.23.0, clap 4.6.4, regorus 0.11.0, and cedar-policy 4.12.0 pins against the branch workspace (cargo update --precise per crate) Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> * test(agent-os): re-add the #3519 output-verdict regressions and census The rebuild dropped the regression tests that pin the anthropic and autogen output-verdict enforcement (the adapter-side fix itself is in the tree) and the discarded-verdict census in test_transform_never_dropped.py. Re-add them with the per-point stub adapted to the ACS evaluate_intervention_point protocol; all pass against the branch adapters. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> --------- Signed-off-by: Liam Crumm <liamcrumm@gmail.com> Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Signed-off-by: GitHub <noreply@github.com> Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Co-authored-by: copilot-swe-agent[bot] <198982749+Copilot@users.noreply.github.com> Copilot-Session: 384bda36-161f-4689-9b2c-0a6d5c90776f
Oluwajuwon Omotayo (kingztech2019)
added a commit
to kingztech2019/agent-governance-toolkit
that referenced
this pull request
Jul 30, 2026
… packs to ACS v5 Upstream main shipped two breaking-change PRs (microsoft#3444, microsoft#3451) while this PR was in review, replacing the legacy v4 PolicyDocument YAML format with native ACS manifests across the entire examples/policies/african-regulatory directory. agent_os/policies/schema.py (PolicyDocument/PolicyAction/ PolicyOperator) no longer exists on main. These 4 packs were the only files left in the old format, which is why the v4 removal ratchet correctly failed after merging upstream/main. - YAML: replaced the old rules:/condition:/action: PolicyDocument shape with an ACS manifest (agent_control_specification_version, policies, intervention_points), matching the exact shape of the 13 other already- migrated packs (Kenya, NDPA, CBN, BVN/NIN, NFIU, POS-geofencing, POPIA, Uganda, Tanzania, Ethiopia, and the 3 universal agent-safety packs). - Rego: appended the same "Native ACS result adapters" block used by all 13 comparison packs (acs_input_result, acs_pre_tool_call_result, acs_output_result via the shared agt_policies.acs adapter). The existing deny/escalate/audit/decision rules are completely unchanged. Verified: - All 4 manifests validate against policy-engine/spec/schema/manifest.schema.json - All 4 packs' acs_input_result/acs_pre_tool_call_result/acs_output_result manually exercised end-to-end against realistic ACS snapshots - opa check clean, 144/144 opa test pass (unchanged, since the underlying rules were never touched) - v4 removal ratchet passes (grand total 4, matches baseline — resolved by the format change itself, no baseline edits needed) - jurisdiction-router.rego already correctly wired (points at .decision, which is untouched), no changes needed Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com>
Dipika Ranabhat (qubeena07)
added a commit
to qubeena07/agent-governance-toolkit
that referenced
this pull request
Jul 30, 2026
…enforcement (microsoft#3068) Adds an AppArmor profile (agt-sandbox) that enforces the sandbox command denylist at the kernel level for Docker containers, closing the gap where the existing denylist was advisory only. DockerSandboxProvider now applies the profile when it is loaded in enforce mode and falls back to docker-default with a warning otherwise. Addresses review feedback from the original PR: - Deny rules use mrx/mx so ld.so-based invocations and copy-then-exec from writable directories (/tmp, /workspace) are blocked, not just direct exec. - Removed a ptrace read rule that contradicted the docker-default allow form. - The profile header now accurately scopes the enforcement to exec-level denylist only; in-process network access is explicitly called out as unblocked. - _apparmor_profile_loaded only returns True for a profile in (enforce) mode; (complain) mode logs but does not block, so it is not a safe confinement signal. - _fallback_apparmor_profile emits a one-time warning when falling back to docker-default instead of failing silently. Rebased onto current main; agent-governance-rust/agentmesh's prompt_injection.rs and tests/prompt_injection.rs are untouched by this change (verified via diff against origin/main). docker_config_from_policy was dropped from provider.py during the rebase since main's v4-removal refactor (microsoft#3444, microsoft#3451) deleted every caller of it along with the v4 policy model it served; it had become dead code. Signed-off-by: qubeena07 <qubeena7@gmail.com>
4 tasks done
This was referenced Jul 31, 2026
Merged
This was referenced Aug 4, 2026
MohammadHaroonAbuomar
pushed a commit
that referenced
this pull request
Aug 5, 2026
… Egypt, Mauritius) (#3243) * feat(examples): add West/North Africa policy packs (Ghana, Rwanda, Egypt, Mauritius) Extends the African regulatory policy pack from 6 to 10 jurisdictions, adding four new YAML + Rego + test bundles under examples/policies/african-regulatory/. ─── Ghana (GH) — Data Protection Act 2012 (Act 843) ─────────────────── - s.17: Data minimisation — blocks disproportionate record exports (>1000), escalates moderate exports (101-1000) - s.37: Special personal data — escalates health, ethnic, religious, political, criminal data; blocks biometric transmission - s.38: Cross-border transfer — blocks transfers to non-adequacy regions; requires documented consent for non-GH destinations - s.40: Breach suppression — deny rule covering all suppression/delay patterns - NIA Act 707: Ghana Card (GHA-XXXXXXXXX-X) regex — blocks national ID in agent output - 28 OPA tests ─── Rwanda (RW) — Law No. 058/2021 on Personal Data Protection ───────── - Art. 3(2)/10: Sensitive data — biometric, health, racial, religious, political, criminal; blocks transmission, escalates processing - Art. 21: Right against automated decisions — escalates all automated decision actions (auto_approve, auto_reject, algorithmic_decision, etc.) for human oversight; first AI-agent-specific control in the African section - Art. 43: 48-hour NCSA breach notification — deny on suppression (strictest breach timeline on the continent) - Art. 48/49: Cross-border transfer — blocks to non-RW without consent or contractual safeguards - Art. 50: Storage location — blocks large exports, escalates moderate - NIDA: Rwanda National ID (16-digit) regex — blocks in output - 34 OPA tests ─── Egypt (EG) — Personal Data Protection Law No. 151/2020 ──────────── - Art. 2: Financial data as sensitive — Egypt is the only African jurisdiction that classifies credit scores, account balances, and loan history as sensitive data (same tier as health/biometric); escalates on detection - Art. 2: Children's data as sensitive — escalates minor/children data processing patterns - Art. 2: Biometric, health, genetic, religious, political, criminal data — full sensitive category coverage - Art. 14: PDPC operating licence — blocks agents that acknowledge operating without PDPC licence - Art. 18: 72-hour PDPC breach notification — deny on suppression - Art. 20: Cross-border transfer — blocks to non-EG/non-adequacy regions; requires PDPC approval basis - Egypt NID: 14-digit national ID (birth-date-embedded format) regex - 39 OPA tests ─── Mauritius (MU) — Data Protection Act 2017 (Act 20/2017) ─────────── - Most GDPR-aligned African DPA; mandatory DPO for ALL controllers with no size threshold (stricter than EU GDPR) - s.31: Mandatory registration — blocks registration bypass patterns - s.31: DPO requirement — escalates DPO bypass/omission patterns - s.24: Special categories — health, genetic, biometric, racial, ethnic, political, religious, trade union, criminal; blocks biometric transmission, escalates special category processing - s.39: Breach notification — deny on suppression - s.45: Cross-border transfer — adequacy or consent required - s.49: Automated decision-making transparency — escalates automated decision actions for human oversight - Mauritius NIC: [A-Z][0-9]{6,7} format regex — blocks in output - 39 OPA tests ─── Infrastructure ────────────────────────────────────────────────────── - rego/jurisdiction-router.rego: GH, RW, EG, MU added to jurisdiction_policies and policy_queries - README.md: jurisdiction list, architecture diagram, policy table, and test count (384 → 140) updated All 140 OPA tests pass: opa test rego/ -v → PASS: 140/140 Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(examples): address PR #3243 reviewer corrections to African policy packs Ghana DPA (Act 843): - Correct cross-border section cite: s.38 → s.18(2) (s.38 is the religious-belief exemption to s.37, not the transfer provision) - Correct security/breach sections: s.40 → s.28-30 (safeguards) / s.31 (breach notification to DPC + data subjects) - Correct principles range: s.17-33 → s.17-26 Egypt PDPL: - Fix Executive Regulations decree: "Prime Ministerial Decree No. 1699/2021" → "Ministerial Decree No. 816/2025" (published 1 Nov 2025) Mauritius DPA: - Fix NIC regex: [A-Z][0-9]{6,7} → [A-Z][0-9]{13} (real format is 1 letter + 13 digits, 14 chars total, e.g. A1234567890123) - Align mauritius-skip-dpo action: rego escalate → deny to match yaml - Update existing NIC test inputs to 14-char format; add test confirming the corrected regex matches, and a negative test that old 7-char codes no longer trigger the rule Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(schema): add escalate to PolicyAction enum and JSON schema escalate is semantically distinct from block — it routes execution to a human-in-the-loop review rather than hard-stopping. The action was already implemented in integrations/escalation.py but missing from the PolicyAction enum and policy_schema.json, causing cli validate to reject any YAML that used it (including the Ghana, Rwanda, Egypt, and Mauritius packs in this PR, plus the already-merged Ethiopia, Tanzania, and Uganda packs). Resolves the CHANGES_REQUESTED on PR #3243. Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * Revert "fix(schema): add escalate to PolicyAction enum and JSON schema" This reverts commit 6ca588a. Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(examples): convert action:escalate to block in west/north Africa packs Per MohammadHaroonAbuomar review on PR #3243: this PR is examples-only and must not pull in unapproved runtime schema changes. The schema revert (c8c4d0a) removed PolicyAction.ESCALATE from schema.py and policy_schema.json. These 28 rules across Ghana, Rwanda, Egypt, and Mauritius YAMLs are now set to block (matching PR #3254's interim approach) until escalate is canonically wired to escalation.py with tests and docs in a dedicated PR. Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(ci): add West/North Africa regulatory spell-check terms Adds jurisdiction-specific vocabulary needed for the Ghana, Rwanda, Egypt and Mauritius policy pack files to pass cspell: - British data-protection spellings: minimisation, pseudonymisation - Nationality adjectives: Ghanaian, Rwandan - Place name: Kigali - Administrative term: Governorate (Egypt ID format) - Legal usage: Gazetted (Rwanda gazette reference), Cybercrimes (Egypt law) - Regulatory bodies: NCSA, NIDA, RISA (Rwanda); PDPC, PDPL, MCIT (Egypt); NIA, DPC (Ghana) - Currency/country codes: GHA, RWF, EGP, MUR, YYMMDD Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(examples): address Copilot review findings on PR #3243 - ghana-dpa.yaml/rego: require an actual GHA-formatted ID after the "Ghana Card"/"Ghana ID" label instead of matching the bare label, which caused false-positive denies on innocuous mentions - add regression test for the label-only (no ID) case - fix stale README test count and wrong "Run:" path in the four *_test.rego header comments - fix test_escalate_dpo_bypass (Mauritius) checking the wrong action: the rule was already aligned to deny per earlier review, the test was never updated to match Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(examples): address follow-up Copilot review on PR #3243 - ghana-dpa.yaml/rego: the GHA regex was unanchored, so GHA[0-9]{10} matched as a substring inside unrelated identifiers (probe-confirmed by reviewer: "SGHA1721234567", an order reference, triggered a priority-1 deny). Added \b boundaries and scoped the case-insensitive flag to just the label alternative, keeping the raw ID match uppercase-only. Verified identically in Python's re (the engine that enforces the YAML) and Go's RE2 (opa eval) against both the false positive and the existing true-positive cases. - Added a regression test for the unrelated-reference false positive. - .cspell-repo-terms.txt: add "hostnames" (flagged in an added comment line) and "SGHA" (the fictitious order-reference prefix used in the new regression test). Left the Rwanda bare HIV/race substring-match nit alone per the reviewer's own note that it's a pre-existing, repo-wide convention to address in a separate follow-up, not specific to this PR. Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * refactor(african-regulatory): migrate Ghana, Rwanda, Egypt, Mauritius packs to ACS v5 Upstream main shipped two breaking-change PRs (#3444, #3451) while this PR was in review, replacing the legacy v4 PolicyDocument YAML format with native ACS manifests across the entire examples/policies/african-regulatory directory. agent_os/policies/schema.py (PolicyDocument/PolicyAction/ PolicyOperator) no longer exists on main. These 4 packs were the only files left in the old format, which is why the v4 removal ratchet correctly failed after merging upstream/main. - YAML: replaced the old rules:/condition:/action: PolicyDocument shape with an ACS manifest (agent_control_specification_version, policies, intervention_points), matching the exact shape of the 13 other already- migrated packs (Kenya, NDPA, CBN, BVN/NIN, NFIU, POS-geofencing, POPIA, Uganda, Tanzania, Ethiopia, and the 3 universal agent-safety packs). - Rego: appended the same "Native ACS result adapters" block used by all 13 comparison packs (acs_input_result, acs_pre_tool_call_result, acs_output_result via the shared agt_policies.acs adapter). The existing deny/escalate/audit/decision rules are completely unchanged. Verified: - All 4 manifests validate against policy-engine/spec/schema/manifest.schema.json - All 4 packs' acs_input_result/acs_pre_tool_call_result/acs_output_result manually exercised end-to-end against realistic ACS snapshots - opa check clean, 144/144 opa test pass (unchanged, since the underlying rules were never touched) - v4 removal ratchet passes (grand total 4, matches baseline — resolved by the format change itself, no baseline edits needed) - jurisdiction-router.rego already correctly wired (points at .decision, which is untouched), no changes needed Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> * fix(african-regulatory): address liamcrumm review on PR #3243 - ghana-dpa.rego, rwanda-dpa.rego, egypt-pdpl.rego: the health/special- category escalate rules matched bare "HIV" and "race" with no word boundary, so ordinary words false-positived ("archive" contains "hiv", "embrace"/"grace period"/"stack trace" contain "race"). Anchored HIV with \b and replaced bare "race" with "racial\s+origin", matching the pattern mauritius-dpa.rego:176 already used correctly. Added 6 regression tests (2 per jurisdiction). - .cspell-repo-terms.txt: removed 20 entries (CloudEvents block, Docker/ BuildKit terms, pytest identifiers) that duplicated entries already present earlier in the file — picked up during the earlier merge- conflict resolution against upstream/main, not actually new. Left hostnames/dorny in place since those aren't duplicated and are still needed. - README: test count 143 -> 150 (accounts for all regression tests added this session). Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com> --------- Signed-off-by: oluwajuwon omotayo <ginuxtechacademy@gmail.com>
Open
3 tasks done
MohammadHaroonAbuomar
added a commit
that referenced
this pull request
Sep 18, 2026
The Secret Scanning workflow has failed on every push and scheduled run on main since 2026-07-30. Its non-PR mode scans the full history, and two historic commits trip it: - 33eeccb (squash merge of #3955) carries the test fixture sk-abcdefghijklmnopqrstuvwxyz0123 in agent-governance-python/agent-os/tests/test_mcp_pii_and_response_gateway.py. The fingerprint added in that PR names the pre-squash commit 4013555, which is not in main's history. - 2c43622 (#2399) carries a placeholder JWT in the cedarling-agentmesh README. Its fingerprint was dropped from .gitleaksignore by #3444 on 2026-07-30, and the file itself was deleted by #3451, so only the historic commit remains. Add the two fingerprints for the commits that are actually on main and drop the dead pre-squash line. Verified with gitleaks 8.24.3: a full history scan of main reports no leaks with these lines and exactly the two findings without them. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com>
MohammadHaroonAbuomar
added a commit
that referenced
this pull request
Sep 24, 2026
The Secret Scanning workflow has failed on every push and scheduled run on main since 2026-07-30. Its non-PR mode scans the full history, and two historic commits trip it: - 33eeccb (squash merge of #3955) carries the test fixture sk-abcdefghijklmnopqrstuvwxyz0123 in agent-governance-python/agent-os/tests/test_mcp_pii_and_response_gateway.py. The fingerprint added in that PR names the pre-squash commit 4013555, which is not in main's history. - 2c43622 (#2399) carries a placeholder JWT in the cedarling-agentmesh README. Its fingerprint was dropped from .gitleaksignore by #3444 on 2026-07-30, and the file itself was deleted by #3451, so only the historic commit remains. Add the two fingerprints for the commits that are actually on main and drop the dead pre-squash line. Verified with gitleaks 8.24.3: a full history scan of main reports no leaks with these lines and exactly the two findings without them. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com>
MohammadHaroonAbuomar
added a commit
that referenced
this pull request
Sep 26, 2026
The Secret Scanning workflow has failed on every push and scheduled run on main since 2026-07-30. Its non-PR mode scans the full history, and two historic commits trip it: - 33eeccb (squash merge of #3955) carries the test fixture sk-abcdefghijklmnopqrstuvwxyz0123 in agent-governance-python/agent-os/tests/test_mcp_pii_and_response_gateway.py. The fingerprint added in that PR names the pre-squash commit 4013555, which is not in main's history. - 2c43622 (#2399) carries a placeholder JWT in the cedarling-agentmesh README. Its fingerprint was dropped from .gitleaksignore by #3444 on 2026-07-30, and the file itself was deleted by #3451, so only the historic commit remains. Add the two fingerprints for the commits that are actually on main and drop the dead pre-squash line. Verified with gitleaks 8.24.3: a full history scan of main reports no leaks with these lines and exactly the two findings without them. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com>
Yuvraj Singh (yuvrajsingh2428)
pushed a commit
to yuvrajsingh2428/agent-governance-toolkit
that referenced
this pull request
Oct 1, 2026
…osoft#4033) The Secret Scanning workflow has failed on every push and scheduled run on main since 2026-07-30. Its non-PR mode scans the full history, and two historic commits trip it: - 33eeccb (squash merge of microsoft#3955) carries the test fixture sk-abcdefghijklmnopqrstuvwxyz0123 in agent-governance-python/agent-os/tests/test_mcp_pii_and_response_gateway.py. The fingerprint added in that PR names the pre-squash commit 4013555, which is not in main's history. - 2c43622 (microsoft#2399) carries a placeholder JWT in the cedarling-agentmesh README. Its fingerprint was dropped from .gitleaksignore by microsoft#3444 on 2026-07-30, and the file itself was deleted by microsoft#3451, so only the historic commit remains. Add the two fingerprints for the commits that are actually on main and drop the dead pre-squash line. Verified with gitleaks 8.24.3: a full history scan of main reports no leaks with these lines and exactly the two findings without them. Signed-off-by: MohammadHaroonAbuomar <40180927+MohammadHaroonAbuomar@users.noreply.github.com> Signed-off-by: yuvrajsingh2428 <offcyuvi2428@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Removes the v4 policy language from the Python runtime and puts ACS v5 in its place. This is the first of seven PRs that finish the ACS port; the v4 syntax survives only inside the one-way migration tool by the end of the stack.
Problem
The original ACS port kept a compatibility bridge that translated v4 policy documents into ACS v5 at runtime. That left two policy languages live at once. Every adapter carried a translation path, the evaluator had to accept both shapes, and the v4 semantics quietly constrained what ACS could express. The bridge was meant to be temporary and had no removal date.
Rewriting the callers onto ACS directly is a breaking change for anyone still authoring v4 policy documents. That is intentional and is the point of the stack.
Changes
agent-governance-python/agent-os/src/agent_os/integrations/agent-governance-python/agent-os/src/agent_os/integrations/base.pyAdapterExecutionStateand the shared adapter runtime accessor that the adapters build on.agent-governance-python/agt-policies/src/agt/policies/result.pydrops theEvaluationResultcompatibility model;runtime.pyevaluates ACS manifests without the resolution shim.agent-governance-python/agent-mesh/,agent-compliance/,agent-marketplace/policy-engine/.github/workflows/demos.ymlagent-governance-python/requirements/ci-policy-test.txtpygments, whichpytest==9.0.3needs. Without it the hashed install fails, and the policy test job only runs on PRs that touch policy paths, so the gap stayed hidden..github/workflows/ci.ymlagt-policiesafter.[dev], so those tests run against this commit's ACS API rather than the last published wheel.docs/security/audits/2026-07-27-v4-policy-language-removal.mddocs/dependency-audits/2026-07-27-v4-policy-language-removal.mdpygmentspin, the ACS npm package, and the lockfiles deleted with their examples.scripts/v4_ratchet_baseline.json01_policy_enforcement_101.ipynbAgtRuntime. It was run; the pattern rules and tool-call budget give the same allow and deny decisions the notebook described. Thequickstartnotebook moves with the examples PR, and the ratchet gains.ipynbscanning in the migration-tool PR.agent-os/tests/test_adapter_enforcement.py,test_adapter_interception.py,test_governed_clients.py,test_citadel_integration.pyagent_os.integrationsat 50% line coverage.mainmeasures 87%, so this PR still lands below it; the module shrank 25% while covered lines fell 57%. The mediation contract test that survived matches adapter source for the seam strings, so it never runs the code. These 139 tests drive the runtime seam instead of per-provider payload shapes: allow, deny, transform, tool-call budget, message redaction, the governed proxies, the AutoGen handler, the ADK tool callback, and the Citadel identity and bundle paths. Coverage recovers to 63%, and the stack tracks the rest of the gap.pydantic_ai_adapter.py,openai_agents_sdk.pyhealth_check()read a wrapped-agent dict this stack removed, so every call raisedAttributeError. Both now report from the session state each kernel still keeps. Neither had a test reaching that line.langgraph_adapter.pyself._evaluatordefensively when building a graph fingerprint. That attribute is assigned nowhere here or onmain, so the path raised whenever it ran._native_adapter_runtime.py, 5 sandbox providers, 4 model adapters, 13 framework gates, 43 guarded rewritestransformverdict permits, but it permits a rewritten action and carries the replacement. Code reading bareallowedtherefore ran the original while the policy believed it was rewritten, so a redaction policy did not redact and nothing reported it. Sites that can apply the replacement rewrite fromtransformed_value; sites with nowhere to put it now gate onpermits_unchangedand refuse. The sandbox gate refuses outright, since it cannot rewrite code it is about to execute. Four censuses intest_transform_never_dropped.pykeep the class closed: a verdict consumed without applying or refusing, a replacement of the wrong shape, a write whose failure is swallowed, and a write skipped because the target takes nothing. Each was written to fail before its fix.agt-policies/tests/scenarios/test_gemini_*,test_mistral_*inputandpre_tool_callbut notoutput, while the adapter evaluates output after every call. The runtime denies an unconfigured point, so they now bind it, as the langchain and llamaindex manifests already did.BREAKING_CHANGES.mdagent-os/src/agent_os/integrations/registry.pysrc/referenced it here or onmain, it is not exported, and the stack removed its only test.Deletions dominate the diff because the bridge, its tests, and the duplicated v4 fixtures all go away.
Testing
agt-policies: 350 passed, 1 skippedagent-os: 2921 passed, 36 skipped;agent_os.integrationsline coverage 63%, up from 50% on this branch and still below 87% onmainscripts/check_v4_ratchet.py: passes, and the repo-wide v4 marker count drops from 3790 to 1415scripts/tests/test_check_v4_ratchet.py: 35 passedpython scripts/docs/check_links.py --root .: 0 new brokenscripts/ci/no-stubs.sh: cleanagent-sandbox: 499 passed, 64 skippedtests/ci: 73 passedruff check --select E,F,W --ignore E501on the 150 changed Python files reports 72 findings, all intest_cli.pyandtest_cli_extended_coverage.py. This PR edits both files but adds none of the findings:main's copies report the same 72.Stack
This is PR 1 of 7. The examples and the agentmesh integrations are split into their own PRs so this one stays under GitHub's 300-file display limit. Each merges cleanly once the one below it lands:
agt.policiesremoval