Skip to content

fix(agent-os): prevent invisible-character evasion in conversation guardian - #4124

Open
Ricky Gummadi (Ricky-G) wants to merge 2 commits into
mainfrom
ricky-g-unicode-evasion-hardening
Open

Ricky Gummadi (Ricky-G) wants to merge 2 commits into
mainfrom
ricky-g-unicode-evasion-hardening

Conversation

@Ricky-G

@Ricky-G Ricky Gummadi (Ricky-G) commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Harden Conversation Guardian against invisible Unicode characters inserted into detection keywords, and apply the same normalization strategy to escalation, offensive-intent, and retry-loop checks. Preserve existing matches, original audit content, and scoring thresholds while adding 447 regression cases.

Related Issue

Fixes #3500.

Related prior proposal: #3501. Attribution is included below and in the implementation.

Problem & Solution

The previous normalizer removed only five zero-width characters. Soft hyphens, variation selectors, Hangul fillers, and other invisible characters could split keywords and evade word-boundary-based detection. Retry-loop checks did not normalize messages at all.

This change:

  • Covers all 4,174 Unicode 17.0.0 Default_Ignorable_Code_Point characters, plus the three interlinear annotation controls. Removal happens before NFKD, so compatibility decomposition cannot turn a filler into a surviving character.
  • Checks deletion and space-substitution views: deletion repairs split keywords, while spaces preserve boundaries between keywords.
  • Retains normalized views with and without leetspeak conversion, including a punctuation-preserving view, so combinations such as ur\u00adg3nt! and numeric errors such as 4\u034f03 remain detectable.
  • Retains the previous normalization result to prevent regressions when old joiners and newly covered separators occur together.
  • Prepares the views once per guardian message and shares the tuple across all three detectors. Each pattern contributes its weight only once, across at most eight distinct views, regardless of how many invisible characters the message contains.
  • Exports detection_texts(text) for callers that need the detection-only views. A compiled regex substitution replaces the punctuation-preserving per-character Python loop.

Verified soft-hyphen reproductions now match their unobfuscated equivalents:

Signal Clean message Same message with soft hyphens inside every word
Escalation example from #3500 0.80 0.80
Offensive-intent example from #3500 1.00 1.00
Offensive example through the guardian critical / quarantine critical / quarantine

Changes

File Change
agent-governance-python/agent-os/src/agent_os/integrations/conversation_guardian.py Property-based invisible-character coverage and shared, bounded detection views prepared once across all three detectors.
agent-governance-python/agent-os/tests/test_conversation_guardian_unicode.py 447 regression cases covering the complete Unicode snapshot, detector score parity, word boundaries, mixed evasions, retry loops, audit preservation, multilingual text, view-count bounds, single preparation, and punctuation semantics.
docs/tutorials/09-prompt-injection-detection.md Document the public helper, coverage, compatibility safeguards, numeric retry trade-off, and defense-in-depth limitations.
.cspell-repo-terms.txt Add the Unicode and normalization terms identified in review.
.cspell.json Exclude exact four-/eight-digit Unicode escape sequences without excluding surrounding prose.

Impact on Your Work

Closes the reported normalization bypass without changing existing public method signatures, detection patterns, configured thresholds, or transcript content. The public detection_texts helper is additive. No dependencies, workflow changes, or unrelated package changes are included. Invisible characters alone do not trigger an alert.

Numeric retry trade-off: 4\u200b0\u200b1 and id 40\u00ad3 normalize to existing 401/403 error matches. Even benign-looking identifiers can therefore contribute to the retry limit; the tutorial and end-to-end tests explicitly cover this behavior.

This remains a heuristic detection layer, not a guarantee against every prompt injection. An AlertAction.NONE result is not authorization; deterministic tool permissions, policy enforcement, and approval controls remain necessary.

Timeline

None.

Alternatives Considered

  • Add a few more characters: leaves the same class of omissions. The implementation uses a versioned Unicode property snapshot, independently checked by tests.
  • Strip an entire Unicode category: misses characters in other categories and can remove legitimate visible text.
  • Use only deletion: can glue adjacent keywords and lose existing matches. Multiple bounded views address that regression without enumerating combinations of every character position.
  • Replace the broader SDK normalization architecture: out of scope for this targeted guardian fix.

Type of Change

  • Bug fix (non-breaking change that fixes an issue)
  • New feature
  • Breaking change
  • Documentation update
  • Maintenance
  • Security fix

Package(s) Affected

  • agent-os
  • agent-governance-toolkit-core (packages the agent_os source tree)
  • docs / root

Testing

Validation environment: Windows, Python 3.14.7, with PYTHONPATH pointing to the local agent-governance-python\agent-os\src directory.

Check Result
Existing guardian tests and new Unicode regression suite 526 passed: 79 existing tests plus 447 new cases.
Guardian coverage with branch measurement enabled 94% combined statement/branch coverage.
Regression-first run The initial regression suite produced 319 failures before the implementation was changed.
Deterministic compatibility probe 10,000 mixed-Unicode inputs retained the previous normalized view and stayed within eight views.
Regex optimization parity probe 10,000 punctuation/Unicode inputs retained the prior character-loop punctuation view.
New test file, full configured Ruff rules Passed.
Production file, E,F,W checks excluding E501 Passed.
Changed tutorial, link and strict frontmatter checks Passed: 15 links checked, no findings.
Unicode spell exclusions JavaScript assertions verify exact escape lengths and preservation of surrounding prose. Full local cspell execution is blocked by the configured npm remote-package restriction; CI must validate the spelling changes.
git diff --check Passed.

Unit Testing

The new tests independently enumerate the Unicode property snapshot, verify every code point is removed inside a keyword, cover all escalation/offensive pattern groups, and assert exact score and matched-pattern parity. Additional cases verify critical/quarantine output, obfuscated retry-loop breaking, unchanged original hashes/previews, benign multilingual/emoji input, non-duplicated scoring, public helper behavior, and exactly one view preparation per guardian message.

Reproduce the focused suite from the repository root in PowerShell:

$env:PYTHONPATH = "$PWD\agent-governance-python\agent-os\src"
python -m pytest agent-governance-python\agent-os\tests\test_conversation_guardian.py agent-governance-python\agent-os\tests\test_conversation_guardian_unicode.py -o addopts= -q --tb=short --cov=agent_os.integrations.conversation_guardian --cov-branch --cov-report=term-missing

Manual Testing

Reran the issue's escalation and offensive examples with soft hyphens inserted throughout each keyword and verified score parity and quarantine behavior.

Review follow-up performance probe: five runs of the same 200,000-character input with 10% invisibles on the same local environment. Median analyze_message time decreased from approximately 1,769 ms to 907 ms; standalone retry classification decreased from approximately 420 ms to 229 ms. These are local comparisons, not universal latency guarantees.

The broader Agent OS package suite, excluding test_mcp_server.py as directed by the package instructions, reports 3,548 passed, 44 skipped, 1 failed, 8 errors. The same non-passing cases were previously reproduced using an untouched archive of base commit 4b9f41ff:

  • test_policy_gen.py::test_strict_runtime_allows_reads_and_denies_unknown: the read operation returns deny instead of allow.
  • Four parameterizations of test_credential_redactor.py::test_trailing_lookahead_patterns_handle_adversarial_input_quickly: pytest's generated test IDs exceed Windows' 32,767-character environment-variable limit, producing setup and teardown errors.

Full configured production lint retains three pre-existing diagnostics (UP015, B905, C401). MyPy retains two pre-existing missing annotations in get_stats() (by_action, by_severity), also verified against the base source. No unrelated fixes are bundled into this PR.

Checklist

  • I have linked a related issue above
  • My code follows all project style checks without findings (targeted checks pass; pre-existing findings documented above)
  • I have added tests that prove the fix works
  • All new and existing package tests pass (unrelated baseline failures documented above)
  • I have updated documentation as needed
  • I have signed the Microsoft CLA
  • I have personally reviewed the changes at 20cbd5a7 and am satisfied with them.
  • I understand and can explain the meaningful changes and their trade-offs.
  • This contribution does not implement patent-pending or patent-encumbered techniques.
  • This contribution does not require an NDA or separate licensing agreement to understand or use.
  • The AI tools used have terms compatible with the MIT License.

The commits include DCO signoffs and Copilot co-author trailers. The five author confirmations above were explicitly provided by Ricky Gummadi on September 24, 2026.

Attribution & Prior Art

  • This contribution does not contain code copied or derived from other projects without attribution
  • Sources that inspired the design are credited in code comments and below

Prior art / related work: LHMQ878 reported the invisible-character bypass in #3500 and proposed property-based coverage and deletion/space-substitution matching in #3501. This implementation builds on those ideas and review findings, adds shared retry-loop handling, punctuation/numeric preservation, legacy-match compatibility, bounded view generation, and expanded regression coverage.

Unicode coverage was verified against Unicode 17.0.0 DerivedCoreProperties.txt. The test snapshot is versioned explicitly; it does not automatically track future Unicode releases.

Cover Unicode default-ignorables before compatibility normalization and preserve bounded detection views for word boundaries, punctuation, numeric errors, and legacy matches across all guardian detectors.

Add 435 regression cases and document detection-only normalization and defense-in-depth limits. Fixes #3500; credits prior work in #3501.

Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation tests size/L Large PR (< 500 lines) labels Sep 24, 2026
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@github-actions

Copy link
Copy Markdown

PR Review Summary

Check Status Details
🔍 Code Review ⚠️ Missing No current-run comment
🛡️ Security Scan ⚠️ Missing No current-run comment
🔄 Breaking Changes ⚠️ Missing No current-run comment
📝 Docs Sync ⚠️ Missing No current-run comment
🧪 Test Coverage ⚠️ Missing No current-run comment

Verdict: ⚠️ AI review incomplete; ready for human review

AI review comments are untrusted advisory output. The summary reports workflow-generated completion status only, not model-authored pass/fail claims.

@MohammadHaroonAbuomar MohammadHaroonAbuomar left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • .cspell-repo-terms.txt: Spell Check is red on 29 added-line words: real terms (Bidi, Halfwidth, ignorables, maketrans, homoglyphs, leetspeak, NFKD, invisibles, midword, LHMQ) plus fragments cspell carves out of the \u00ad-escaped test literals (scalate, privileg, rsonate, nied, rized, and so on). Add the real terms to the terms file and an ignoreRegExpList entry in .cspell.json for \\u[0-9a-fA-F]{4} and \\U[0-9a-fA-F]{8} escapes, which the list does not have yet. The detection change itself is verified: all 27 probed in-scope code points that main misses (soft hyphen, variation selectors, Hangul fillers, bidi controls, tag characters, Mongolian vowel separator) are caught at the clean-text score; the strip table matches the 4,174 Unicode 17 Default_Ignorable code points exactly; 30,000 fuzzed inputs never score lower than main and main's matched set is always a subset; family and flag emoji, Korean fillers, soft-hyphen prose and RLM text stay benign with transcript hash and preview unchanged.
  • PR body checklist: the CLA box, "I can explain every meaningful change", "I have reviewed the specific AI-produced changes" and the IP boxes are all unchecked, and the body says they were left unchecked rather than inferred. This is a Copilot-written security change; please attest before merge.

Comment thread agent-governance-python/agent-os/tests/test_conversation_guardian_unicode.py Outdated
Comment thread docs/tutorials/09-prompt-injection-detection.md
Prepare detection views once per guardian message and use regex substitution for punctuation-preserving leetspeak. Expose detection_texts, verify public behavior, and document numeric retry classification.

Add Unicode spelling terms and precise escape exclusions; construct intentionally obfuscated test words from readable text.

Signed-off-by: Ricky Gummadi <ricky.gummadi@outlook.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@github-actions github-actions Bot added size/XL Extra large PR (500+ lines) and removed size/L Large PR (< 500 lines) labels Sep 24, 2026

@MohammadHaroonAbuomar MohammadHaroonAbuomar left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • Code, tests, docs and spelling are all verified at 20cbd5a: the views are computed once per message and shared across the three detectors, the punctuation view is a regex with zero mismatches against the old loop over 200,000 random strings, detection_texts is public, the tutorial states the numeric retry trade-off, the spell check is green, and the detection results from round one still hold (27 of 27 in-scope code points caught, benign emoji, Korean, RTL and soft-hyphen prose unchanged, transcript hash identical, 5,000 fuzzed inputs never lower than main). One item is left and it is not one a reviewer can supply. The body still leaves unchecked: "I can explain every meaningful change and its tradeoffs", "I have reviewed the specific AI-produced changes before submission", and the three IP boxes, and says they were left for the reviewer to request. Those are the author's attestations. This is a Copilot-written change to a security control, and it needs a person on the submitting side to say they read it. Please tick them yourself, or a code owner can decide to waive them; I will not approve without one of the two.

@Ricky-G

Copy link
Copy Markdown
Contributor Author

Those are the author's attestations.

Thanks for the careful review and independent validation. Confirming as the author: I have personally reviewed the changes at 20cbd5a7, understand the implementation and its trade-offs, and am satisfied with the change.

I also confirm that this contribution does not implement patent-pending or patent-encumbered techniques, requires no NDA or separate licensing agreement, and that the AI tools used have terms compatible with the MIT License. I have recorded these five confirmations in the existing PR checklist.

The focused guardian suite was rerun: all 526 tests pass, with 94% coverage. There have been no code changes since your technical review.

@Ricky-G

Copy link
Copy Markdown
Contributor Author

Hi MohammadHaroonAbuomar, thanks again for the careful review and independent validation. The technical feedback has been addressed, all three review threads are resolved, and the requested personal-review and IP/licensing confirmations are recorded in my earlier comment. CI is green, with no code changes since your last technical review.

Could you please take another look when you have a chance and, if everything looks good, update your review to approval? I appreciate your help getting this fix over the line. Thank you!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation size/XL Extra large PR (500+ lines) tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

conversation_guardian: invisible characters inside a keyword bypass all detection in normalize_text

2 participants