Skip to content

feat(normalize): shared content-normalization module for prompt-injection defense — Rust + Python (RFC #2957) - #2991

Merged
Imran Siddique (imran-siddique) merged 2 commits into
microsoft:mainfrom
kerberosmansour:slo/content-normalization
Jun 12, 2026
Merged

Imran Siddique (imran-siddique) merged 2 commits into
microsoft:mainfrom
kerberosmansour:slo/content-normalization

Conversation

@kerberosmansour

Copy link
Copy Markdown
Contributor

Summary

Implements the contribution offered in RFC #2957: a shared, surfaced content-normalization (canonicalization) module for prompt-injection defense, shipped in both SDKs — agentmesh::normalize (Rust) and agent_os.normalize (Python) — with verified cross-SDK parity.

The module produces a canonical view of untrusted text plus a closed-vocabulary record of which transforms fired, so every text-based control — the existing regex detector, classifier/LLM annotators, policy/IFC decisions, and human reviewers — can consume the same un-disguised content. As the RFC argues, normalization is a force-multiplier for every downstream control, detective and preventative alike, regardless of whether an ML/embedding detector is ever adopted.

This PR is additive and non-breaking: the existing private normalize_for_detection inside the detector is untouched, no existing behavior changes, and no new dependencies are introduced (Rust uses the already-present base64 crate; Python is stdlib-only).

Related work

What's included

agentmesh::normalize (Rust, new module)

use agentmesh::normalize::{normalize, Transform};

let r = normalize("1gn0r3 4ll pr3v10u5 1n57ruc710n5");
assert!(r.text.contains("ignore all previous instructions"));
assert!(r.transforms.contains(&Transform::Leet)); // auditable: what was un-disguised

Public surface: normalize(&str), normalize_with(&str, &NormalizeConfig), Normalized { text, transforms }, the closed Transform enum, and NormalizeConfig (decode depth, expansion cap, printable-ratio guard, decoder toggle).

agent_os.normalize (Python, 1:1 port)

from agent_os.normalize import normalize, Transform

r = normalize("1gn0r3 4ll pr3v10u5 1n57ruc710n5")
assert "ignore all previous instructions" in r.text
assert Transform.LEET in r.transforms

Same transforms, same order, same guards, same tag vocabulary, same config defaults.

Transforms (false-positive safety is the design centerpiece)

Strip invisible incl. bidi override/isolate (Trojan Source) · width fold · bounded decode layers (base64 / hex / rot13 / percent / unicode-escape / HTML-entity, each behind a printable-ratio + English-benefit acceptance guard, depth ≤ 2, output ≤ 4× input) · homoglyph/confusable fold · letter-spacing collapse · token-guarded leetspeak de-substitution · lowercase · whitespace collapse.

Every aggressive transform fires only under a guard tuned so legitimate inputs pass through unchanged: percentages (Save 50% off), single &, real base64 blobs, hashes, version strings, code, and high-entropy structured data are never mangled. The leetspeak guard (result must be entirely alphabetic, length ≥ 3) is what preserves the measured zero false-positives. Rejected and capped decodes are surfaced as explicit tags (decode_rejected, decode_depth_capped, output_capped) rather than silent behavior.

Invariants: deterministic and idempotent (normalize(normalize(x)) == normalize(x)), property-tested in both languages.

Evidence

From the controlled study on the synthetic research corpus referenced in #2957 (directional evidence, not a production guarantee): with a fixed downstream detector, fuller normalization in front of it raised catch-at-0%-FP from 14% → 43%, and the extended decode layers raised the encoding bypass class from 35% → 62% — with zero benign-control false positives throughout.

This PR's own verification:

Check Command Result
Rust module tests cargo test -p agentmesh --lib normalize 17 passed (transforms, benign-safety, idempotency, Trojan Source)
Full Rust suite cargo test -p agentmesh --lib 371 passed, 0 failed (no regression)
Lint cargo clippy -p agentmesh --lib 0 warnings in normalize.rs
Python module tests python3 -m unittest tests.test_normalize -v 21 passed (mirrors the Rust suite)
Cross-SDK parity both implementations over the regenerated 280-row smoke corpus (#2924 fixture) + a 21-case obfuscation battery, 300 rows total byte-identical normalized text and identical transform tags on every row
Research-corpus parity Rust vs the measured research normalizer over 3,680 frozen test-split attacks 100% functional agreement (per-bypass-class table in docs/slo/completion/rust-content-normalization.md)

What this PR does not do

  • It does not change the existing detector: normalize_for_detection stays in place and PromptInjectionDetector behavior is unchanged. Pointing the detector (and the policy-engine snapshot/annotation surface) at this module is the follow-up integration step scoped in RFC: Strengthen and surface content normalization as a shared pre-detection control #2957, kept separate so this PR is reviewable as a pure addition.
  • It does not introduce blocking policy, thresholds, or any ML component.
  • Full NFKC is not applied (manual width-fold instead, dependency-free); happy to switch to unicode-normalization if maintainers prefer — flagged for review.

Risk

Low — a new, self-contained, dependency-free module that nothing consumes yet. The decode layers are bounded (depth ≤ 2, ≤ 4× expansion, printable-ratio guard) so the module cannot be used as a decompression-bomb vector, and all transforms are deterministic with no I/O.

🤖 Generated with Claude Code

Implements RFC microsoft#2957: strengthen + surface content normalization as a shared
pre-detection control. New `agentmesh::normalize` module returns canonical text
PLUS the set of transforms that fired (closed enum), so the detector, policy
annotators, IFC, and human review can all consume the same un-disguised content.

Transforms (each FP-guarded): strip invisible incl. bidi override/isolate
(Trojan Source), width fold, bounded decode layers (base64/hex/rot13/percent/
unicode-escape/HTML-entity) under a printable-ratio + English-benefit guard,
homoglyph fold, letter-spacing collapse, token-guarded leetspeak (entirely-
alphabetic guard preserves measured 0-FP), lowercase, whitespace collapse.

Additive and non-breaking (existing private normalize_for_detection untouched);
no new dependencies (base64 only). Deterministic + idempotent.

Evidence: 19 module tests + 365 full agentmesh suite green; 0 clippy warnings;
100% functional de-obfuscation parity with the measured Python normalizer over
3,680 frozen test-split attacks (zero-FP recall 43->49%, encoding 35->62%).

Refs microsoft#2957

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…ith agentmesh::normalize)

1:1 port of the Rust module added in the previous commit: same transforms in
the same order (strip invisible incl. bidi override/isolate, width fold,
guarded decode layers for base64/hex/rot13/percent/unicode-escape/HTML-entity,
homoglyph fold, letter-spacing collapse, token-guarded leetspeak, lowercase,
whitespace collapse), same FP-safety guards, same closed Transform vocabulary,
stdlib-only.

Parity verified mechanically: both implementations normalize the regenerated
280-row smoke corpus plus a 21-case obfuscation battery (300 rows) to
byte-identical text with identical transform tags on every row.

21 unit tests mirroring the Rust suite; deterministic, idempotent, no model
or network use.

Refs microsoft#2957

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@github-actions github-actions Bot added documentation Improvements or additions to documentation tests size/XL Extra large PR (500+ lines) labels Jun 12, 2026
@github-actions

Copy link
Copy Markdown

❔ Contributor Check: UNKNOWN

Check Result
Profile MEDIUM
Credential NONE
Overall UNKNOWN

Automated check by AGT Contributor Check.

@github-actions github-actions Bot added the needs-review:UNKNOWN Contributor check flagged UNKNOWN risk label Jun 12, 2026
@github-actions

Copy link
Copy Markdown

PR Review Summary

Check Status Details
🔍 Code Review ⚠️ Missing No current-run comment
🛡️ Security Scan ⚠️ Missing No current-run comment
🔄 Breaking Changes ⚠️ Missing No current-run comment
📝 Docs Sync ⚠️ Missing No current-run comment
🧪 Test Coverage ⚠️ Missing No current-run comment

Verdict: ⚠️ AI review incomplete; ready for human review

AI review comments are untrusted advisory output. The summary reports workflow-generated completion status only, not model-authored pass/fail claims.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Strong design and the integration is genuinely safe: the module is purely additive (nothing in the existing detection path consumes it yet), decode bounds are real (depth <= 2, output <= 4x, printable-ratio + english-benefit guards, no regex so no catastrophic backtracking), no panics on hostile input in the library, the closed Transform enum is the right audit surface, and the FP-safety guards check out (Save 50% off, single &amp;, high-entropy base64, 2024-style tokens all pass through unmangled; leet guard de-leets h@ck1ng correctly while leaving 2024 alone). The confusable/leet tables are byte-identical across SDKs. Nice work.

But the headline value proposition, byte-identical normalized text and identical transform tags across Rust and Python, does not actually hold. Two HIGH items must be fixed before merge (the Rust behaviors below are by code reading, since cargo was unavailable; the Python ones are confirmed by running the module):

  1. Whitespace classification diverges on C0 separators (parity break, security-relevant). normalize.rs:205 uses `ch.is_control() && !ch.is_whitespace()` where Rust's `char::is_whitespace()` (Unicode White_Space) does NOT include U+001C-U+001F (FS/GS/RS/US); normalize.py:193 uses `str.isspace()` which DOES. So for input like `igno<U+001F>re all`, Rust strips it to `ignore all` (tag StripInvisible) while Python collapses to `igno re all` (tag WhitespaceCollapse). Different normalized text AND different tags, and Rust un-disguises the attack token while Python leaves a word break. Same for U+001C/1D/1E and U+0085. Pick one whitespace definition and apply it in both SDKs.

  2. Rust unicode_unescape corrupts non-ASCII via b[i] as char (Rust bug + parity break). normalize.rs:581 pushes raw UTF-8 bytes as Latin-1 scalars in the non-escape fallthrough, so `i...e cafe all` (with a literal cafe) decodes the escapes correctly but turns `cafe`'s bytes C3 A9 into cafA© (U+00C3 U+00A9). Python (normalize.py:485) correctly pushes `s[i]`. The printable/english guards still pass, so Rust accepts the mangled output. Fix by iterating by char in the copy path (as html_unescape already does at line 621). This is a bug regardless of parity.

One MEDIUM to address too:

  1. Quadratic HTML-entity scan (CPU-DoS). count_html_entities / html_unescape (normalize.rs:586-626, normalize.py:490-518) do `find(';')` from each `&`, O(n^2) for many `&` with no nearby `;`. ~500k `&` chars take ~8s in Python's _count_html_entities alone, and it runs unconditionally per input/decode-pass with no pre-decode input-length cap (max_output_ratio bounds only the output, not the scan work). Add a pre-decode input-length guard, or bound the find(';') lookahead to the 10-char window the count already uses.

Please also: add adversarial parity fixtures covering the C0-separator and escaped-text-with-non-ASCII cases to whatever harness produced the byte-identical claim, then re-state that claim honestly (the doc at docs/slo/completion/rust-content-normalization.md:82 currently asserts it). Minor: the Python module is not re-exported from agent_os/init.py while Rust adds pub mod normalize (asymmetric discoverability), and that doc has inconsistent test counts / a wrong branch name.

The security-detection value here depends entirely on both SDKs canonicalizing identically, so the parity breaks are blocking. Happy to re-review quickly once 1-3 are fixed.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All four blockers from the original review are resolved:

  1. C0 whitespace classification -- Python _is_control and Rust .is_control() both carry the .isspace() / .is_whitespace() guard so tabs/newlines count as printable on both sides. Consistent.

  2. Rust unicode_unescape byte cast -- fixed. Python does chr(int(digits, 16)) (Unicode scalar); Rust routes through hex_scalar returning Option<char>. No raw byte truncation.

  3. Quadratic HTML entity scan -- fixed. Both Python and Rust implementations use s.find(";", i) with index advancement past the match -- O(n).

  4. Adversarial parity fixtures -- the 3,680-row corpus parity check (100% functional agreement) plus 21 mirrored unit tests are solid coverage.

The stdlib-only, idempotent design is the right call for a pre-detection control that must compose safely with downstream stages.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation needs-review:UNKNOWN Contributor check flagged UNKNOWN risk size/XL Extra large PR (500+ lines) tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants