Repository navigation
Proposal: optional embedding/kNN prompt-injection signal from research corpus #2918
Description
Activity
🟡 Contributor Check: MEDIUM
Check Result Profile MEDIUM Credential NONE Overall MEDIUM Automated check by AGT Contributor Check.
- addedneeds-review:MEDIUMContributor check flagged MEDIUM riskContributor check flagged MEDIUM risk
on Jun 9, 2026 imran-siddique commented
on Jun 13, 2026 CollaboratorMore actionsThe embedding+kNN direction was validated by PR #2990 (merged) which added the cosine dimension check and embedding alignment work. The optional, default-off kNN signal approach described here is the right model -- it doesn't replace rules but adds a tunable evidence layer.
For anyone looking to contribute the next step: the remaining gap is connecting the kNN signal to the detection pipeline as an optional backend (ADR-0015 style). See #2957 for the content normalization RFC that would sit upstream of this.
- added 3 commits that reference this issue
on Jun 13, 2026 - addedhelp wantedExtra attention is neededExtra attention is neededand removedneeds-review:MEDIUMContributor check flagged MEDIUM riskContributor check flagged MEDIUM risk
on Jun 15, 2026 - added a commit that references this issue
on Jun 15, 2026 - addedresearchAcademic papers and researchAcademic papers and researchsecuritySecurity-related issuesSecurity-related issuesrustPull requests that update rust codePull requests that update rust code
on Jun 16, 2026 Keep this optional and default-off. The next step is an ADR-0015-style backend path with benchmark and smoke coverage.
kerberosmansour commented
on Jun 16, 2026 ContributorAuthorMore actionsThanks Ricky - The default off approach was implemented in PR #3014 -Sherif Mansour…On Tue, 16 Jun 2026 at 12:06, Ricky Gummadi ***@***.***> wrote: *Ricky-G* left a comment (microsoft/agent-governance-toolkit#2918) <#2918 (comment)> Keep this optional and default-off. The next step is an ADR-0015-style backend path with benchmark and smoke coverage. — Reply to this email directly, view it on GitHub <#2918?email_source=notifications&email_token=ADGPVQWEQ3TWIL7FM3WJ3XL5AESZVA5CNFSNUABFM5UWIORPF5TWS5BNNB2WEL2JONZXKZKDN5WW2ZLOOQXTINZRHAYDKOBTHE42M4TFMFZW63VGMF2XI2DPOKSWK5TFNZ2KYZTPN52GK4S7MNWGSY3L#issuecomment-4718058399>, or unsubscribe <https://github.com/notifications/unsubscribe-auth/ADGPVQQN7W3GFT5CRRICKZ35AESZVAVCNFSNUABGKJSXA33TNF2G64TZHMYTCNZRGEYTCNBWG45US43TOVSTWNBWGI3DGNBYGA2TFILWAI> . You are receiving this because you authored the thread.Message ID: ***@***.***>#3014 is merged and it closes the gap this issue tracked. It connects the existing local kNN/embedding signal to the
PromptInjectionDetectorpipeline as an optional, default-off, evidence-only backend in both Python (agent-os) and Rust (agentmesh), with the design recorded in ADR-0031 following the ADR-0015 pluggable-backend pattern. That matches the hypothesis here: an additive evidence layer that is default-off, tunable, and auditable, with the rules baseline unchanged.Content normalization (RFC #2957 / #2991) stays upstream of this and is tracked separately, so it is out of scope for this issue.
Closing as completed. Thanks Kerberosmansour (@kerberosmansour) for the research corpus and the experiment that motivated it. If a benchmark or smoke-coverage follow-up comes up, let us open a fresh issue against ADR-0031 rather than reopening this one.
- added a commit that references this issue
on Jun 17, 2026
Project hypothesis
This started from a narrow prompt-injection detection hypothesis:
The goal was not to replace AGT's rules or governance layer. The goal was to test whether a semantic similarity signal could add measurable value over the current rules-only baseline, especially for attacks that do not share the exact trigger words the rules expect.
To test that, I built a synthetic annotated prompt-injection research corpus, ran the existing AGT Rust prompt-injection rules as the baseline, then compared that baseline with a local embedding/kNN scoring path. The research repo is here:
https://github.com/kerberosmansour/AGT-Embeddings-Experiment
I am opening this issue first to ask whether this is something you would like contributed upstream. If it is useful, I am happy to submit a PR in the form you prefer.
What the research repo contains
The central artifact is a large annotated prompt-injection evaluation dataset:
The repo also includes smoke reproduction, metadata-only artifact validators, claims mapping, and reports.
Result snapshot
Using the migrated AGT Rust prompt-injection rules as a rules-only baseline, the research repo reports:
The conservative zero-FP operating point is the most interesting comparison: on the frozen test split, it raises observed attack catch rate from about 1% to about 14% while keeping observed benign false positives at 0 in this synthetic corpus.
Method at a glance
BAAI/bge-small-en-v1.5fastembed/onnxruntime-localqdrant/bge-small-en-v1.5-onnx-qk=5threshold_tau=0.08026763573288917What I am proposing
I am not proposing that AGT make this a default-blocking detector.
The useful contribution could be one of several shapes, depending on what maintainers want:
Important boundaries
This work is research evidence only:
The embedding signal is best viewed as optional, tunable, auditable evidence that may augment existing AGT policy or review routing.
Question for maintainers
Would you like this contributed to
microsoft/agent-governance-toolkit?If yes, I can prepare a PR and would appreciate guidance on the preferred scope: evaluation corpus only, harness/reporting only, or an optional default-off embedding signal.