Skip to content

FIX: keep Prompt Shield metadata in a shape the Score model accepts - #3045

Merged
Roman Lutz (romanlutz) merged 3 commits into
microsoft:mainfrom
feiiiiii5:fix/prompt-shield-metadata
Oct 9, 2026
Merged

Roman Lutz (romanlutz) merged 3 commits into
microsoft:mainfrom
feiiiiii5:fix/prompt-shield-metadata

Conversation

@feiiiiii5

@feiiiiii5 Chen Yufeiyang (feiiiiii5) commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Description

FIX: PromptShieldScorer raised a ValidationError for every real Prompt Shield response, so the scorer could not be used against the endpoint at all.

_score_piece_async parses the endpoint body and puts it in score_metadata:

# pyrit/score/true_false/prompt_shield_scorer.py:88
meta = json.loads(response)
...
score_metadata=meta,  # type: ignore[ty:invalid-argument-type]

Score.score_metadata is a flat dict[str, str | int | float] (pyrit/models/score/score.py:99), and the real body is nested — {"userPromptAnalysis": {...}, "documentsAnalysis": [{...}]}, which is the shape the scorer's own test fixture uses. So Score(...) fails with six validation errors, and score_async wraps that into RuntimeError: Error in scorer PromptShieldScorer: ...:

>>> await scorer.score_text_async("hello")
RuntimeError: Error in scorer PromptShieldScorer: 6 validation errors for Score
score_metadata.userPromptAnalysis.str  Input should be a valid string
...

The scorer was already passing the parsed response when Score changed from a regular Python class to a validating Pydantic model in b3b018ff16 (#1891). That migration began enforcing the existing flat metadata type at runtime, but this scorer was not updated. The endpoint still returns valid JSON; it is the nested Python object produced by json.loads() that violates the metadata contract. The symmetric construction path in the same package already does it right: _build_unvalidated_score normalises metadata to primitives and drops anything else ("Unrecognized metadata shape; drop to avoid downstream errors", pyrit/score/response_handler.py:154-162).

What changed: the response is stored as the JSON text it arrived as: score_metadata={"raw": response}. This retains the old fallback key while satisfying the model's flat metadata contract. The parser docstring documents access through score_metadata["raw"], and callers can use json.loads() to recover the response. The # type: ignore is gone because the value now satisfies the model.

Related: #2925 fixed the same class of problem for SelfAskLikertScorer's metadata; this is the PromptShieldScorer instance of it.

Tests and Documentation

tests/unit/score/test_prompt_shield_scorer.py::test_prompt_shield_scorer_metadata_is_the_response_text drives score_text_async against a mocked PromptTarget that returns the real-shaped body, and asserts the score and the metadata survive the Score model, and that json.loads on the stored value round-trips the body. It also checks that the raw metadata key and response text survive SQLite storage/read-back. It fails on main with the ValidationError above and passes on this branch.

pytest tests/unit/score/test_prompt_shield_scorer.py -q      3 passed
pytest tests/unit/score -q                                   42 failed, 3041 passed, 51 skipped

The 42 failures are identical on main (3d279a84) with the same names — they come from optional dependencies this environment does not have (transformers, and similar), not from this change. ruff check and ruff format --check (0.16.10) are clean on both changed files. I did not run JupyText or the doc notebooks, and no documentation page changes were needed. The scorer docstring now names the raw metadata key explicitly.

Validation of the raw key follow-up:

  • uv run --frozen --no-sync pytest tests\unit\score\test_prompt_shield_scorer.py -q: passed, 3 tests.
  • uv run --frozen --no-sync ruff check pyrit\score\true_false\prompt_shield_scorer.py tests\unit\score\test_prompt_shield_scorer.py: passed.
  • uv run --frozen --no-sync ruff format --check pyrit\score\true_false\prompt_shield_scorer.py tests\unit\score\test_prompt_shield_scorer.py: passed.
  • uv run --frozen --no-sync ty check pyrit\score\true_false\prompt_shield_scorer.py tests\unit\score\test_prompt_shield_scorer.py: passed.
  • Commit-time pre-commit hooks, including the production type check: passed.

PromptShieldScorer stored the parsed endpoint body in score_metadata, but Score
only accepts flat str/int/float values, so every real Prompt Shield response
raised a ValidationError that score_async re-raised as a RuntimeError. Store the
response as the JSON text it arrived as, which is what the scorer's docstring
already says the metadata attribute holds.
@romanlutz Roman Lutz (romanlutz) self-assigned this Oct 9, 2026
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
@romanlutz
Roman Lutz (romanlutz) merged commit f1d610f into microsoft:main Oct 9, 2026
55 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants