Skip to content

FIX: keep the response handler's metadata in SelfAskLikertScorer - #2925

Merged
Roman Lutz (romanlutz) merged 1 commit into
microsoft:mainfrom
feiiiiii5:fix/likert-preserve-response-metadata
Sep 30, 2026
Merged

Roman Lutz (romanlutz) merged 1 commit into
microsoft:mainfrom
feiiiiii5:fix/likert-preserve-response-metadata

Conversation

@feiiiiii5

Copy link
Copy Markdown
Contributor

Description

SelfAskLikertScorer drops the metadata that the response handler parsed out of the judge's reply.

UnvalidatedScore.to_score installs the parsed metadata on the Score (pyrit/models/score/score.py:406), and then _convert_score replaces that dict:

# pyrit/score/float_scale/self_ask_likert_scorer.py:472
score = unvalidated.to_score(...)                     # installs score_metadata
score.score_metadata = {"likert_value": ...}           # throws it away

So a judge that reports its own metadata loses it, and only likert_value survives on the persisted score:

judge reply: {"score_value": "3", ..., "metadata": {"verdict_confidence": 0.9, "raw_judge_output": "level 3"}}

SelfAskLikertScorer     -> 0.75  {'likert_value': 3}
SelfAskScaleScorer      -> 0.75  {'verdict_confidence': 0.9, 'raw_judge_output': 'level 3'}

Both rows are the same payload through two scorers in the same family. The other two siblings, SelfAskGeneralFloatScaleScorer and InsecureCodeScorer, also keep it.

response_handler is a documented argument of this scorer, so a caller-supplied handler that reports metadata could never have it survive. The fix extends the dict instead of replacing it, which is the merge idiom already used in true_false/regex/agent_threat_rules_scorer.py:207 and true_false/wildguard_scorer.py:360. The loss also propagates: FloatScaleThresholdScorer forwards aggregate_score.metadata into the score it builds, so the keys are missing there too.

Tests and Documentation

Test: pytest tests/unit/score/test_self_ask_likert.py -q — the new case fails on ea9d0b43 with assert {'likert_value': 3} == {'verdict_confidence': 0.9, 'raw_judge_output': 'level 3', 'likert_value': 3} and passes on this branch, where the file reads 44 passed. The two existing metadata assertions at :102 and :129 use replies with no metadata key, so they pass either way — with nothing to merge, the result is still exactly {"likert_value": N}. pytest tests/unit/score tests/unit/scenario -q gives 4199 passed, 43 skipped. ruff check and ruff format --check are clean on both changed files. Not run: ty (its pre-commit hook needs the project's --extra all venv), and JupyText (no notebooks or documentation pages changed).

Details

unvalidated.score_metadata is typed dict | None, so the merge guards for None the same way the sibling call sites do (**(x.score_metadata or {})).

This also applies to the judgment-replay path, which routes through the same _convert_score.

`_convert_score` assigned `score.score_metadata` outright, replacing the
dict that `UnvalidatedScore.to_score` had just installed from the judge's
reply. Any metadata the response handler parsed was therefore dropped from
the persisted score, leaving only `likert_value`.

`response_handler` is a documented argument, so a caller-supplied handler
that reports metadata could never have it survive. The sibling float-scale
scorers -- `SelfAskScaleScorer`, `SelfAskGeneralFloatScaleScorer` and
`InsecureCodeScorer` -- all leave the parsed metadata in place for the same
payload. Extending the dict instead of replacing it matches them, and
matches the merge idiom used elsewhere in the score package.
@romanlutz
Roman Lutz (romanlutz) added this pull request to the merge queue Sep 30, 2026
Merged via the queue into microsoft:main with commit 3b89acb Sep 30, 2026
49 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants