When the objective scorer returns no score, several attacks record FAILURE, which reads as "measured and defended". AttackOutcome.UNDETERMINED exists for exactly this case.
prompt_sending.py _determine_attack_outcome: if response: return FAILURE
multi_prompt_sending.py: same
red_teaming.py:411: ... if context.last_score else AttackOutcome.FAILURE
chunked_request.py: "No score returned from scorer" maps to FAILURE
This happens e.g. when a text-only scorer gets an image response. Crescendo and TAP raise in the same situation instead, so the attacks don't agree either.
PromptSendingAttack: outcome=FAILURE score=None reason='Failed to achieve objective after 1 attempts'
MultiPromptSendingAttack: outcome=FAILURE score=None reason='Failed to achieve objective'
It understates ASR for modality-mismatched runs, since unscored rows count as defended. Proposal: when a scorer is configured, a response exists and the score is None, return UNDETERMINED with "Objective scorer returned no score". Filing first since it changes reported outcomes.
When the objective scorer returns no score, several attacks record FAILURE, which reads as "measured and defended".
AttackOutcome.UNDETERMINEDexists for exactly this case.prompt_sending.py_determine_attack_outcome:if response: return FAILUREmulti_prompt_sending.py: samered_teaming.py:411:... if context.last_score else AttackOutcome.FAILUREchunked_request.py: "No score returned from scorer" maps to FAILUREThis happens e.g. when a text-only scorer gets an image response. Crescendo and TAP raise in the same situation instead, so the attacks don't agree either.
It understates ASR for modality-mismatched runs, since unscored rows count as defended. Proposal: when a scorer is configured, a response exists and the score is None, return UNDETERMINED with "Objective scorer returned no score". Filing first since it changes reported outcomes.