Skip to content

fix(scorers): preserve default LLM judge categories - #30

Merged
knisar merged 1 commit into
wandb:mainfrom
YusefSyed:fix/preserve-llm-judge-category
Sep 2, 2026
Merged

knisar merged 1 commit into
wandb:mainfrom
YusefSyed:fix/preserve-llm-judge-category

Conversation

@YusefSyed

Copy link
Copy Markdown
Contributor

Summary

  • let BaseScorer retain an LLM judge subclass's declared category when construction does not supply an override
  • preserve both non-empty and explicit empty-string category overrides
  • add focused regression coverage for all three constructor paths

Fixes #28.

Why

LLMJudgeScorer.__init__ previously defaulted category to "" and always passed that value to BaseScorer. Because BaseScorer uses None to mean “keep the class attribute,” direct construction replaced categories such as FactualityJudge.category == "MIT-3.1" with an empty string.

Changing the default to None uses the existing sentinel contract. Callers that explicitly pass a category—including category=""—retain the previous override behavior.

Validation

  • uv run --extra dev pytest tests/test_llm_judge_categories.py -q — 3 passed
  • uv run --extra dev pytest tests/test_groundedness_scorer.py -q — 6 passed
  • uv run --extra dev pytest -q — 9 passed
  • uv run python -m compileall -q rai_toolkit tests — passed
  • uvx --from 'reuse[charset-normalizer]' reuse lint — REUSE 3.3 compliant

The existing test run reports 11 asyncio.iscoroutinefunction deprecation warnings from rai_toolkit/_tracing.py; this change does not introduce or modify them.

AI assistance disclosure

AI tools assisted with repository search, implementation drafting, and an independent fresh-context review. I directed the work, reproduced the bug on the pinned upstream revision, reviewed the source and final two-file diff, added the compatibility edge case raised during review, and ran the validation listed above.

@knisar

knisar commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator

This is the right fix and the right size. Using the existing None sentinel rather than adding a new one keeps category consistent with how BaseScorer already handles name, description and threshold, and keeping the explicit category="" override working means nobody's existing call site changes meaning. Three tests for three constructor paths is exactly the coverage this needed.

Verified locally: your 3 tests pass, the PR head suite is green at 9, and merging this into current main after #24 gives 18 passing. It merges cleanly without a rebase.

Worth stating for anyone reading this later, since it is a behavior change even though it is the repair: any judge constructed directly without an explicit category now reports its real MIT category instead of an empty string. So pipeline summaries key on MIT-3.1 rather than falling back to the class name, several direct judges in the same category now aggregate together, category-based policy triggers and composite weights start working, and trace names pick up the category suffix. No report field or type changes. Where same-category judges have different assessed-row counts, their now-correct co-aggregation can move an overall score and not just its label. That is what #28 was about.

Thanks for the clean writeup and for catching this. Squashing and merging. Fixes #28.

@knisar
knisar merged commit 3a31767 into wandb:main Sep 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LLMJudgeScorer.__init__ wipes the class-level category when constructed directly

2 participants