Repository navigation
fix(scorers): preserve default LLM judge categories - #30
Conversation
|
This is the right fix and the right size. Using the existing Verified locally: your 3 tests pass, the PR head suite is green at 9, and merging this into current main after #24 gives 18 passing. It merges cleanly without a rebase. Worth stating for anyone reading this later, since it is a behavior change even though it is the repair: any judge constructed directly without an explicit category now reports its real MIT category instead of an empty string. So pipeline summaries key on Thanks for the clean writeup and for catching this. Squashing and merging. Fixes #28. |
Summary
BaseScorerretain an LLM judge subclass's declared category when construction does not supply an overrideFixes #28.
Why
LLMJudgeScorer.__init__previously defaultedcategoryto""and always passed that value toBaseScorer. BecauseBaseScorerusesNoneto mean “keep the class attribute,” direct construction replaced categories such asFactualityJudge.category == "MIT-3.1"with an empty string.Changing the default to
Noneuses the existing sentinel contract. Callers that explicitly pass a category—includingcategory=""—retain the previous override behavior.Validation
uv run --extra dev pytest tests/test_llm_judge_categories.py -q— 3 passeduv run --extra dev pytest tests/test_groundedness_scorer.py -q— 6 passeduv run --extra dev pytest -q— 9 passeduv run python -m compileall -q rai_toolkit tests— passeduvx --from 'reuse[charset-normalizer]' reuse lint— REUSE 3.3 compliantThe existing test run reports 11
asyncio.iscoroutinefunctiondeprecation warnings fromrai_toolkit/_tracing.py; this change does not introduce or modify them.AI assistance disclosure
AI tools assisted with repository search, implementation drafting, and an independent fresh-context review. I directed the work, reproduced the bug on the pinned upstream revision, reviewed the source and final two-file diff, added the compatibility edge case raised during review, and ran the validation listed above.