Repository navigation
Add HR industry preset - #24
Conversation
|
This is in good shape. Both checks pass, the preset wiring stays inside the existing registries instead of adding an abstraction, and the tests cover the enum, the MIT categories, the BBQ/BOLD bundle, the severity gate and the CLI filter. Leaving the preset-to-policy selection gap out was the right call. That one is #23 and it changes policy loading for every existing preset, so it does not belong in a preset PR. Thanks for the assistance disclosure as well. The only thing blocking me is that it is still marked as a draft, so GitHub will not let me merge. Click Ready for review and I will squash it in. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7cc41b45e5
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
|
The second commit is a good catch and it is the right fix. Draining the first split meant a small demo sample was effectively one demographic category, so round-robin across the splits is what the demo needed to be worth running. Scoring BBQ rows against an answer-accuracy rubric with an empty scorer context is also correct: the ambiguous context is the point of the benchmark, and handing it to the factuality judge as ground truth would have produced confident nonsense. Your inline comment explaining that is exactly the kind of note that saves the next reader an hour. I checked out f7cc8a2 and ran it. The 9 tests in One thing worth naming for the record: the BBQ loader is global, so the new sampling and rubric behavior applies to any caller of Squashing and merging. Thanks for seeing this through, and thanks again for the assistance disclosure. Closes #10. |
Summary
This adds
hras the toolkit's first industry preset beyond healthcare,financial services, and government. The preset selects employment-relevant MIT
risk categories, exposes the existing BBQ and BOLD datasets as its opt-in demo
bundle, and reuses the existing fairness baseline through the current default
policy-directory behavior.
It also applies the same conservative operational defaults used by the other
high-impact presets: a red-team severity gate of 3 and a recommended 30-day
reassessment cadence. The workflow enum, filtered dataset CLI, compliance API
documentation, and README examples now recognize the new preset.
Context and user impact
Before this change, the repository already contained the BBQ and BOLD loaders
and a fairness policy baseline, but users could not select an HR preset or ask
the demo-dataset CLI for an HR bundle. The preset wiring is intentionally
distributed across the existing registries and defaults rather than introducing
a new abstraction, so existing preset behavior and public data formats remain
unchanged.
Users can now run an opt-in HR fairness/bias demo with:
The broader preset-to-policy selection gap is deliberately excluded because it
would change policy-loading behavior for every existing preset.
Validation
python3 -m pytest -q— 13 passedpython3 -m compileall -q rai_toolkit tests— passedgit diff --cached --check— passed before commitThe seven new focused tests cover enum serialization, exact MIT categories,
the BBQ/BOLD bundle, workflow scoping with the existing fairness baseline,
red-team severity, monitoring cadence, and CLI filtering without downloading
datasets.
AI assistance disclosure
This contribution was prepared with assistance from OpenAI Codex. I reviewed
the implementation, tests, documentation, and final diff and can explain or
rework the submitted changes.
Closes #10