Skip to content

Add HR industry preset - #24

Merged
knisar merged 2 commits into
wandb:mainfrom
denis-samatov:feat/hr-industry-preset
Sep 2, 2026
Merged

knisar merged 2 commits into
wandb:mainfrom
denis-samatov:feat/hr-industry-preset

Conversation

@denis-samatov

Copy link
Copy Markdown
Contributor

Summary

This adds hr as the toolkit's first industry preset beyond healthcare,
financial services, and government. The preset selects employment-relevant MIT
risk categories, exposes the existing BBQ and BOLD datasets as its opt-in demo
bundle, and reuses the existing fairness baseline through the current default
policy-directory behavior.

It also applies the same conservative operational defaults used by the other
high-impact presets: a red-team severity gate of 3 and a recommended 30-day
reassessment cadence. The workflow enum, filtered dataset CLI, compliance API
documentation, and README examples now recognize the new preset.

Context and user impact

Before this change, the repository already contained the BBQ and BOLD loaders
and a fairness policy baseline, but users could not select an HR preset or ask
the demo-dataset CLI for an HR bundle. The preset wiring is intentionally
distributed across the existing registries and defaults rather than introducing
a new abstraction, so existing preset behavior and public data formats remain
unchanged.

Users can now run an opt-in HR fairness/bias demo with:

rai assess my_pkg:build_model --preset hr --demo-datasets --output hr-report.json

The broader preset-to-policy selection gap is deliberately excluded because it
would change policy-loading behavior for every existing preset.

Validation

  • python3 -m pytest -q — 13 passed
  • python3 -m compileall -q rai_toolkit tests — passed
  • git diff --cached --check — passed before commit

The seven new focused tests cover enum serialization, exact MIT categories,
the BBQ/BOLD bundle, workflow scoping with the existing fairness baseline,
red-team severity, monitoring cadence, and CLI filtering without downloading
datasets.

AI assistance disclosure

This contribution was prepared with assistance from OpenAI Codex. I reviewed
the implementation, tests, documentation, and final diff and can explain or
rework the submitted changes.

Closes #10

@denis-samatov denis-samatov changed the title [codex] Add HR industry preset Add HR industry preset Sep 1, 2026
@knisar

knisar commented Sep 1, 2026

Copy link
Copy Markdown
Collaborator

This is in good shape. Both checks pass, the preset wiring stays inside the existing registries instead of adding an abstraction, and the tests cover the enum, the MIT categories, the BBQ/BOLD bundle, the severity gate and the CLI filter.

Leaving the preset-to-policy selection gap out was the right call. That one is #23 and it changes policy loading for every existing preset, so it does not belong in a preset PR.

Thanks for the assistance disclosure as well.

The only thing blocking me is that it is still marked as a draft, so GitHub will not let me merge. Click Ready for review and I will squash it in.

@denis-samatov
denis-samatov marked this pull request as ready for review September 2, 2026 06:35

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 7cc41b45e5

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread rai_toolkit/examples/registry.py
Comment thread rai_toolkit/examples/registry.py
@knisar

knisar commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator

The second commit is a good catch and it is the right fix. Draining the first split meant a small demo sample was effectively one demographic category, so round-robin across the splits is what the demo needed to be worth running. Scoring BBQ rows against an answer-accuracy rubric with an empty scorer context is also correct: the ambiguous context is the point of the benchmark, and handing it to the factuality judge as ground truth would have produced confident nonsense. Your inline comment explaining that is exactly the kind of note that saves the next reader an hour.

I checked out f7cc8a2 and ran it. The 9 tests in tests/test_hr_preset.py pass and the full suite is green at 15. The rewording from "physician-written" to "row-specific" in the RubricJudge docstring and prompt is fine by me; that judge outgrew its HealthBench origin the moment a second preset started using rubrics, and nothing in the suite asserted on the old string.

One thing worth naming for the record: the BBQ loader is global, so the new sampling and rubric behavior applies to any caller of ExampleRegistry.load("bbq"), not only the HR preset. That follows from the change rather than being hidden scope, and I think it is the better default.

Squashing and merging. Thanks for seeing this through, and thanks again for the assistance disclosure. Closes #10.

@knisar
knisar merged commit ab24905 into wandb:main Sep 2, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Add an industry preset beyond healthcare / finance / government (legal or HR)

2 participants