Skip to content

FIX: match XL-SafetyBench cultural category labels to upstream - #3006

Open
Edson (Zhuoxi2000) wants to merge 2 commits into
microsoft:mainfrom
Zhuoxi2000:fix-xlsb-cultural-category-commas
Open

Edson (Zhuoxi2000) wants to merge 2 commits into
microsoft:mainfrom
Zhuoxi2000:fix-xlsb-cultural-category-commas

Conversation

@Zhuoxi2000

Copy link
Copy Markdown

Description

Fixes #2987

Three XLSafetyBenchCulturalCategory values were missing the commas that the upstream dataset uses in the category column of every data/cultural/<country>/scenario_prompts.csv (AIM-Intelligence/XL-SafetyBench on Hugging Face). The cultural loader filters rows by exact string match, so filtering by any of these members matched no rows and raised ValueError: No XL-SafetyBench cultural scenarios matched the configured filters. This PR changes the three values to the upstream labels:

  • FOOD_DIETARY_LAW_AND_HOSPITALITY: Food Dietary Law & Hospitality -> Food, Dietary Law & Hospitality
  • DEATH_GRIEF_AND_FUNERAL_PRACTICES: Death Grief & Funeral Practices -> Death, Grief & Funeral Practices
  • HIERARCHY_ADDRESS_AND_SOCIAL_DEFERENCE: Hierarchy Address & Social Deference -> Hierarchy, Address & Social Deference

This also corrects the class-level harm_categories list on _XLSafetyBenchCulturalDataset, which is built from these values. Member names are unchanged. The category filter only accepts enum members, so callers that pass members need no change. Only code that compares against the old .value strings would notice, and those strings never matched any upstream row.

Tests and Documentation

New parametrized test test_cultural_category_filter_matches_upstream_labels_with_commas in tests/unit/datasets/test_xl_safety_bench_dataset.py. It feeds a row with each upstream comma label through the category filter.

Before the fix (the new test run against main at 4659ebe):

$ python -m pytest tests/unit/datasets/test_xl_safety_bench_dataset.py -q
FAILED ...::test_cultural_category_filter_matches_upstream_labels_with_commas[Food Dietary Law & Hospitality-Food, Dietary Law & Hospitality]
FAILED ...::test_cultural_category_filter_matches_upstream_labels_with_commas[Death Grief & Funeral Practices-Death, Grief & Funeral Practices]
FAILED ...::test_cultural_category_filter_matches_upstream_labels_with_commas[Hierarchy Address & Social Deference-Hierarchy, Address & Social Deference]
3 failed, 34 passed

After the fix:

$ python -m pytest tests/unit/datasets/test_xl_safety_bench_dataset.py -q
37 passed
$ python -m pytest tests/unit/datasets/ -q
4960 passed

Against the real dataset (France), every XLSafetyBenchCulturalCategory member now returns scenarios: 15 each, and 25 for Legal Landmines. Before the fix, the three affected members raised ValueError.

pre-commit run --files on both changed files passes, including ruff format, ruff check (v0.16.9, as pinned in .pre-commit-config.yaml) and ty. No documentation change is needed, and no notebooks or docs are touched, so JupyText was not run.

Three XLSafetyBenchCulturalCategory values were missing the commas used
in the category column of every upstream
data/cultural/<country>/scenario_prompts.csv ('Food, Dietary Law &
Hospitality', 'Death, Grief & Funeral Practices', 'Hierarchy, Address &
Social Deference'). The cultural loader filters rows by exact string
match, so filtering by any of these members matched no rows and raised
ValueError. The class-level harm_categories metadata carried the same
wrong strings.
@hannahwestra25 hannahwestra25 self-assigned this Oct 8, 2026

@hannahwestra25 hannahwestra25 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

thanks for contributing

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

XLSafetyBenchCulturalCategory: three values are missing the upstream commas, so filtering by them never matches

2 participants