Skip to content

fix(backends): keep sentencepiece spaces in the MLXLM Outlines Core vocabulary - #2072

Open
BlueX888 wants to merge 1 commit into
dottxt-ai:mainfrom
BlueX888:fix/prep-mlxlm-outlines-core-vocabulary-strips-sentencepiece-spaces
Open

BlueX888 wants to merge 1 commit into
dottxt-ai:mainfrom
BlueX888:fix/prep-mlxlm-outlines-core-vocabulary-strips-sentencepiece-spaces

Conversation

@BlueX888

Copy link
Copy Markdown

Problem

The MLXLM branch of OutlinesCoreBackend.__init__ (src/outlines/backends/outlines_core.py:199-204 on main) built the Outlines Core vocabulary from model.mlx_tokenizer with a raw convert_tokens_to_string lambda. For sentencepiece-style tokenizers that call strips the space information:

>>> tokenizer.convert_tokens_to_string(["<0x20>"])
''
>>> tokenizer.convert_tokens_to_string(["▁"])
''
>>> tokenizer.convert_tokens_to_string(["▁b"])
'b'

The regex FSM is built from these strings (the comment in create_outlines_core_vocabulary, src/outlines/backends/outlines_core.py:286-289, notes this conversion step exists "in particular for spacing"), so space-carrying tokens become invisible to the FSM: they are accepted in every state while the model actually emits a space, and the text generated under constraint can violate the regex/JSON schema.

The Transformers branch (src/outlines/backends/outlines_core.py:187-192) already avoids this by using TransformerTokenizer.convert_token_to_string (src/outlines/models/transformers.py:63-73), which re-adds the leading space for tokens starting with SPIECE_UNDERLINE or equal to <0x20>. MLXLM.__init__ already builds exactly that wrapper as self.tokenizer (src/outlines/models/mlxlm.py:135); the MLXLM branch just bypassed it.

Concretely, with a sentencepiece tokenizer (hf-internal-testing/tiny-random-LlamaForCausalLM) and a [0-9]+ logits processor, greedily taking the first allowed token at each step: the MLXLM branch emits the <0x20> token forever (7 spaces, re.fullmatch("[0-9]+", ...) is None) while the Transformers branch of the same backend emits "0".

Changes

  1. src/outlines/backends/outlines_core.py: the MLXLM branch now reads the tokenizer through model.tokenizer (the TransformerTokenizer wrapper) and uses tokenizer.convert_token_to_string, mirroring the Transformers branch. The vocabulary, eos_token_id and eos_token are unchanged (the wrapper reads them from the same underlying tokenizer).
  2. tests/backends/test_outlines_core.py: two regression tests using a sentencepiece-style tokenizer, skipped when mlx is unavailable:
    • test_outlines_core_mlxlm_vocabulary_keeps_sentencepiece_spaces asserts <0x20> and ▁ map to " " and ▁b maps to " b" in the backend vocabulary.
    • test_outlines_core_mlxlm_processor_output_respects_regex asserts the text produced under a [0-9]+ logits processor fullmatches the regex.

The branch was previously # pragma: no cover and the only MLX model covered by the suite (SmolLM, byte-BPE) has no sentencepiece tokens, so the existing tests could not catch this. The pragma is kept because CI runs on Linux, where these tests are skipped.

Testing

Both new tests fail on main:

FAILED tests/backends/test_outlines_core.py::test_outlines_core_mlxlm_vocabulary_keeps_sentencepiece_spaces
>       assert vocab["<0x20>"] in space_token_ids
E       assert 35 in [259]
FAILED tests/backends/test_outlines_core.py::test_outlines_core_mlxlm_processor_output_respects_regex
>       assert re.fullmatch("[0-9]+", text) is not None
E       AssertionError: assert None is not None
E        +  where None = <function fullmatch at 0x...>('[0-9]+', '       ')

and pass with this change:

tests/backends/test_outlines_core.py::test_outlines_core_mlxlm_vocabulary_keeps_sentencepiece_spaces PASSED
tests/backends/test_outlines_core.py::test_outlines_core_mlxlm_processor_output_respects_regex PASSED
====================== 2 passed, 11 deselected in 11.53s =======================

The full module gives 12 passed with this change; the one remaining local failure (test_outlines_core_processor_mlx, AttributeError: GPT2Tokenizer has no attribute vocabulary with the local transformers 5.17.0) also fails on main in this environment and is unrelated to this change.

pre-commit run --all-files passes (check-merge-conflict, debug-statements, end-of-file-fixer, trailing-whitespace, mypy, ruff), and the new tests skip cleanly when mlx is not importable, as on CI.

Contributor Checklist

  • We should be able to understand what the PR does from its title only
  • There is a high-level description of the changes
  • There are links to all the relevant issues (no existing issue covers this bug)
  • The branch is rebased on the latest main commit
  • Commit messages follow these guidelines
  • One commit per logical change
  • The code respects the current naming conventions
  • Docstrings follow the numpy style guide
  • pre-commit is installed and configured on your machine, and you ran it before opening the PR
  • There are tests covering the changes
  • The documentation is up-to-date

The MLXLM branch of OutlinesCoreBackend.__init__ built the Outlines Core
vocabulary with the raw convert_tokens_to_string, which strips the
leading space of sentencepiece-style tokens: "<0x20>" and "▁" mapped to
the empty string (allowed in every FSM state while the model actually
emits a space) and "▁b" mapped to "b" instead of " b". The FSM was
therefore built over wrong token strings and the text emitted under
constraint could violate the regex.

Use the TransformerTokenizer wrapper that MLXLM already provides
(model.tokenizer), whose convert_token_to_string preserves the space,
as the Transformers branch of the same backend already does.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant