Repository navigation
fix(backends): keep sentencepiece spaces in the MLXLM Outlines Core vocabulary - #2072
Open
BlueX888 wants to merge 1 commit into
Conversation
The MLXLM branch of OutlinesCoreBackend.__init__ built the Outlines Core vocabulary with the raw convert_tokens_to_string, which strips the leading space of sentencepiece-style tokens: "<0x20>" and "▁" mapped to the empty string (allowed in every FSM state while the model actually emits a space) and "▁b" mapped to "b" instead of " b". The FSM was therefore built over wrong token strings and the text emitted under constraint could violate the regex. Use the TransformerTokenizer wrapper that MLXLM already provides (model.tokenizer), whose convert_token_to_string preserves the space, as the Transformers branch of the same backend already does.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
The MLXLM branch of
OutlinesCoreBackend.__init__(src/outlines/backends/outlines_core.py:199-204 onmain) built the Outlines Core vocabulary frommodel.mlx_tokenizerwith a rawconvert_tokens_to_stringlambda. For sentencepiece-style tokenizers that call strips the space information:The regex FSM is built from these strings (the comment in
create_outlines_core_vocabulary, src/outlines/backends/outlines_core.py:286-289, notes this conversion step exists "in particular for spacing"), so space-carrying tokens become invisible to the FSM: they are accepted in every state while the model actually emits a space, and the text generated under constraint can violate the regex/JSON schema.The Transformers branch (src/outlines/backends/outlines_core.py:187-192) already avoids this by using
TransformerTokenizer.convert_token_to_string(src/outlines/models/transformers.py:63-73), which re-adds the leading space for tokens starting withSPIECE_UNDERLINEor equal to<0x20>.MLXLM.__init__already builds exactly that wrapper asself.tokenizer(src/outlines/models/mlxlm.py:135); the MLXLM branch just bypassed it.Concretely, with a sentencepiece tokenizer (
hf-internal-testing/tiny-random-LlamaForCausalLM) and a[0-9]+logits processor, greedily taking the first allowed token at each step: the MLXLM branch emits the<0x20>token forever (7 spaces,re.fullmatch("[0-9]+", ...) is None) while the Transformers branch of the same backend emits"0".Changes
src/outlines/backends/outlines_core.py: the MLXLM branch now reads the tokenizer throughmodel.tokenizer(theTransformerTokenizerwrapper) and usestokenizer.convert_token_to_string, mirroring the Transformers branch. The vocabulary,eos_token_idandeos_tokenare unchanged (the wrapper reads them from the same underlying tokenizer).tests/backends/test_outlines_core.py: two regression tests using a sentencepiece-style tokenizer, skipped whenmlxis unavailable:test_outlines_core_mlxlm_vocabulary_keeps_sentencepiece_spacesasserts<0x20>and▁map to" "and▁bmaps to" b"in the backend vocabulary.test_outlines_core_mlxlm_processor_output_respects_regexasserts the text produced under a[0-9]+logits processor fullmatches the regex.The branch was previously
# pragma: no coverand the only MLX model covered by the suite (SmolLM, byte-BPE) has no sentencepiece tokens, so the existing tests could not catch this. The pragma is kept because CI runs on Linux, where these tests are skipped.Testing
Both new tests fail on
main:and pass with this change:
The full module gives
12 passedwith this change; the one remaining local failure (test_outlines_core_processor_mlx,AttributeError: GPT2Tokenizer has no attribute vocabularywith the local transformers 5.17.0) also fails onmainin this environment and is unrelated to this change.pre-commit run --all-filespasses (check-merge-conflict, debug-statements, end-of-file-fixer, trailing-whitespace, mypy, ruff), and the new tests skip cleanly whenmlxis not importable, as on CI.Contributor Checklist
maincommitpre-commitis installed and configured on your machine, and you ran it before opening the PR