Repository navigation
fix: Resolve flaky text_only_forward_produces_finite_logits in granite4_vision and hunyuan_vl parity tests #997
Description
Activity
- addedtype:bugBug fixes, error corrections, or issue resolutionsBug fixes, error corrections, or issue resolutionspriority:mediumMedium priorityMedium prioritystatus:readyReady to be worked onReady to be worked on
on Aug 2, 2026 Reproduced on current
main(e909998f), with a sharper characterization that rules out one of the hypotheses in the report.The granite4 test fails in isolation too, so this is not contention between test targets.
The issue currently records that both tests "pass every time when their target is run in isolation". That no longer holds for granite4. Running the single test alone, in its own process, with nothing else on the machine:
cargo test --release --features metal,accelerate --test granite4_vision_parity text_only_forward20 runs: 16 pass, 4 fail (20%). Failure pattern across the run sequence was
..F.F...F....F......, so it is not a warm-up or first-run effect either.That leaves no other test binary, no other test in the same binary, and no concurrent GPU work to blame. Whatever produces the non-finite logit is inside a single load-plus-one-forward of
granite-4.0-3b-vision-4bit.Whole-suite run on the same commit,
cargo test --release --features metal,accelerate --no-fail-fast, ended with exactly:error: 2 targets failed: `--test granite4_vision_parity` `--test hunyuan_vl_parity`Nothing else in the workspace failed. Note that
fused_moe_geglu_kernel_matches_references_gemma4_shape(#964) passed in this run, which is its own separate flakiness and not relevant here.hunyuan_vl still matches the original description. It failed in the whole-suite run and passed when its target was run alone in the same session. So the two tests may not share a cause, and the report's framing of them as one phenomenon is worth holding loosely. Granite 4 is a hybrid (
granitemoehybrid, Mamba plus Transformer); the Hunyuan VL text backbone shows no SSM path. A recurring-state initialization bug would explain a nondeterministic NaN in the first forward of the former and would not explain the latter at all.That last paragraph is a hypothesis, not a finding. What is measured here is the 20% isolated-failure rate and the target list above.
Context: found while running the full suite as the integration check for epic #909. None of that epic's changes touch these paths, and the failure predates it (this issue was filed from a different branch), so this is not a regression from it.
New data point from the #1074 verification runs
While verifying PR #1078 on top of
main, two consecutivemake verify-testruns (cargo test --workspace --profile test-fast --features metal,accelerate --no-fail-fast) each produced exactly one failure of this class, but a different test each time:- run 1:
granite4_vision_parity::text_only_forward_produces_finite_logits - run 2 (same source, 7595 passed / 1 failed):
hunyuan_vl_parity::text_only_forward_produces_finite_logits
Both passed when re-run in isolation, including running their whole test binary. Neither failure is attributable to #1078, whose change is confined to the
dequantizeFFI shim.This suggests the trigger is whole-workspace concurrency rather than anything specific to one family: whichever real-checkpoint parity binary happens to contend loses. That is a different shape from the observation already recorded here that granite4 fails 4 times in 20 even as a single test in a single process, so the two may be separate mechanisms sharing one symptom, or the same resource pressure sampled at different concurrency levels. Recording it rather than concluding.
Machine context: this box carries persistent background load, so a load-sensitive assertion is expected to be unstable here in a way it may not be on a quiet machine.
- run 1:
- addedstatus:in-progressCurrently being worked onCurrently being worked onstatus:readyReady to be worked onReady to be worked onand removedstatus:readyReady to be worked onReady to be worked onstatus:in-progressCurrently being worked onCurrently being worked on
on Aug 10, 2026 New evidence that narrows this: the hunyuan_vl case now reproduces in isolation, which contradicts the "passes every time when their target is run in isolation" premise in the issue body.
Measured 2026-08-16 on Apple Silicon, release profile,
DEVELOPER_DIRset to Xcode 26.6.0, while verifying an unrelated PR:cargo test --release --test hunyuan_vl_parity --features metal,accelerate \ text_only_forward_produces_finite_logitsFive consecutive standalone runs of that single test, nothing else in the process: 4 passed, 1 failed. The failure is the same
tests/hunyuan_vl_parity.rs:105assertion,text-only logits must be finite.That is roughly a 20 percent standalone failure rate, which matches the granite4 rate the issue already records (4 in 20) rather than the "isolation always passes" behavior attributed to hunyuan_vl. So the two tests now look like the same phenomenon at the same rate, and the cross-target contention hypothesis is ruled out for hunyuan_vl too, exactly as it already was for granite4.
Worth updating the issue body, since "passes in isolation" is currently doing load-bearing work in the diagnosis and it no longer holds.
Context for why this was measured: it surfaced as the only non-ffmpeg failure in a full workspace run for PR #1174. Attribution was checked before reporting: hunyuan loads through
src/loading/vlm_hunyuan_vl.rs, which that PR does not touch, and the standalone repeats above were run on the PR branch, so the flake is present there independently of the whole-suite context. Not a regression from that PR.The box carried background load during these runs, which is worth recording but does not explain a non-finite logit.
A thread-count-controlled measurement at
5dfcb390, and a ruling-outFiled #1211 for the granite4 half before finding this issue; closing that as a duplicate and moving what was new here.
Measured 2026-08-18 on an Apple M5 Max (
Mac17,7, 18 logical CPUs, macOS 26.6.1) at5dfcb390onmain, running the compiledgranite4_vision_paritybinary directly under[profile.test-fast]with--features metal,accelerate, 40 runs per arm, machine otherwise idle:libtest threads runs failures rate default (18) 40 6 15% --test-threads=140 4 10% The comparison is the point, not the rates. Serializing libtest does not move it: both arms fail at the same order of magnitude, 10 of 80 combined. Together with the 2026-08-03 isolated 4-in-20 already in this thread, thread count and cross-target contention are both ruled out for granite4 at two commits three weeks apart.
This is not #1092 and PR #1210 does not fix it. #1210 serializes the macOS gate because the
mlxcel-corebinary tookSIGSEGVunder 18 concurrent libtest workers. The table above was collected specifically to keep the two apart: under exactly the configuration #1210 lands, this assertion still fails 4 of 40. Do not close this when #1210 merges.Three things worth knowing before someone starts bisecting, from reading the path rather than from measurement:
- No vision code is involved.
src/vision/granite4_vision.rs:263-271implementsLanguageModel::forwardas a straight delegation toself.text_model.forward. DeepStack and window-QFormer injection are only reached throughforward_with_embeddings(:273-281), which this test never calls. This is agranitemoehybridtext-backbone question. - The external caches are not the state that varies.
src/models/granitemoehybrid.rs:1514-1526ignores theKVCacheslice the test passes and routes throughsequence_state.with_sequence_state(None, ...). The per-layer mixed cache (:66-68, built at:1049) holdsMamba2Cacherecurrent conv/SSM state for the mamba layers. With fixed weights and a fixed input, that recurrent state is the one thing in this path a freshKVCacheis not, which makes it the natural first suspect. - The failing value has never been captured. The assertion only asks
is_finite(), so nobody knows whether it isNaNorInf, or how many entries of the row are affected.NaNpoints at0 * Inf,Inf - Inf, or a badexp/rsqrtinput; a rawInfpoints at overflow or a zero denominator. Different suspects. Printing the value and the non-finite count in the assertion message costs nothing and answers it on the next reproduction, and it is worth doing before anything else.
Not touching the
hunyuan_vl_parity.rs:105half; whether the two share a root cause is still open.- No vision code is involved.
Fresh occurrence data from a 13-unit sequential implementation run on 2026-08-23, offered because it adds a frequency measurement and a hardware/profile data point the issue does not yet have.
Environment. Apple M1 Ultra, macOS,
cargo test --workspace --profile test-fast --features metal,accelerate. Note this is thetest-fastprofile, not thereleaseprofile the original evidence table used, and--workspacerather than the root package alone. The failure reproduces on both.Frequency. Over 9 full-workspace gate runs across 5 different feature branches,
hunyuan_vl_parity::text_only_forward_produces_finite_logitsfailed on 3 of them. The granite4_vision sibling did not fail in this batch. Each failure aborted the whole run at that binary, so the surrounding suite reported roughly 6200 of the expected ~8360 tests, which is what makes the failure expensive rather than merely noisy: a re-run costs a full gate cycle.Isolation, per failure. Each time, the target passed 3 of 3 in isolation immediately afterwards:
cargo test --profile test-fast --features metal,accelerate --test hunyuan_vl_parity run 1: exit=0 test result: ok. 4 passed; 0 failed run 2: exit=0 test result: ok. 4 passed; 0 failed run 3: exit=0 test result: ok. 4 passed; 0 failedIndependence from the branches under test. The three failing runs were on branches for #1355 (llama3 rope_scaling), #1324 (internlm dynamic NTK) and #1346 (prompt cache gating). None of the three touches any file a hunyuan path reaches, verified per branch with
git diff --name-only origin/main...HEAD | grep -iE 'hunyuan|vision'returning nothing. Two of the three branches also had an earlier green full-workspace run on the same commit range, so the same code both passed and failed the same gate.A detail that may narrow the cause. The failures clustered in runs where the gate followed other GPU work in the same session (a release build, or a prior full gate). The runs that came after an idle period passed. That is consistent with the memory-pressure hypothesis rather than a specific test-ordering interaction, though this batch did not vary the two independently and so does not establish it.
Practical impact for anyone running the documented gate locally. With the flake at roughly 1 in 3 full-workspace runs on this machine, a merge gate that treats a single red run as decisive is unreliable in both directions: it blocks on noise, and it trains the reader to discount a red result. The working procedure used in this run, offered in case it is useful before a real fix lands, was to treat a failure as inconclusive rather than as either pass or fail, and to resolve it with four checks: confirm the branch touches no file the failing test's path reaches, run the target in isolation, confirm the same commit range passed the gate earlier, and then re-run the full gate. Every occurrence here resolved as a flake under that procedure, but the procedure is what made "flake" a conclusion rather than an assumption.
Four findings that narrow this, and one that contradicts a premise in the evidence table above. All on Apple M5 Max 128GB,
--features metal,accelerate,test-fastprofile, at1b6760d7,granite-4.0-3b-vision-4bitonly.The value is NaN, never an infinity. The issue text says "non-finite (NaN or inf)" because it had not been captured; a probe on the assertion reports
nan=true, inf=falseon every failure observed. That points away from overflow and toward a value produced by an invalid operation or read from memory that was not written.It reproduces in isolation, so the whole-suite framing is wrong.
--test granite4_vision_parity text_only_forwardalone failed 1 run in 5, then 1 in 8, on fresh processes with no other test in the binary. The table above records three consecutive isolated passes, which now reads as too small a sample rather than as a property of isolation. Ordering and cross-test contention are not required.There is a much cheaper reproducer. Twenty forwards in one process, same tokens, hit NaN on 1, 4 and 3 iterations in three runs. That is roughly a hundred times faster to iterate on than one failure per six processes, and it holds whether
make_cachesis called once before the loop or on every iteration, so replacing the internal sequence state is not the trigger.The two sibling tests pass.
smolvlm_parityandmolmo_paritycarry a test of the same name and shape and passed 6 of 6 each. Shared VLM test scaffolding would have moved all three.The real generation path looks unaffected.
mlxcel generate -m granite-4.0-3b-vision-4bit -p "Hello, world." -n 16 --temp 0produced identical output on 6 of 6 runs. So the fault is reached by callingLanguageModel::forwarddirectly and not by the prefill and decode drivers, which is worth settling before this is treated as user-visible.One structural note for whoever picks this up:
GraniteMoeHybrid'smake_cachesdiscards theVec<KVCache>it returns to the caller and replaces model-internal state instead, so thecachesthe test threads through are not the state the forward reads. The internal per-layer state isMamba2Cache::new()for mamba mixers andKVCache::new()for attention mixers, and both startNone, so nothing is uninitialized at construction. The nondeterminism is downstream of that.- added a commit that references this issue
on Sep 9, 2026
Problem / Background
Two real-model VLM parity tests, both named
text_only_forward_produces_finite_logits, fail intermittently when the whole root-package test suite runs, and pass every time when their target is run in isolation.tests/granite4_vision_parity.rs:100(modelgranite-4.0-3b-vision-4bit), assertiontext-only logits must be finite.tests/hunyuan_vl_parity.rs:105(modelhunyuanocr-mlx-4bit), same assertion.Both tests follow the same shape:
mlxcel::load_model(&dir), tokenize"Hello, world.",LanguageModel::make_caches,LanguageModel::forward,eval, slice the last position,max_all, then assertitem_f32(&max).is_finite(). The observed failure is therefore a non-finite (NaN or inf) maximum logit on a text-only forward pass.These failures are unrelated to the fix in #996 (issue #962). They surfaced while running
cargo test --release --features metal,accelerate --no-fail-fastas that issue's fourth acceptance criterion, and are deliberately not folded into that PR.Evidence
All observations on the same machine (Apple Silicon, macOS,
--features metal,accelerate, release profile), against branchfix/issue-962-nightly-verify-redat commitd32f684ca.granite4_vision_parityhunyuan_vl_paritycargo test --release --features metal,accelerate --no-fail-fast(73 result lines)--test granite4_vision_parity--test hunyuan_vl_parityTwo things that are worth ruling out up front:
tests/cli_help_consistency.rsandtests/common/mod.rs. Neither failing test spawns a binary; both reachcommon::repo_model_dir, which that change did not touch. The change is not a plausible cause.glm_ocr_parity,granite_vision_parity,internvl_parity) were green in both runs.Not urgent
Both tests self-skip when
models/<name>is absent:model_dir()returnsNoneand the test returns early. The nightly-verify self-hosted runner carries no model weights, as.github/workflows/nightly-verify.ymldocuments, so neither test executes there and this cannot make the nightly red. It only reproduces on a developer machine that has the checkpoints locally.Hypothesis (unverified)
Tests within one integration target run in parallel threads, and each of these two targets calls
mlxcel::load_modelfrom more than one test against the same real checkpoint. The repository has prior art for parallel real-model loads producing NaN under memory pressure, which is why several parity suites document--test-threads=1as a requirement (for exampletests/speculative_parity.rs: "real-model tests share GPU memory and concurrent loads will OOM on smaller (32-48 GB) Apple Silicon hosts", andtests/ernie4_5_moe_vl_parity.rs).That makes concurrent load plus Metal memory pressure the first thing to check, but it is a hypothesis, not a conclusion. Whoever picks this up should confirm:
--test-threads=1for these two targets,Acceptance Criteria
cargo test --release --features metal,accelerate --no-fail-fastis green across repeated runs on a machine that has both checkpoints.