Skip to content

fix: Resolve flaky text_only_forward_produces_finite_logits in granite4_vision and hunyuan_vl parity tests #997

Description

@inureyes

Problem / Background

Two real-model VLM parity tests, both named text_only_forward_produces_finite_logits, fail intermittently when the whole root-package test suite runs, and pass every time when their target is run in isolation.

  • tests/granite4_vision_parity.rs:100 (model granite-4.0-3b-vision-4bit), assertion text-only logits must be finite.
  • tests/hunyuan_vl_parity.rs:105 (model hunyuanocr-mlx-4bit), same assertion.

Both tests follow the same shape: mlxcel::load_model(&dir), tokenize "Hello, world.", LanguageModel::make_caches, LanguageModel::forward, eval, slice the last position, max_all, then assert item_f32(&max).is_finite(). The observed failure is therefore a non-finite (NaN or inf) maximum logit on a text-only forward pass.

These failures are unrelated to the fix in #996 (issue #962). They surfaced while running cargo test --release --features metal,accelerate --no-fail-fast as that issue's fourth acceptance criterion, and are deliberately not folded into that PR.

Evidence

All observations on the same machine (Apple Silicon, macOS, --features metal,accelerate, release profile), against branch fix/issue-962-nightly-verify-red at commit d32f684ca.

Run Command granite4_vision_parity hunyuan_vl_parity Rest of suite
Whole-suite 1 cargo test --release --features metal,accelerate --no-fail-fast (73 result lines) 3 passed, 0.97s 4 passed, 1.25s green
Whole-suite 2 identical command 2 passed / 1 failed, 0.82s 3 passed / 1 failed, 0.67s green (5216 passed, 274 ignored)
Isolated x3 --test granite4_vision_parity 3/3 passed, three consecutive attempts
Isolated x3 --test hunyuan_vl_parity 4/4 passed, three consecutive attempts

Two things that are worth ruling out up front:

  • The only source difference between whole-suite run 1 and run 2 was an edit to tests/cli_help_consistency.rs and tests/common/mod.rs. Neither failing test spawns a binary; both reach common::repo_model_dir, which that change did not touch. The change is not a plausible cause.
  • Neighbouring targets in the run order (glm_ocr_parity, granite_vision_parity, internvl_parity) were green in both runs.

Not urgent

Both tests self-skip when models/<name> is absent: model_dir() returns None and the test returns early. The nightly-verify self-hosted runner carries no model weights, as .github/workflows/nightly-verify.yml documents, so neither test executes there and this cannot make the nightly red. It only reproduces on a developer machine that has the checkpoints locally.

Hypothesis (unverified)

Tests within one integration target run in parallel threads, and each of these two targets calls mlxcel::load_model from more than one test against the same real checkpoint. The repository has prior art for parallel real-model loads producing NaN under memory pressure, which is why several parity suites document --test-threads=1 as a requirement (for example tests/speculative_parity.rs: "real-model tests share GPU memory and concurrent loads will OOM on smaller (32-48 GB) Apple Silicon hosts", and tests/ernie4_5_moe_vl_parity.rs).

That makes concurrent load plus Metal memory pressure the first thing to check, but it is a hypothesis, not a conclusion. Whoever picks this up should confirm:

  • whether the non-finite value is NaN or inf,
  • whether it reproduces under --test-threads=1 for these two targets,
  • whether it is specific to the 4-bit quantized path.

Acceptance Criteria

  • Root cause identified, distinguishing a genuine numerical bug in the granite4_vision / hunyuan_vl text-only forward path from a test-harness concurrency or memory-pressure artifact.
  • If it is a harness artifact, the two targets are made deterministic without weakening what they assert.
  • If it is a real numerical bug, it is fixed and a test pins it.
  • cargo test --release --features metal,accelerate --no-fail-fast is green across repeated runs on a machine that has both checkpoints.

Activity

  1. inureyes commented on Aug 3, 2026

    @inureyes
    MemberAuthor

    Reproduced on current main (e909998f), with a sharper characterization that rules out one of the hypotheses in the report.

    The granite4 test fails in isolation too, so this is not contention between test targets.

    The issue currently records that both tests "pass every time when their target is run in isolation". That no longer holds for granite4. Running the single test alone, in its own process, with nothing else on the machine:

    cargo test --release --features metal,accelerate --test granite4_vision_parity text_only_forward
    

    20 runs: 16 pass, 4 fail (20%). Failure pattern across the run sequence was ..F.F...F....F......, so it is not a warm-up or first-run effect either.

    That leaves no other test binary, no other test in the same binary, and no concurrent GPU work to blame. Whatever produces the non-finite logit is inside a single load-plus-one-forward of granite-4.0-3b-vision-4bit.

    Whole-suite run on the same commit, cargo test --release --features metal,accelerate --no-fail-fast, ended with exactly:

    error: 2 targets failed:
        `--test granite4_vision_parity`
        `--test hunyuan_vl_parity`
    

    Nothing else in the workspace failed. Note that fused_moe_geglu_kernel_matches_references_gemma4_shape (#964) passed in this run, which is its own separate flakiness and not relevant here.

    hunyuan_vl still matches the original description. It failed in the whole-suite run and passed when its target was run alone in the same session. So the two tests may not share a cause, and the report's framing of them as one phenomenon is worth holding loosely. Granite 4 is a hybrid (granitemoehybrid, Mamba plus Transformer); the Hunyuan VL text backbone shows no SSM path. A recurring-state initialization bug would explain a nondeterministic NaN in the first forward of the former and would not explain the latter at all.

    That last paragraph is a hypothesis, not a finding. What is measured here is the 20% isolated-failure rate and the target list above.

    Context: found while running the full suite as the integration check for epic #909. None of that epic's changes touch these paths, and the failure predates it (this issue was filed from a different branch), so this is not a regression from it.

  2. inureyes commented on Aug 7, 2026

    @inureyes
    MemberAuthor

    New data point from the #1074 verification runs

    While verifying PR #1078 on top of main, two consecutive make verify-test runs (cargo test --workspace --profile test-fast --features metal,accelerate --no-fail-fast) each produced exactly one failure of this class, but a different test each time:

    • run 1: granite4_vision_parity::text_only_forward_produces_finite_logits
    • run 2 (same source, 7595 passed / 1 failed): hunyuan_vl_parity::text_only_forward_produces_finite_logits

    Both passed when re-run in isolation, including running their whole test binary. Neither failure is attributable to #1078, whose change is confined to the dequantize FFI shim.

    This suggests the trigger is whole-workspace concurrency rather than anything specific to one family: whichever real-checkpoint parity binary happens to contend loses. That is a different shape from the observation already recorded here that granite4 fails 4 times in 20 even as a single test in a single process, so the two may be separate mechanisms sharing one symptom, or the same resource pressure sampled at different concurrency levels. Recording it rather than concluding.

    Machine context: this box carries persistent background load, so a load-sensitive assertion is expected to be unstable here in a way it may not be on a quiet machine.

  3. inureyes commented on Aug 15, 2026

    @inureyes
    MemberAuthor

    New evidence that narrows this: the hunyuan_vl case now reproduces in isolation, which contradicts the "passes every time when their target is run in isolation" premise in the issue body.

    Measured 2026-08-16 on Apple Silicon, release profile, DEVELOPER_DIR set to Xcode 26.6.0, while verifying an unrelated PR:

    cargo test --release --test hunyuan_vl_parity --features metal,accelerate \
      text_only_forward_produces_finite_logits
    

    Five consecutive standalone runs of that single test, nothing else in the process: 4 passed, 1 failed. The failure is the same tests/hunyuan_vl_parity.rs:105 assertion, text-only logits must be finite.

    That is roughly a 20 percent standalone failure rate, which matches the granite4 rate the issue already records (4 in 20) rather than the "isolation always passes" behavior attributed to hunyuan_vl. So the two tests now look like the same phenomenon at the same rate, and the cross-target contention hypothesis is ruled out for hunyuan_vl too, exactly as it already was for granite4.

    Worth updating the issue body, since "passes in isolation" is currently doing load-bearing work in the diagnosis and it no longer holds.

    Context for why this was measured: it surfaced as the only non-ffmpeg failure in a full workspace run for PR #1174. Attribution was checked before reporting: hunyuan loads through src/loading/vlm_hunyuan_vl.rs, which that PR does not touch, and the standalone repeats above were run on the PR branch, so the flake is present there independently of the whole-suite context. Not a regression from that PR.

    The box carried background load during these runs, which is worth recording but does not explain a non-finite logit.

  4. inureyes commented on Aug 18, 2026

    @inureyes
    MemberAuthor

    A thread-count-controlled measurement at 5dfcb390, and a ruling-out

    Filed #1211 for the granite4 half before finding this issue; closing that as a duplicate and moving what was new here.

    Measured 2026-08-18 on an Apple M5 Max (Mac17,7, 18 logical CPUs, macOS 26.6.1) at 5dfcb390 on main, running the compiled granite4_vision_parity binary directly under [profile.test-fast] with --features metal,accelerate, 40 runs per arm, machine otherwise idle:

    libtest threads runs failures rate
    default (18) 40 6 15%
    --test-threads=1 40 4 10%

    The comparison is the point, not the rates. Serializing libtest does not move it: both arms fail at the same order of magnitude, 10 of 80 combined. Together with the 2026-08-03 isolated 4-in-20 already in this thread, thread count and cross-target contention are both ruled out for granite4 at two commits three weeks apart.

    This is not #1092 and PR #1210 does not fix it. #1210 serializes the macOS gate because the mlxcel-core binary took SIGSEGV under 18 concurrent libtest workers. The table above was collected specifically to keep the two apart: under exactly the configuration #1210 lands, this assertion still fails 4 of 40. Do not close this when #1210 merges.

    Three things worth knowing before someone starts bisecting, from reading the path rather than from measurement:

    • No vision code is involved. src/vision/granite4_vision.rs:263-271 implements LanguageModel::forward as a straight delegation to self.text_model.forward. DeepStack and window-QFormer injection are only reached through forward_with_embeddings (:273-281), which this test never calls. This is a granitemoehybrid text-backbone question.
    • The external caches are not the state that varies. src/models/granitemoehybrid.rs:1514-1526 ignores the KVCache slice the test passes and routes through sequence_state.with_sequence_state(None, ...). The per-layer mixed cache (:66-68, built at :1049) holds Mamba2Cache recurrent conv/SSM state for the mamba layers. With fixed weights and a fixed input, that recurrent state is the one thing in this path a fresh KVCache is not, which makes it the natural first suspect.
    • The failing value has never been captured. The assertion only asks is_finite(), so nobody knows whether it is NaN or Inf, or how many entries of the row are affected. NaN points at 0 * Inf, Inf - Inf, or a bad exp/rsqrt input; a raw Inf points at overflow or a zero denominator. Different suspects. Printing the value and the non-finite count in the assertion message costs nothing and answers it on the next reproduction, and it is worth doing before anything else.

    Not touching the hunyuan_vl_parity.rs:105 half; whether the two share a root cause is still open.

  5. inureyes commented on Aug 23, 2026

    @inureyes
    MemberAuthor

    Fresh occurrence data from a 13-unit sequential implementation run on 2026-08-23, offered because it adds a frequency measurement and a hardware/profile data point the issue does not yet have.

    Environment. Apple M1 Ultra, macOS, cargo test --workspace --profile test-fast --features metal,accelerate. Note this is the test-fast profile, not the release profile the original evidence table used, and --workspace rather than the root package alone. The failure reproduces on both.

    Frequency. Over 9 full-workspace gate runs across 5 different feature branches, hunyuan_vl_parity::text_only_forward_produces_finite_logits failed on 3 of them. The granite4_vision sibling did not fail in this batch. Each failure aborted the whole run at that binary, so the surrounding suite reported roughly 6200 of the expected ~8360 tests, which is what makes the failure expensive rather than merely noisy: a re-run costs a full gate cycle.

    Isolation, per failure. Each time, the target passed 3 of 3 in isolation immediately afterwards:

    cargo test --profile test-fast --features metal,accelerate --test hunyuan_vl_parity
    run 1: exit=0 test result: ok. 4 passed; 0 failed
    run 2: exit=0 test result: ok. 4 passed; 0 failed
    run 3: exit=0 test result: ok. 4 passed; 0 failed
    

    Independence from the branches under test. The three failing runs were on branches for #1355 (llama3 rope_scaling), #1324 (internlm dynamic NTK) and #1346 (prompt cache gating). None of the three touches any file a hunyuan path reaches, verified per branch with git diff --name-only origin/main...HEAD | grep -iE 'hunyuan|vision' returning nothing. Two of the three branches also had an earlier green full-workspace run on the same commit range, so the same code both passed and failed the same gate.

    A detail that may narrow the cause. The failures clustered in runs where the gate followed other GPU work in the same session (a release build, or a prior full gate). The runs that came after an idle period passed. That is consistent with the memory-pressure hypothesis rather than a specific test-ordering interaction, though this batch did not vary the two independently and so does not establish it.

    Practical impact for anyone running the documented gate locally. With the flake at roughly 1 in 3 full-workspace runs on this machine, a merge gate that treats a single red run as decisive is unreliable in both directions: it blocks on noise, and it trains the reader to discount a red result. The working procedure used in this run, offered in case it is useful before a real fix lands, was to treat a failure as inconclusive rather than as either pass or fail, and to resolve it with four checks: confirm the branch touches no file the failing test's path reaches, run the target in isolation, confirm the same commit range passed the gate earlier, and then re-run the full gate. Every occurrence here resolved as a flake under that procedure, but the procedure is what made "flake" a conclusion rather than an assumption.

  6. inureyes commented on Sep 8, 2026

    @inureyes
    MemberAuthor

    Four findings that narrow this, and one that contradicts a premise in the evidence table above. All on Apple M5 Max 128GB, --features metal,accelerate, test-fast profile, at 1b6760d7, granite-4.0-3b-vision-4bit only.

    The value is NaN, never an infinity. The issue text says "non-finite (NaN or inf)" because it had not been captured; a probe on the assertion reports nan=true, inf=false on every failure observed. That points away from overflow and toward a value produced by an invalid operation or read from memory that was not written.

    It reproduces in isolation, so the whole-suite framing is wrong. --test granite4_vision_parity text_only_forward alone failed 1 run in 5, then 1 in 8, on fresh processes with no other test in the binary. The table above records three consecutive isolated passes, which now reads as too small a sample rather than as a property of isolation. Ordering and cross-test contention are not required.

    There is a much cheaper reproducer. Twenty forwards in one process, same tokens, hit NaN on 1, 4 and 3 iterations in three runs. That is roughly a hundred times faster to iterate on than one failure per six processes, and it holds whether make_caches is called once before the loop or on every iteration, so replacing the internal sequence state is not the trigger.

    The two sibling tests pass. smolvlm_parity and molmo_parity carry a test of the same name and shape and passed 6 of 6 each. Shared VLM test scaffolding would have moved all three.

    The real generation path looks unaffected. mlxcel generate -m granite-4.0-3b-vision-4bit -p "Hello, world." -n 16 --temp 0 produced identical output on 6 of 6 runs. So the fault is reached by calling LanguageModel::forward directly and not by the prefill and decode drivers, which is worth settling before this is treated as user-visible.

    One structural note for whoever picks this up: GraniteMoeHybrid's make_caches discards the Vec<KVCache> it returns to the caller and replaces model-internal state instead, so the caches the test threads through are not the state the forward reads. The internal per-layer state is Mamba2Cache::new() for mamba mixers and KVCache::new() for attention mixers, and both start None, so nothing is uninitialized at construction. The nondeterminism is downstream of that.

  7. added a commit that references this issue on Sep 9, 2026
    ec9db8b
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:mediumMedium prioritystatus:readyReady to be worked ontype:bugBug fixes, error corrections, or issue resolutions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions