Skip to content

fix: Resolve flaky text_only_forward_produces_finite_logits in granite4_vision and hunyuan_vl parity tests #997

Description

@inureyes

Problem / Background

Two real-model VLM parity tests, both named text_only_forward_produces_finite_logits, fail intermittently when the whole root-package test suite runs, and pass every time when their target is run in isolation.

  • tests/granite4_vision_parity.rs:100 (model granite-4.0-3b-vision-4bit), assertion text-only logits must be finite.
  • tests/hunyuan_vl_parity.rs:105 (model hunyuanocr-mlx-4bit), same assertion.

Both tests follow the same shape: mlxcel::load_model(&dir), tokenize "Hello, world.", LanguageModel::make_caches, LanguageModel::forward, eval, slice the last position, max_all, then assert item_f32(&max).is_finite(). The observed failure is therefore a non-finite (NaN or inf) maximum logit on a text-only forward pass.

These failures are unrelated to the fix in #996 (issue #962). They surfaced while running cargo test --release --features metal,accelerate --no-fail-fast as that issue's fourth acceptance criterion, and are deliberately not folded into that PR.

Evidence

All observations on the same machine (Apple Silicon, macOS, --features metal,accelerate, release profile), against branch fix/issue-962-nightly-verify-red at commit d32f684ca.

Run Command granite4_vision_parity hunyuan_vl_parity Rest of suite
Whole-suite 1 cargo test --release --features metal,accelerate --no-fail-fast (73 result lines) 3 passed, 0.97s 4 passed, 1.25s green
Whole-suite 2 identical command 2 passed / 1 failed, 0.82s 3 passed / 1 failed, 0.67s green (5216 passed, 274 ignored)
Isolated x3 --test granite4_vision_parity 3/3 passed, three consecutive attempts
Isolated x3 --test hunyuan_vl_parity 4/4 passed, three consecutive attempts

Two things that are worth ruling out up front:

  • The only source difference between whole-suite run 1 and run 2 was an edit to tests/cli_help_consistency.rs and tests/common/mod.rs. Neither failing test spawns a binary; both reach common::repo_model_dir, which that change did not touch. The change is not a plausible cause.
  • Neighbouring targets in the run order (glm_ocr_parity, granite_vision_parity, internvl_parity) were green in both runs.

Not urgent

Both tests self-skip when models/<name> is absent: model_dir() returns None and the test returns early. The nightly-verify self-hosted runner carries no model weights, as .github/workflows/nightly-verify.yml documents, so neither test executes there and this cannot make the nightly red. It only reproduces on a developer machine that has the checkpoints locally.

Hypothesis (unverified)

Tests within one integration target run in parallel threads, and each of these two targets calls mlxcel::load_model from more than one test against the same real checkpoint. The repository has prior art for parallel real-model loads producing NaN under memory pressure, which is why several parity suites document --test-threads=1 as a requirement (for example tests/speculative_parity.rs: "real-model tests share GPU memory and concurrent loads will OOM on smaller (32-48 GB) Apple Silicon hosts", and tests/ernie4_5_moe_vl_parity.rs).

That makes concurrent load plus Metal memory pressure the first thing to check, but it is a hypothesis, not a conclusion. Whoever picks this up should confirm:

  • whether the non-finite value is NaN or inf,
  • whether it reproduces under --test-threads=1 for these two targets,
  • whether it is specific to the 4-bit quantized path.

Acceptance Criteria

  • Root cause identified, distinguishing a genuine numerical bug in the granite4_vision / hunyuan_vl text-only forward path from a test-harness concurrency or memory-pressure artifact.
  • If it is a harness artifact, the two targets are made deterministic without weakening what they assert.
  • If it is a real numerical bug, it is fixed and a test pins it.
  • cargo test --release --features metal,accelerate --no-fail-fast is green across repeated runs on a machine that has both checkpoints.

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:mediumMedium prioritystatus:readyReady to be worked ontype:bugBug fixes, error corrections, or issue resolutions

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions