Skip to content

v1.3.4

Latest

Choose a tag to compare

@FKKimura FKKimura released this 01 Oct 07:58
83e58d8

[v1.3.4] 2026-10-01

Bug Fix

  • Resolve all quantizer-specific tensor state in the quantized-layer replacement path before generic state-dict key remapping. This prevents Gemma4 multimodal parameters such as audio_tower.rel_pos_enc.inv_timescales from capturing a GPTQ scales tensor, and supports GPTQ/DBF/MDBF/OneBit checkpoints with VLM/text wrapper prefix differences.
  • Fix scale layout handling for groupsize=-1 in RTN fallback. This fallback is used when an MoE expert receives no routed calibration tokens, because GPTQ cannot compute activation-based statistics for that expert. The fix keeps the fallback result compatible with GPTQ's per-channel dequantization path.
  • Fix MPS loading of large sharded checkpoints by loading weights on CPU before moving the model to MPS.
  • Reject chunked calibration (CalibrationConfig(batch_size=...)) on MPS, where it is not supported.
  • Fix MoE fusion for Gemma 4 MoE and Qwen MoE models by preserving the up_proj and down_proj weight dtypes when allocating fused tensors.(unfuse_moe.py)
  • Support UserDict-based shared KV states when QEP processes transformer blocks one at a time.

Documentation

  • Clarified that OneComp is released under the MIT License and that licenses for dependency OSS may change when dependencies are updated.
  • Clarify the confirmed Qwen3.6 save_format="full_wrapper" workflows, including vLLM serving and the current GGUF export workflow.

Tests (CI)

  • Make cluster CI use an explicitly prepared, content-verified C4 calibration
    cache from shared storage. This removes its dependency on Hugging Face Hub
    connectivity and generated config hashes, which previously caused C4 loading
    to fail on offline compute nodes despite unrelated cached configs being
    present. Add an intentional cache preparation/verification script and fail
    fast when ONECOMP_CALIB_CACHE/c4 is missing or invalid.
  • Fix parallel cluster test jobs racing on the shared repository's
    .git/config.lock. The cluster test orchestrator now fetches directly from
    the authenticated CI URL under the existing lock without temporarily
    rewriting the origin remote.