You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
This commit was created on GitHub.com and signed with GitHub’s verified signature.
[v1.3.4] 2026-10-01
Bug Fix
Resolve all quantizer-specific tensor state in the quantized-layer replacement path before generic state-dict key remapping. This prevents Gemma4 multimodal parameters such as audio_tower.rel_pos_enc.inv_timescales from capturing a GPTQ scales tensor, and supports GPTQ/DBF/MDBF/OneBit checkpoints with VLM/text wrapper prefix differences.
Fix scale layout handling for groupsize=-1 in RTN fallback. This fallback is used when an MoE expert receives no routed calibration tokens, because GPTQ cannot compute activation-based statistics for that expert. The fix keeps the fallback result compatible with GPTQ's per-channel dequantization path.
Fix MPS loading of large sharded checkpoints by loading weights on CPU before moving the model to MPS.
Reject chunked calibration (CalibrationConfig(batch_size=...)) on MPS, where it is not supported.
Fix MoE fusion for Gemma 4 MoE and Qwen MoE models by preserving the up_proj and down_proj weight dtypes when allocating fused tensors.(unfuse_moe.py)
Support UserDict-based shared KV states when QEP processes transformer blocks one at a time.
Documentation
Clarified that OneComp is released under the MIT License and that licenses for dependency OSS may change when dependencies are updated.
Clarify the confirmed Qwen3.6 save_format="full_wrapper" workflows, including vLLM serving and the current GGUF export workflow.
Tests (CI)
Make cluster CI use an explicitly prepared, content-verified C4 calibration
cache from shared storage. This removes its dependency on Hugging Face Hub
connectivity and generated config hashes, which previously caused C4 loading
to fail on offline compute nodes despite unrelated cached configs being
present. Add an intentional cache preparation/verification script and fail
fast when ONECOMP_CALIB_CACHE/c4 is missing or invalid.
Fix parallel cluster test jobs racing on the shared repository's .git/config.lock. The cluster test orchestrator now fetches directly from
the authenticated CI URL under the existing lock without temporarily
rewriting the origin remote.