·
23 commits
to master
since this release
Automated Unsloth llama.cpp CUDA + ROCm + Vulkan + macOS + CPU prebuild for upstream b10639, merged with:
- Carry ggml-org#24423 (DiffusionGemma) merged onto b10630 (unslothai/llama.cpp#107, commit 74acc40)
- Carry ggml-org#25731 (TML Inkling) merged onto b10630 (unslothai/llama.cpp#108, commit 2c5b000)
- kimi-k3 : the MoonViT-3d vision tower and full-size loading fixes (unslothai/llama.cpp#70, commit edfd4c1)
- IQ1_XS, IQ1_XXS, IQ1_XXXS on an upstream base so the nightly can pin them (unslothai/llama.cpp#91, commit c86ed26)
- sampling: index penalties by token id instead of scanning every candidate (unslothai/llama.cpp#95, commit 3db8cb5)
- Carry ggml-org#27742 (Qwen3.8-Flash-Next) merged onto b10632 (unslothai/llama.cpp#114, commit 950f135)
- Carry ggml-org#27754 (GLM-5-Next) composed onto the qwen4exp b10639 carry (unslothai/llama.cpp#125, commit f48b99e)