Skip to content

v0.22.0

Latest

Choose a tag to compare

@github-actions github-actions released this 25 Aug 08:44
34dc0e5

Overview

This release adds tensor-split support to the multi-backend (meta) backend with improved split-state propagation, reworks the Metal kernels into per-op sources with parallel compilation, and fixes ggml_clamp to be a proper non-in-place op. It also brings new ops (POOL_1D, PAD_REFLECT_1D), Q2_K SYCL kernels, MoE bias fusion on OpenCL, and assorted fixes across the CUDA, Metal, SYCL, Vulkan, OpenCL and WebGPU backends.

API changes

  • ggml_clamp is now a non-in-place op returning a new tensor; the previous in-place (view) behavior is available via the new ggml_clamp_inplace (llama.cpp#27644)

Core changes

  • Fix ggml_clamp to be a proper non-in-place op, add ggml_clamp_inplace, and allow in-place clamp in the allocator (llama.cpp#27644)
  • Optimize CPU concat by replacing per-element copies with row-level memcpy (llama.cpp#24575)
  • Meta backend: add tensor-split mode, including delayed all-reduce for independent shared-expert branches (llama.cpp#26490)
  • Meta backend: fix tensor split-state propagation for mul_mat, reshape, flash-attn and other ops (llama.cpp#27574)
  • Shorten virtual device naming in the CUDA and Metal backends (llama.cpp#27608)

Backend changes

CUDA

Metal

SYCL

Vulkan

OpenCL

  • Fuse MoE per-expert bias adds into the GLU and combine epilogues (opt-out via env vars) (llama.cpp#26431)

WebGPU

  • Fix handling of infinity values in ARGSORT and TOP_K (llama.cpp#27538)
  • Reorder includes to avoid a V/K macro conflict in the KV-overlap case (llama.cpp#27545)

More info

Changelog since v0.21.0

34dc0e5 ggml : bump version to 0.22.0 (#1607)
5ef9ad9 sync : llama.cpp
e19a660 sycl : mark tq2_0 as not supported (llama/27660)
29f0560 webgpu : fix handling of infinity values during ARGSORT and TOP_K (llama/27538)
d737987 metal : per-device tuned (Q, NE) for flash-attn vec (llama/26570)
674b15c metal: per-op source split + parallel compile (llama/26561)
12c8fb0 ggml : shorten virtual device naming in CUDA and Metal (llama/27608)
03555f6 webgpu : reorder includes since V that appears in common_decls.tmpl may be defined as K in flash_attn_decls.tmpl if KV_OVERLAP (llama/27545)
709321d ggml : fix ggml_clamp (llama/27644)
3526cdc Deepseek 4: -sm tensor (llama/26490)
70ff0f7 Fix meta tensor split state propagation (llama/27574)
0aab870 cuda : add POOL_1D support (llama/27573)
6af3397 vulkan : added the PAD_REFLECT_1D operation (llama/26586)
47c985d ggml: optimize concat op by replacing per-element memcpy with row-level memcpy (llama/24575)
04f16ef sycl : add Q2_K reordered MMVQ and ESIMD kernels (again) (llama/27490)
4a57f1a opencl: fold the gpt-oss MoE per-expert bias adds into the epilogue (op/kernel fusion) (llama/26431)
d99724f readme : update following llama.cpp style (#1599)
43724a6 ci : run builds only on source and build file changes (#1600)