Releases: KoopahTManiac/speed.llama.cpp
Release list
b10373
speculative: never-worse scheduler - verify-all cold start, robust fits
Field regression on a 5090 (coding 600 -> 350-400 t/s) exposed three
defects in the self-calibrating scheduler:
- cold sessions ran admission on the head's raw pessimistic
confidences; worse, the resulting truncation censored the very
full-block rounds calibration needs, so the first fit was ~6000
rounds away. Uncalibrated mode now admits everything (same
convention as sglang's flat-SPS verify-all default): full fixed-depth
speed from round one, and every round feeds calibration, putting the
first fit ~1 long prompt away. - early small-window fits produced extreme temperatures (T=0.087 seen,
crushing deep-position survival). Fits now require >=32 samples of
each outcome class per position (identity otherwise) and shrink
geometrically toward identity by window size. - the admission scan now demands a meaningful predicted gain before
truncating (4% margin on the causal early stop, and full admission
is preferred whenever it scores within the margin of an interior
argmax): near-flat theta curves - easy content on fast GPUs - admit
everything, encoding the measured lesson that trading accepted
tokens for noise-level predicted savings always loses.
Also: profiled SPS curve forced monotone (Windows timing outliers put
size 1 above size 2), and the calibration cache version bumped to v2
so existing caches from the noisy-fit era are discarded on load.
Isolation on RTX PRO 6000 (GSM k8, fresh caches, 4 rounds): OFF 588
t/s mean vs sched+adapt 587 (parity, accept len 7.18 vs 7.08); cold
first round 600.6 (was 350-400 on the affected 5090 setup).
Website:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI:
b10372
speculative: self-calibrating scheduler with persistent cache
--spec-sts required a manual offline fitting pipeline; calibration is
now automatic. The server collects (confidence vector, realized
accepted length) pairs from full-block rounds into a ring buffer -
with an exploration round every 16th drafted round verifying the full
block, so scheduled truncation never censors deep positions out of the
data - and refits the paper-3.2.1 sequential cumprod-ECE grid search
in-process every 256 rounds (~3ms on the most recent 1024 rounds).
STS temperatures and the online cost model (t_fix, c_tok) persist to a
per-draft-model cache file (--spec-sched-cache, default 'auto' in the
llama cache dir), so later boots start warm. --spec-sts still accepts
fixed temperatures and disables auto-refit.
Smoke-tested: cold boot refits at 384/640 rounds and writes the cache;
warm boot loads it (sts=8, t_fix=6.2ms).
Website:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI:
b10365
ci: point release download links at the publishing repository
The release-notes template hardcoded the upstream repo, so every link
in a fork's release body pointed at ggml-org instead of the release's
own assets.
Website:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI:
b10364
ci: drop the s390x release build
The s390x partner runner routinely queues for hours and the release job
needs the whole ubuntu-cpu matrix, so it stalled every release.
Website:
macOS/iOS:
- macOS Apple Silicon (arm64)
- macOS Apple Silicon (arm64, KleidiAI enabled) DISABLED
- macOS Intel (x64)
- iOS XCFramework
Linux:
- Ubuntu x64 (CPU)
- Ubuntu arm64 (CPU)
- Ubuntu s390x (CPU)
- Ubuntu x64 (Vulkan)
- Ubuntu arm64 (Vulkan)
- Ubuntu x64 (ROCm 7.2)
- Ubuntu x64 (OpenVINO)
- Ubuntu x64 (SYCL FP32)
- Ubuntu x64 (SYCL FP16)
Android:
Windows:
- Windows x64 (CPU)
- Windows arm64 (CPU)
- Windows arm64 (OpenCL Adreno)
- Windows x64 (CUDA 12) - CUDA 12.4 DLLs
- Windows x64 (CUDA 13) - CUDA 13.3 DLLs
- Windows x64 (Vulkan)
- Windows x64 (OpenVINO)
- Windows x64 (SYCL)
- Windows x64 (HIP)
openEuler:
- DISABLED
- openEuler x86 (310p)
- openEuler x86 (910b, ACL Graph)
- openEuler aarch64 (310p)
- openEuler aarch64 (910b, ACL Graph)
UI: