llama.cpp v0.4.0 with CUDA
llama.cpp v0.4.0 with CUDA Support
Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.
Source: https://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Commit: 427291b
CUDA Versions
- CUDA 12.8 - GPU compute capability: 6.1
Host architectures
Tarballs are published per host CPU architecture (Linux):
-amd64.tar.gz— x86_64 (most desktops, servers, cloud VMs)-arm64.tar.gz— aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)
GPU compute capability reference
- 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)
Usage
Download the tarball matching your host CPU arch and CUDA version, then extract:
# amd64 host
tar -xzf llama.cpp-v0.4.0-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-v0.4.0-cuda-12.8-arm64.tar.gz
./llama-cli --help