Skip to content

Releases: Yanichar/llama.cpp-cuda

llama.cpp v0.5.0 with CUDA

Choose a tag to compare

@github-actions github-actions released this 24 Sep 02:27

llama.cpp v0.5.0 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/v0.5.0
Commit: d2e5458

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-v0.5.0-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-v0.5.0-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp v0.4.1 with CUDA

Choose a tag to compare

@github-actions github-actions released this 15 Sep 02:38

llama.cpp v0.4.1 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/v0.4.1
Commit: 391fac1

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-v0.4.1-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-v0.4.1-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp v0.4.0 with CUDA

Choose a tag to compare

@github-actions github-actions released this 05 Sep 02:35

llama.cpp v0.4.0 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/v0.4.0
Commit: 427291b

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-v0.4.0-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-v0.4.0-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp v0.3.0 with CUDA

Choose a tag to compare

@github-actions github-actions released this 26 Aug 00:55

llama.cpp v0.3.0 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/v0.3.0
Commit: c1d0e7a

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-v0.3.0-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-v0.3.0-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp v0.2.0 with CUDA

Choose a tag to compare

@github-actions github-actions released this 22 Aug 00:53

llama.cpp v0.2.0 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/v0.2.0
Commit: 5a32f7b

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-v0.2.0-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-v0.2.0-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp b10531 with CUDA

Choose a tag to compare

@github-actions github-actions released this 21 Aug 00:56

llama.cpp b10531 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/b10531
Commit: f20395d

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-b10531-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-b10531-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp b10502 with CUDA

Choose a tag to compare

@github-actions github-actions released this 20 Aug 00:50

llama.cpp b10502 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/b10502
Commit: 0adcc3b

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-b10502-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-b10502-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp b10488 with CUDA

Choose a tag to compare

@github-actions github-actions released this 19 Aug 00:54

llama.cpp b10488 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/b10488
Commit: 9d77fa1

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-b10488-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-b10488-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp b10472 with CUDA

Choose a tag to compare

@github-actions github-actions released this 18 Aug 00:51

llama.cpp b10472 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/b10472
Commit: 60eeeb6

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-b10472-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-b10472-cuda-12.8-arm64.tar.gz
./llama-cli --help

llama.cpp b10453 with CUDA

Choose a tag to compare

@github-actions github-actions released this 17 Aug 00:54

llama.cpp b10453 with CUDA Support

Pre-built binaries of llama.cpp with CUDA support for multiple CUDA versions.

Source: https://github.com/ggml-org/llama.cpp/releases/tag/b10453
Commit: 3cb7ffb

CUDA Versions

  • CUDA 12.8 - GPU compute capability: 6.1

Host architectures

Tarballs are published per host CPU architecture (Linux):

  • -amd64.tar.gz — x86_64 (most desktops, servers, cloud VMs)
  • -arm64.tar.gz — aarch64 (Grace Hopper / Grace Blackwell / DGX Spark / Ampere Altra)

GPU compute capability reference

  • 6.1: Titan XP, Tesla P40, GTX 10xx (Pascal)

Usage

Download the tarball matching your host CPU arch and CUDA version, then extract:

# amd64 host
tar -xzf llama.cpp-b10453-cuda-12.8-amd64.tar.gz
# arm64 host (e.g. Grace Blackwell)
tar -xzf llama.cpp-b10453-cuda-12.8-arm64.tar.gz
./llama-cli --help