Skip to content

b4400 Updated version from the end of 2024

Choose a tag to compare

@kreier kreier released this 05 Apr 15:34
· 44 commits to main since this release

Much newer than previously available versions (b2275 from 2024-02-26 or b1618 from 2023-12-07) this GPU assisted llama.cpp is actually 20% faster than pure CPU builds. The following 6 files are to edited:

  • CMakeLists.txt 14
  • ggml/CMakeLists.txt 274
  • ggml/src/ggml-cuda/common.cuh 455
  • ggml/src/ggml-cuda/fattn-common.cuh 632
  • ggml/src/ggml-cuda/fattn-vec-f32.cuh 71
  • ggml/src/ggml-cuda/template-instances/../fattn-vec-f16.cuh 73

And the first cmake command gets additional flags:

cmake -B build -DGGML_CUDA=ON -DLLAMA_CURL=ON -DCMAKE_CUDA_STANDARD=14 -DCMAKE_CUDA_STANDARD_REQUIRED=true -DGGML_CPU_ARM_ARCH=armv8-a -DGGML_NATIVE=off