Repository navigation
b4400 Updated version from the end of 2024
Much newer than previously available versions (b2275 from 2024-02-26 or b1618 from 2023-12-07) this GPU assisted llama.cpp is actually 20% faster than pure CPU builds. The following 6 files are to edited:
- CMakeLists.txt 14
- ggml/CMakeLists.txt 274
- ggml/src/ggml-cuda/common.cuh 455
- ggml/src/ggml-cuda/fattn-common.cuh 632
- ggml/src/ggml-cuda/fattn-vec-f32.cuh 71
- ggml/src/ggml-cuda/template-instances/../fattn-vec-f16.cuh 73
And the first cmake command gets additional flags:
cmake -B build -DGGML_CUDA=ON -DLLAMA_CURL=ON -DCMAKE_CUDA_STANDARD=14 -DCMAKE_CUDA_STANDARD_REQUIRED=true -DGGML_CPU_ARM_ARCH=armv8-a -DGGML_NATIVE=off