b1618 First working release 2023-12-07
There are 5 lines to be added in ggml-cuda.cu. Detailed description in the article LLAMA.CPP on NVIDIA Jetson Nano: A Complete Guide.
Using nano ggml-cuda.cu add the following after line 82:
#if CUDA_VERSION < 1100
#define CUBLAS_TF32_TENSOR_OP_MATH CUBLAS_TENSOR_OP_MATH
#define CUBLAS_COMPUTE_16F CUDA_R_16F
#define CUBLAS_COMPUTE_32F CUDA_R_32F
#endifThen execute:
mkdir build && cd build
cmake .. -DLLAMA_CUBLAS=ON
make -j 2