Repository navigation
--gpus all shows the machine's NVIDIA GPUs inside an environment, on Linux as root, with the same flag as docker run --gpus all.
sudo range run --gpus all ghcr.io/ggml-org/llama.cpp:light-cuda-b11206 \
--mount hf://unsloth/gemma-3-270m-it-GGUF:/model -- \
llama-cli -m /model/gemma-3-270m-it-Q4_K_M.gguf -ngl 99 -st -p "Hi"- Range shows the host driver's libraries and tools, such as
libcuda.so.1andnvidia-smi, read-only at/usr/local/nvidia, where CUDA images look for them. - Range keeps the kernels CUDA compiles for your GPU. On a Tesla T4, this command answered in 3.1 s from the second run on. With Docker and nvidia-container-toolkit, every run took 112 to 127 s, because CUDA compiled the kernels again in each container.
- From nothing, the same command took 143.1 s with Range and 184.5 s with Docker.
curl -fsSL https://github.com/andreygrehov/range/releases/latest/download/range_$(uname -s)_$(uname -m).tar.gz | tar -xz
./range shell python:3.12