Replies: 2 comments
|
As a Halo Strix user, I'd love to see this implemented. Many headaches and sanity checks have been made trying to do this myself a s a non-coder. I love this project and really would also love to see full support for AMD! |
0 replies
|
yeah this is more or less a must have feature... considering how great the 7900XTX is for running local models with ROCm it is a shame it is half baked when using this frontend |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Before I continue hammering away at
my sanitythe following problems and creating pull requests, I want to have some clarification on how the maintainers would want these issues to be fixed.ROCM
llama.cpp is unable to build with ROCM support
Only the GPU hardware is passed through the compose files (
/dev/kfd,/dev/dri/renderD128indocker/gpu.amd.yml), but no ROCm software is present. I'm aware this isn't a bug — the backend gate atroutes/cookbook_helpers.py:986only fires the HIP branch when a ROCm toolchain is detected, and the codebase deliberately reports"rocm"only when hipconfig exists (services/hwfit/hardware.py:292-295,routes/cookbook_routes.py:3016-3034). So a container with no ROCm userspace is supposed to fall back to Vulkan.The problem is that installing the ROCm stack still does not produce a working build — it just unblocks configure far enough to hit the two blockers below.
Solution(s)
apt-get install --no-install-recommends hipcc libhipblas-dev librocblas-dev # +4,5GBWith this, the HIP compiler is found (Clang 17.0.6) but configure still fails — see below.
SELinux blocks
/dev/kfdso ROCm cannot enumerate the GPUOn SELinux-enforcing hosts (Fedora/Bazzite/RHEL) the AMD overlay passes
/dev/kfdthrough, but the container domain (container_t) has no rule for the device's label (hsa_device_t), sorocminfois deniedmapand no ROCm agent is visible inside the container:This is separate from the HIP version gate below and would block a correctly built ROCm binary too. The Vulkan path is unaffected because
/dev/dri/renderD128isdri_device_t, which the container policy does allow.Solution(s)
The ROCm and Red Hat guidance for AMD GPUs in Podman is to allow containers to use host devices (
container_use_devicesis off by default):The bootstrap assumes the upstream
/opt/rocmlayout for the HIP compilerOnce hipconfig is on PATH the HIP branch fires, and configure then dies because the bootstrap derives the compiler as
HIPCXX="$(hipconfig -l)/clang"(cookbook_helpers.py:989). That only holds for AMD's/opt/rocmlayout; on Debian,hipconfig -lreturns/usr/binbut the amdgpu clang actually lives at/usr/lib/llvm-17/bin/clang, so configure fails with:Solution(s)
With that, the whole toolchain (hipcc, hipblas, rocblas, hsa-rocr, comgr, device libs at
/usr/lib/llvm-17/lib/clang/17/amdgcn/bitcode) resolves cleanly and configure proceeds to the version gate below.Current llama.cpp requires HIP version 6.1 which isn't in repos
As of writing this, Debian Trixie only has HIP 5.7 (5.7.1). Even with the HIPCXX fix, configure aborts at
ggml/src/ggml-hip/CMakeLists.txt:54-55:Solution(s)
A: pin llama.cpp to a version that requires HIP 5.7
B: Install official AMD ROCm ≥ 6.1 instead of the Debian package. The upstream
/opt/rocmlayout then satisfies the previous blockers at once:/opt/rocm/bin/hipconfigon PATH trips thecookbook_helpers.py:986gate, and$(hipconfig -l)/clangresolves to the real ROCm clang — so the HIPCXX/ln workaround above becomes unnecessary.AMD's apt repo (7.2.4; AMD's Debian 13 instructions map to the Ubuntu
noblerepo) can be registered inside the container with the steps from AMD's Debian native installation guide:With this, llama.cpp rebuilds with ROCm support and serves a working model, resolving both of the earlier build problems (no ROCm userspace in the image, and the HIPCXX
/opt/rocmlayout assumption).Unfortunately this is a large install (~8 GB). ROCm distributes prebuilt GEMM/solver kernels, and the bulk is two hard dependencies that upstream's
find_package(hipblas/rocblas REQUIRED)makes load-bearing:hipblaslt(~4.8 GB, depended on byrocblas) androcsolver(~1.7 GB, depended on byhipblas).Adding this in would require a refactor of the Dockerfiles/composefiles, eg. adding an Odysseus-ROCM image variant or something like that in order to not make the main image heavier than needed.
Vulkan
llama.cpp is unable to build with Vulkan support
The build starts, but
cmake configurefails due to missingglslc, SPIRV-Headers,libvulkan-dev.Solution(s)
I was able to build llama.cpp for Vulkan with the following:
apt-get install libvulkan-dev glslc spirv-headers # +44.6 MBBuild log seen later:
llama.cpp is unable to serve models using Vulkan
After fixing the Vulkan build above:
This is caused by missing Vulkan ICD.
Solution(s)
apt-get install mesa-vulkan-drivers # +83.1 MBAt this point I'm able to serve a GGUF model with llama.cpp using the Vulkan build.
The above fixes are only useful to users with AMD GPUs, because NVIDIA ships their own Vulkan ICD with the driver/toolkit.
Splitting the Odysseus container into different versions with one of them being AMD is probably an unavoidable path here.
AMD has their own equivalent of CUDA toolkit container, but it would have to be used by basing the odysseus image on it. This is not feasible, as that container is even bigger than the minimal package set required here.
Alternatively the packages could be put in as a dependency install in the Cookbook, though that can get tedious quick due to no persistence (system packages, not pip packages in /app/data/local) on container recreation. Native llama.cpp build works the same way so a possible consideration. Side note that llama.cpp build with ROCM support is a whole lot heavier/longer than other builds I've seen.
Universal
The build result/return code is ignored - the runner sees no binary and instead uses the default python shim that is a CPU-only wheel.
AI summary of this problem:
Solution(s)
Not ignoring the build failures silently.
Issue breakdown
This post is up to date as of commit 3b6c169.
This post was written half-manually half-assisted by AI.
All reactions