Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.
-
Updated
Aug 23, 2026 - Python
Strix Halo guide for AMD Ryzen AI MAX+ 395 / Radeon 8060S local LLM setup and benchmarks: Ollama, llama.cpp, Vulkan/RADV, ROCm, GGUF, and raw evidence.
Performance-tuned llama.cpp for AMD Strix Halo (gfx1151): FA + MoE-prefill fixes with a bundled current Mesa driver. Vulkan and HIP; portable dir, Docker, and distrobox.
vLLM Qwen 3.6-27B (AWQ-INT4) + DFlash speculative decoding on AMD Strix Halo (gfx1151 iGPU, 128 GB UMA, ROCm 7.13). 24.8 t/s single-stream, vision, tool calling, 256K context, OpenAI-compatible, Docker. Matches DGX Spark FP8+DFlash+MTP at a third of the cost. No CUDA.
Local text to textured GLB on an AMD Strix Halo iGPU (gfx1151): FLUX.2 klein, a Vulkan-only TRELLIS.2 engine, and humanoid auto-rigging with SkinTokens on ROCm. No Blender, no CUDA.
Docker Compose for llama.cpp GGUF servers on AMD Strix Halo: Qwen, Gemma, and Laguna packages (abliterated and quantized), stock Vulkan plus ROCmFP4/MTP and ROCmFPX, parallel slots, with prefill/decode and quality metrics measured on this rig.
vLLM + Qwen3.6-27B (BF16) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). Vision input, 256K context, /v1/responses with separated reasoning, via TheRock ROCm.
DeepSeek V4 Flash 284B on AMD Strix Halo (gfx1151) — up to 32 tok/s decode & ~250 tok/s prefill via ROCmFPX, DSpark & ROCm 7.2
llama.cpp + Qwen3.6-27B (Q8_0 GGUF) OpenAI-compatible inference server on AMD Strix Halo (Ryzen AI Max+ 395, gfx1151). 256K context, ~7.5 t/s decode via TheRock ROCm Docker.
Claude Code skill for AMD Strix Halo (Ryzen AI MAX+ 395) ML setup. Handles PyTorch installation (official wheels don't work with gfx1151), GTT memory config, and environment setup. Enables 30B parameter models.
A turnkey, fully-local AI workstation engineered for the AMD Ryzen AI Max+ 395. LLM inference, voice, document parsing, browser automation, agents — all on-device.
ROCmFPX llama.cpp fork for Windows 🏆 — native build, headless OpenAI-compatible server & benchmarks. Tested on AMD Strix Halo (gfx1151), runs on other GPUs too.
Run the official Kimi K3 MoE checkpoint on one 128 GB AMD Strix Halo box. ROCm-resident static weights, MXFP4 experts streamed from NVMe via io_uring. C engine, chat client, OpenAI-compatible server. Very experimental.
Optimized dual AMD Strix Halo (gfx1151) vLLM MoE inference: TP=2 over USB4, tuned int4 MoE kernel, ~8us interconnect, serialized serving
ComfyUI on AMD Strix Halo (RDNA 3.5 / gfx1151) via Docker. Ubuntu 26.04 LTS + uv-managed Python 3.12 + pinned TheRock ROCm 7.13 wheels. Fixes the silent CPU fallback Debian / Python 3.13 images hit on gfx1151.
Reproducible local-LLM benchmark harness: llama.cpp on AMD Strix Halo (gfx1151, Ryzen AI Max+ 395) and NVIDIA DGX Spark — frozen corpora, quality gates with unit tests, sealed run bundles. Apache-2.0
Production-qualified batch-1 Qwen3.6-35B-A3B BF16 inference engine for AMD Ryzen AI Max+ 395 on Linux
Reproducible RDNA reference for Meta Muse-Glimmer-30B — adapting MI-series ROCm recipes to Ryzen AI and Radeon with vLLM, llama.cpp, DFlash and auditable benchmarks.
Add a description, image, and links to the gfx1151 topic page so that developers can more easily learn about it.
To associate your repository with the gfx1151 topic, visit your repo's landing page and select "manage topics."