bloomery 0.1.0
The first prebuilt build of bloomery for Linux x86-64 with an NVIDIA GPU. Download, extract and run. No CUDA toolkit
or Rust toolchain is needed.
What is in it
bin/bloomery-serve: llama-server's HTTP API (/v1/chat/completions,/completion, streaming). It serves
DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.6-35B-A3B and Qwen3-30B-A3B.bin/bloomery_serve_clef: Cloudflare Clef-Flash's SystemOne API (POST /v1/systemone).- Both take a model file (
-m) or a Hugging Face repo and quant (--hf <repo>[:<quant>]). A download resumes, is
checked against the upload's sha256, and is not fetched again.
tar -xzf bloomery-0.1.0-linux-x86_64-cuda-sm86.tar.gz && cd bloomery-0.1.0-linux-x86_64-cuda-sm86
bin/bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080
bin/bloomery_serve_clef --hf bartowski/Cloudflare_clef-flash-GGUF:Q5_K_M --port 8091Requirements
- glibc 2.34 or newer (Ubuntu 22.04+, Debian 12+, RHEL 9+), and a CPU with AVX2 and FMA (x86-64-v3).
- An NVIDIA GPU of compute capability 8.6 or newer and its driver. The archive carries sm_86 PTX, which the driver
compiles on the first start (10–25 s measured); later starts take a few seconds. - Qwen3-30B and Qwen3.6 at
Q4_K_Mneed a 24 GB card. V4.1, GLM-5.3 and Qwen3.8 keep experts on the CPU and need a
host with 256 GB of RAM. - Cards are found by device:
--place auses the visible card with the most memory, and--place bpadds the next
one as an expert tier.
What was checked
- Run on an RTX A6000 and an RTX 3090 (driver 615.71, Ubuntu 24.04).
- On an RTX 3060 under WSL2 (Windows driver 591.86, Intel i7-12700F), from this archive:
--hffetched Clef-Flash
Q3_K_Sand its head, and the server's answers to 15 requests equal the A6000's byte for byte. - Other cards, drivers and distributions are untested. Please open an issue with
--versionand the error text.
Numbers and how each model is verified: see the README. AI
assistants helped write bloomery and these notes.