Releases: midagedev/bloomery
Release list
bloomery 0.2.0
bloomery 0.2.0 — one server binary, more models per card, concurrent requests, plans that fit the card you have.
Install
Linux x86-64 (glibc 2.34+), NVIDIA GPU of compute capability 8.6+ with its driver; on Windows, inside WSL2:
curl -fsSL https://raw.githubusercontent.com/midagedev/bloomery/main/tools/release/install.sh | sh
bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080or grab the tarball below and run bin/bloomery-serve as is. Homebrew: brew install midagedev/tap/bloomery.
Docker (needs the host's NVIDIA Container Toolkit):
docker run --gpus all -p 8080:8080 -v bloomery-cache:/root/.cache/bloomery \
ghcr.io/midagedev/bloomery --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_MWhat is new
- One binary.
bloomery_serve_clefis folded intobloomery-serveas its decide seat: with a model file, the
file's architecture picks the seat (the--modelword is optional;--helplists them with a working--hf
example each). Clef-Flash serves frombloomery-serve --hf bartowski/Cloudflare_clef-flash-GGUF:Q5_K_M— the
head is fetched from the repo the model card names — or--head <weights>./v1/systemonefollows llama.cpp's
wire: the request'smodelis optional,/v1/modelsanswers, images and videos are a 501,confidenceis
upstream's formula. A ggml-org Clef-layout GGUF is refused by name. Every server's default port is 8080. - Small cards run the big Qwens. Qwen3.6-35B and Qwen3-30B at
Q4_K_Mrun on a 12–16 GB card with--place a
(the routed experts on the CPU), and with no--placeat all the server plans the split itself when the file
does not fit the card's free bytes, its plan line saying why. - Plans that fit the card you actually have. The device census reads each card's free bytes; the expert share
sizes to them, the plan lines name them, and a plan whose trunk alone cannot fit is refused by name with every
term (dense, KV, context, scratch, margin) and the processes holding the card — no more opaque out-of-memory
after minutes of loading. - Concurrent requests on the ds41, glm and qwen3 seats:
--parallel N(two by default; a lone request pays
nothing for the second). Requests take turns of 64 tokens on one model; a later arrival preempts at the next
step. DSpark drafts do not rejoin a turn, so--parallel > 1withBLOOMERY_DRAFT=dsparkis refused by name.
The qwen38 and decide seats serve one request at a time (Qwen3.8's MTP draft does not rejoin a resume yet). - Prefix reuse everywhere it can be kept. V4.1 and Qwen3.8 as before; GLM-5.3 keeps a slot's state in a host
prompt cache (--cache-ram) with its draft rejoining where it left; Qwen3-30B keeps every held position and
Qwen3.6 keeps back to its recurrent checkpoints (one every 512 positions and each prompt call's end) — a 4k
conversation's turn N goes from ~450 ms to ~65 ms, an 8k one ~14x. - Faster GLM-5.3 prompts: the card experts run by tile items (each expert's rows read once per tile instead of
once per slot), pp4096 +8.5 % measured against 0.1.0's binary in one lease (216.9 → 235.4 tok/s, A6000).
The archive
bin/bloomery-serve alone. Requirements as 0.1.0: Linux x86-64 with glibc 2.34+, a CPU with AVX2 and FMA
(x86-64-v3), an NVIDIA GPU of compute capability 8.6 or newer and its driver — no CUDA toolkit, no Rust toolchain.
On Windows, run it inside WSL2 (the archive's README has the three-line install).
tar -xzf bloomery-0.2.0-linux-x86_64-cuda-sm86.tar.gz && cd bloomery-0.2.0-linux-x86_64-cuda-sm86
bin/bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080What was checked
- Built and run on an RTX A6000 and an RTX 3090 (Ubuntu 24.04, driver 615.71); the release binary answered a
SystemOne request from the packed archive before anything was tagged. - On an RTX 3060 12 GB under WSL2 (Windows driver 591.86, i7-12700F), from this archive:
--hffetched
Clef-FlashQ3_K_Sand its head over the network, and a second start resolved everything from the cache — the
answers to all 8 suite requests byte-identical to the A6000's (modulo themodelid,-mnaming the file and
--hfthe repo, and per-machinetimings). - The concurrency clauses: both interleaved requests answer their solo runs' ids exactly — V4.1 by
snapshot/resume, GLM with its draft rejoining, the qwen3 seats by re-prefill. - Every round landed under the repo's gate protocol (owning gates, FAIL-first mutants, PTX scans); the numbers
above come from the lease runners under the quiet-machine protocol.
Numbers and how each model is verified: the README. AI assistants
helped write bloomery and these notes.
bloomery 0.1.0
The first prebuilt build of bloomery for Linux x86-64 with an NVIDIA GPU. Download, extract and run. No CUDA toolkit
or Rust toolchain is needed.
What is in it
bin/bloomery-serve: llama-server's HTTP API (/v1/chat/completions,/completion, streaming). It serves
DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.6-35B-A3B and Qwen3-30B-A3B.bin/bloomery_serve_clef: Cloudflare Clef-Flash's SystemOne API (POST /v1/systemone).- Both take a model file (
-m) or a Hugging Face repo and quant (--hf <repo>[:<quant>]). A download resumes, is
checked against the upload's sha256, and is not fetched again.
tar -xzf bloomery-0.1.0-linux-x86_64-cuda-sm86.tar.gz && cd bloomery-0.1.0-linux-x86_64-cuda-sm86
bin/bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080
bin/bloomery_serve_clef --hf bartowski/Cloudflare_clef-flash-GGUF:Q5_K_M --port 8091Requirements
- glibc 2.34 or newer (Ubuntu 22.04+, Debian 12+, RHEL 9+), and a CPU with AVX2 and FMA (x86-64-v3).
- An NVIDIA GPU of compute capability 8.6 or newer and its driver. The archive carries sm_86 PTX, which the driver
compiles on the first start (10–25 s measured); later starts take a few seconds. - Qwen3-30B and Qwen3.6 at
Q4_K_Mneed a 24 GB card. V4.1, GLM-5.3 and Qwen3.8 keep experts on the CPU and need a
host with 256 GB of RAM. - Cards are found by device:
--place auses the visible card with the most memory, and--place bpadds the next
one as an expert tier.
What was checked
- Run on an RTX A6000 and an RTX 3090 (driver 615.71, Ubuntu 24.04).
- On an RTX 3060 under WSL2 (Windows driver 591.86, Intel i7-12700F), from this archive:
--hffetched Clef-Flash
Q3_K_Sand its head, and the server's answers to 15 requests equal the A6000's byte for byte. - Other cards, drivers and distributions are untested. Please open an issue with
--versionand the error text.
Numbers and how each model is verified: see the README. AI
assistants helped write bloomery and these notes.