Skip to content

Releases: midagedev/bloomery

bloomery 0.2.0

Choose a tag to compare

@midagedev midagedev released this 03 Oct 11:17

bloomery 0.2.0 — one server binary, more models per card, concurrent requests, plans that fit the card you have.

Install

Linux x86-64 (glibc 2.34+), NVIDIA GPU of compute capability 8.6+ with its driver; on Windows, inside WSL2:

curl -fsSL https://raw.githubusercontent.com/midagedev/bloomery/main/tools/release/install.sh | sh
bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080

or grab the tarball below and run bin/bloomery-serve as is. Homebrew: brew install midagedev/tap/bloomery.
Docker (needs the host's NVIDIA Container Toolkit):

docker run --gpus all -p 8080:8080 -v bloomery-cache:/root/.cache/bloomery \
  ghcr.io/midagedev/bloomery --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M

What is new

  • One binary. bloomery_serve_clef is folded into bloomery-serve as its decide seat: with a model file, the
    file's architecture picks the seat (the --model word is optional; --help lists them with a working --hf
    example each). Clef-Flash serves from bloomery-serve --hf bartowski/Cloudflare_clef-flash-GGUF:Q5_K_M — the
    head is fetched from the repo the model card names — or --head <weights>. /v1/systemone follows llama.cpp's
    wire: the request's model is optional, /v1/models answers, images and videos are a 501, confidence is
    upstream's formula. A ggml-org Clef-layout GGUF is refused by name. Every server's default port is 8080.
  • Small cards run the big Qwens. Qwen3.6-35B and Qwen3-30B at Q4_K_M run on a 12–16 GB card with --place a
    (the routed experts on the CPU), and with no --place at all the server plans the split itself when the file
    does not fit the card's free bytes, its plan line saying why.
  • Plans that fit the card you actually have. The device census reads each card's free bytes; the expert share
    sizes to them, the plan lines name them, and a plan whose trunk alone cannot fit is refused by name with every
    term (dense, KV, context, scratch, margin) and the processes holding the card — no more opaque out-of-memory
    after minutes of loading.
  • Concurrent requests on the ds41, glm and qwen3 seats: --parallel N (two by default; a lone request pays
    nothing for the second). Requests take turns of 64 tokens on one model; a later arrival preempts at the next
    step. DSpark drafts do not rejoin a turn, so --parallel > 1 with BLOOMERY_DRAFT=dspark is refused by name.
    The qwen38 and decide seats serve one request at a time (Qwen3.8's MTP draft does not rejoin a resume yet).
  • Prefix reuse everywhere it can be kept. V4.1 and Qwen3.8 as before; GLM-5.3 keeps a slot's state in a host
    prompt cache (--cache-ram) with its draft rejoining where it left; Qwen3-30B keeps every held position and
    Qwen3.6 keeps back to its recurrent checkpoints (one every 512 positions and each prompt call's end) — a 4k
    conversation's turn N goes from ~450 ms to ~65 ms, an 8k one ~14x.
  • Faster GLM-5.3 prompts: the card experts run by tile items (each expert's rows read once per tile instead of
    once per slot), pp4096 +8.5 % measured against 0.1.0's binary in one lease (216.9 → 235.4 tok/s, A6000).

The archive

bin/bloomery-serve alone. Requirements as 0.1.0: Linux x86-64 with glibc 2.34+, a CPU with AVX2 and FMA
(x86-64-v3), an NVIDIA GPU of compute capability 8.6 or newer and its driver — no CUDA toolkit, no Rust toolchain.
On Windows, run it inside WSL2 (the archive's README has the three-line install).

tar -xzf bloomery-0.2.0-linux-x86_64-cuda-sm86.tar.gz && cd bloomery-0.2.0-linux-x86_64-cuda-sm86
bin/bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080

What was checked

  • Built and run on an RTX A6000 and an RTX 3090 (Ubuntu 24.04, driver 615.71); the release binary answered a
    SystemOne request from the packed archive before anything was tagged.
  • On an RTX 3060 12 GB under WSL2 (Windows driver 591.86, i7-12700F), from this archive: --hf fetched
    Clef-Flash Q3_K_S and its head over the network, and a second start resolved everything from the cache — the
    answers to all 8 suite requests byte-identical to the A6000's (modulo the model id, -m naming the file and
    --hf the repo, and per-machine timings).
  • The concurrency clauses: both interleaved requests answer their solo runs' ids exactly — V4.1 by
    snapshot/resume, GLM with its draft rejoining, the qwen3 seats by re-prefill.
  • Every round landed under the repo's gate protocol (owning gates, FAIL-first mutants, PTX scans); the numbers
    above come from the lease runners under the quiet-machine protocol.

Numbers and how each model is verified: the README. AI assistants
helped write bloomery and these notes.

bloomery 0.1.0

Choose a tag to compare

@midagedev midagedev released this 02 Oct 21:45

The first prebuilt build of bloomery for Linux x86-64 with an NVIDIA GPU. Download, extract and run. No CUDA toolkit
or Rust toolchain is needed.

What is in it

  • bin/bloomery-serve: llama-server's HTTP API (/v1/chat/completions, /completion, streaming). It serves
    DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.6-35B-A3B and Qwen3-30B-A3B.
  • bin/bloomery_serve_clef: Cloudflare Clef-Flash's SystemOne API (POST /v1/systemone).
  • Both take a model file (-m) or a Hugging Face repo and quant (--hf <repo>[:<quant>]). A download resumes, is
    checked against the upload's sha256, and is not fetched again.
tar -xzf bloomery-0.1.0-linux-x86_64-cuda-sm86.tar.gz && cd bloomery-0.1.0-linux-x86_64-cuda-sm86
bin/bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080
bin/bloomery_serve_clef --hf bartowski/Cloudflare_clef-flash-GGUF:Q5_K_M --port 8091

Requirements

  • glibc 2.34 or newer (Ubuntu 22.04+, Debian 12+, RHEL 9+), and a CPU with AVX2 and FMA (x86-64-v3).
  • An NVIDIA GPU of compute capability 8.6 or newer and its driver. The archive carries sm_86 PTX, which the driver
    compiles on the first start (10–25 s measured); later starts take a few seconds.
  • Qwen3-30B and Qwen3.6 at Q4_K_M need a 24 GB card. V4.1, GLM-5.3 and Qwen3.8 keep experts on the CPU and need a
    host with 256 GB of RAM.
  • Cards are found by device: --place a uses the visible card with the most memory, and --place bp adds the next
    one as an expert tier.

What was checked

  • Run on an RTX A6000 and an RTX 3090 (driver 615.71, Ubuntu 24.04).
  • On an RTX 3060 under WSL2 (Windows driver 591.86, Intel i7-12700F), from this archive: --hf fetched Clef-Flash
    Q3_K_S and its head, and the server's answers to 15 requests equal the A6000's byte for byte.
  • Other cards, drivers and distributions are untested. Please open an issue with --version and the error text.

Numbers and how each model is verified: see the README. AI
assistants helped write bloomery and these notes.