Skip to content

bloomery 0.1.0

Choose a tag to compare

@midagedev midagedev released this 02 Oct 21:45

The first prebuilt build of bloomery for Linux x86-64 with an NVIDIA GPU. Download, extract and run. No CUDA toolkit
or Rust toolchain is needed.

What is in it

  • bin/bloomery-serve: llama-server's HTTP API (/v1/chat/completions, /completion, streaming). It serves
    DeepSeek-V4.1-Flash, GLM-5.3-Flash, Qwen3.8-Flash-Next, Qwen3.6-35B-A3B and Qwen3-30B-A3B.
  • bin/bloomery_serve_clef: Cloudflare Clef-Flash's SystemOne API (POST /v1/systemone).
  • Both take a model file (-m) or a Hugging Face repo and quant (--hf <repo>[:<quant>]). A download resumes, is
    checked against the upload's sha256, and is not fetched again.
tar -xzf bloomery-0.1.0-linux-x86_64-cuda-sm86.tar.gz && cd bloomery-0.1.0-linux-x86_64-cuda-sm86
bin/bloomery-serve --hf unsloth/Qwen3-30B-A3B-Instruct-2507-GGUF:Q4_K_M --port 8080
bin/bloomery_serve_clef --hf bartowski/Cloudflare_clef-flash-GGUF:Q5_K_M --port 8091

Requirements

  • glibc 2.34 or newer (Ubuntu 22.04+, Debian 12+, RHEL 9+), and a CPU with AVX2 and FMA (x86-64-v3).
  • An NVIDIA GPU of compute capability 8.6 or newer and its driver. The archive carries sm_86 PTX, which the driver
    compiles on the first start (10–25 s measured); later starts take a few seconds.
  • Qwen3-30B and Qwen3.6 at Q4_K_M need a 24 GB card. V4.1, GLM-5.3 and Qwen3.8 keep experts on the CPU and need a
    host with 256 GB of RAM.
  • Cards are found by device: --place a uses the visible card with the most memory, and --place bp adds the next
    one as an expert tier.

What was checked

  • Run on an RTX A6000 and an RTX 3090 (driver 615.71, Ubuntu 24.04).
  • On an RTX 3060 under WSL2 (Windows driver 591.86, Intel i7-12700F), from this archive: --hf fetched Clef-Flash
    Q3_K_S and its head, and the server's answers to 15 requests equal the A6000's byte for byte.
  • Other cards, drivers and distributions are untested. Please open an issue with --version and the error text.

Numbers and how each model is verified: see the README. AI
assistants helped write bloomery and these notes.