Skip to content

Gemma 4 31B FP8 MTP5 dual-B70 profile v0.1.0

Latest

Choose a tag to compare

@rmacy rmacy released this 14 Aug 06:34
· 1 commit to main since this release

First measured dual-B70 release.

  • 51.24 tok/s steady-state C1 median; 0.071 s median TTFT
  • 156.55 tok/s steady-state C4 aggregate median; about 0.174 s median TTFT
  • Cache-resistant cold C4: 96.25 tok/s aggregate; 1.400 s median TTFT
  • Orrery quality: 73.44/85, including full tool, retrieval, safety, and vision credit
  • OpenAI-compatible chat, tools, continuation, and image input verified
  • No model weights, credentials, private hostnames, or local paths included
  • GHCR and GAR resolve to sha256:a73434cc569716e69c5179268c2d95c87357d61e8a44d4dfffc37f2335f1000b

Pull:

docker pull ghcr.io/rmacy/gemma4-b70-vllm:0.1.0

docker pull us-central1-docker.pkg.dev/home-504803/open-models/gemma4-b70-vllm:0.1.0