Fixes first-start full-context initialization on an empty compile cache by changing the published GPU reservation default from 0.72 to 0.74. Model precision, TP=2, 131,072-token context, 4,096-token batch ceiling, DFlash15, XPU graph mode, and bounded vision are unchanged.
Published images (identical digest):
ghcr.io/rmacy/glimmer-b70-vllm:0.1.5us-central1-docker.pkg.dev/home-504803/open-models/glimmer-b70-vllm:0.1.5sha256:8fc1eb0d459a4602e39e0af6551858e667ce9856c0e5a8db41970c732eab5617
Acceptance on the exact public digest with a new cache: ready in 180 seconds, 148,010 KV tokens at full context, deterministic text pass, streaming and nonstreaming ATEM tool-call passes, tool-result continuation pass, and 1,792-by-1,792 maximum-image pass. The image contains no model weights or private configuration.
Trivy 0.73.0: zero secrets; the two critical records are duplicate version matches for CVE-2026-48746, whose upstream fix is backported and covered by the included regression test. See SECURITY.md for the inherited dependency baseline.