Skip to content

Paddock 0.1.9

Choose a tag to compare

@panterlo panterlo released this 24 Sep 18:29

A feature release: a diffusion language model with a structured-read API,
image generation and editing, a small MiniCPM5 model, Claude Code working
again with its current versions, and long-context fixes for Qwen 3.8
Flash-Next. Windows x64, Linux x64 and the NVIDIA DGX Spark. NVIDIA GPUs,
driver 580 or newer. The macOS pre-release for Apple Silicon includes the corrected Paddock
branding; its exact source is recorded below.

New

  • DiffusionGemma 26B A4B. Google's diffusion version of Gemma 4 26B
    writes a whole block of text per step instead of one token at a time.
    27 GB of weights at full quality, text only for now.

  • Structured reads. POST /v1/systemone asks DiffusionGemma a fixed set
    of yes/no, multiple-choice and scored questions about a text and returns
    every answer with its confidence, in the shape of the Jev API. The
    Studio's new Reads page builds the questions, runs them over a pasted or
    attached text and saves question sets for reuse.

  • Qwen-Image 2.1: pictures from a prompt, and edits of your own
    pictures.
    POST /v1/images/generations and POST /v1/images/edits, with
    previews streamed while the picture renders. In the Studio a picture is a
    conversation of its own: ask for one, then ask for changes. Full quality
    needs about 17 GB of VRAM and the compact build about 10 GB, plus room
    that grows with the picture. The Qwen Research License allows research and
    evaluation only, not commercial use.

  • MiniCPM5 2B, a small model with tool calling.

  • Whisper endpoints can load their model on demand. An endpoint set to
    load on first request stays up without holding GPU memory, loads the model
    when a transcription arrives and unloads it after the idle time you
    choose. Set it under Model loading in the endpoint's settings; the change
    applies without a restart. Other models stay loaded as before.

Improved

  • Qwen 3.8 Flash-Next at long context. Past about 2,000 tokens the model
    attends sparsely, and Paddock now serves it that way. Before, longer
    prompts got full attention, which is not the model's own behaviour. The
    full 262K context now loads, a long prompt no longer freezes the other
    conversations while it is read, and a continued conversation resumes from
    its cache at any depth.

  • Long prompts no longer stall other conversations on any model: while
    one conversation reads a long prompt, the others keep getting tokens at a
    steady pace.

  • Agent loops reuse more of their cache. The next turn of an agent
    resumes at the tool call instead of reading the whole previous reply
    again, and the model's earlier reasoning now reaches its template on every
    model, as the model expects.

  • Replies run to the context window when no limit is set. A request
    without max_tokens used to stop at 1,024 tokens, which cut reasoning
    models off before they answered.

  • Streams stay alive during long work. A long prompt or a long tool call
    no longer leaves a stream silent long enough for a client to hang up.

  • Gemma 4 serves 4-bit k-quant GGUF files such as Q4_K_M at 4 bits.

Fixed

  • Claude Code works again. Current Claude Code versions send system
    messages in the middle of a conversation, and every request failed with
    "invalid message role". They are accepted now.

  • Claude Code at long context. Token usage counts the cached prefix
    once, so Claude Code no longer compacts at half the window. A prompt past
    the window returns Anthropic's own "prompt is too long" error, so Claude
    Code compacts and carries on. Requests with tools keep speculating after
    the first turn.

macOS (pre-release)

Apple Silicon, macOS 26 or later. The native Swift app is pre-release;
visual glitches and incomplete features remain. The CLI's shared web Studio
is currently more complete.

  • DMG: drag Paddock to Applications for the native app with its bundled Metal runner.
  • CLI PKG: installs paddock and paddock-runner in /usr/local/bin.
    Run paddock and open http://localhost:11500 for the web Studio.
  • CLI tar.gz: the same command-line tools for manual installation.

The app and command-line tools are Developer ID signed. The app, DMG and
PKG are notarized; the app, DMG and PKG carry stapled tickets. Each download
has a separate .sha256 file.

macOS build 1902 comes from public commit
d4b45480:
v0.1.9 plus the corrected app/menu-bar/Studio branding and a sidebar layout
test correction. The shared v0.1.9 tag and existing Windows/Linux assets
are unchanged. Validated on an M5 Mac running macOS 26.6.1; M1–M4 and a clean
macOS 26.0 installation were not requalified for this build.

  • Image generation and editing with Qwen-Image on Metal.
  • DiffusionGemma on Metal, with the Reads page in the native app.
  • MiniCPM5 on Metal, and faster Bonsai.
  • Usage, activity and cache views, storage per model, saved settings
    profiles and on-demand Whisper loading in Settings.
  • Fixes to documents, transcripts, viewers and scrolling.

Known

  • DiffusionGemma on Metal: structured-read labels pass the release smoke
    tests, but confidence values are not bit-identical across batching or model
    reloads. These checks do not establish full numerical parity.

  • Switching Vision off does not unload the image tower when its file sits
    beside the model: the model still answers images, and the memory estimate
    does not count the tower (about 0.9 GB).

  • On-demand loading covers Whisper only.

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).


Paddock 0.1.9, built from fd34f33a on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

55598602d6c6ea6991de1d7a2faf3947d906fb60638aedd590ce5bbf47b58f4f  paddock-0.1.9-x86_64-windows.zip
96040231d2ebb8757f3b1a32233bc5868067017ef8624a1149f624680b33f151  paddock-0.1.9-x86_64-linux.tar.gz
e784a9c89778155d15ec9b77a888b160f9da72dcade510ba99223415b754d7ce  paddock-0.1.9-aarch64-linux.tar.gz