Releases: truespar/paddock
Release list
Paddock 0.1.12
A tables release: Kumo Tabular, NVIDIA's model that predicts a missing
column from labelled rows, runs natively on NVIDIA GPUs and on the Mac, with
a Tables page in the Studio and the Mac app. Requests with many images or
long audio use far less memory, Gemma 4 no longer fails on the DGX Spark and
RTX 50-series cards with an f16 KV cache, and long Qwen 3.8 Flash-Next
prompts prefill several times faster. Windows x64, Linux x64 and the NVIDIA
DGX Spark. NVIDIA GPUs, driver 580 or newer. The macOS pre-release for Apple
Silicon is built from the same commit.
New
-
Kumo Tabular. Give it labelled rows and it predicts the missing column
of new ones - a class with its confidence, or a number with an 80% range -
with no training: the labelled rows are read in context on every call.
Small, medium and large, each for classification and for regression, in
the F32 weights NVIDIA ships, behind the same/v1/tabular/predictions
API on NVIDIA GPUs and on the Mac. -
Tables in the Studio and the Mac app. Paste or open a CSV, pick the
column to predict, and the rows with it empty come back filled in; copy
the result as CSV, or see the equivalent curl. Every table and its runs
are saved and shared between the web Studio and the Mac app, and each
table has its own link.
Improved
-
Pictures cost far less memory. The images in a prompt are encoded as
each prefill pass needs them, into memory the plan sets aside, instead of
all at once and held twice. On the DGX Spark an 18K-token prompt with a
picture peaks 1.3 GB above idle instead of 13.6 GB. This covers the Qwen
3.5-3.8 models, Gemma 4, Granite Vision and PaddleOCR-VL. -
A Qwen prompt with pictures no longer pauses other conversations while
its pictures are read: on the DGX Spark the longest pause another session
saw went from 15.4 s to 3.3 s. -
Long audio clips use a fixed amount of memory in Qwen3-ASR and Granite
Speech, with the same transcript as before. -
Qwen 3.8 Flash-Next prefills long prompts much faster on its NVFP4
versions - a 16.5K-token document question on the DGX Spark went from
85 s to under 15 s - keeps its prefix cache across agent sessions, and an
exact re-send of a long prompt resumes from its last walk instead of from
the start. -
An idle server hands the memory it freed back instead of holding it
until the next request. -
Quieter logs. A warning now means a client was held up or a real
fallback happened; long prompts and normal startup no longer warn.
Fixed
-
A Qwen 3.8 server on the DGX Spark could run out of memory after about a
day: startup set aside up to 20 GB it never used. Under memory pressure
the system now stops the runner, not the manager. -
Gemma 4 with an f16 KV cache failed every request on the DGX Spark and
RTX 50-series cards. It now serves there. -
A long Gemma 4 prompt with a picture could get a wrong answer when the
picture fell across two prefill passes. -
Safetensors checkpoints start from the Studio. The NVFP4 and FP8
versions of Qwen 3.8 Flash-Next, Nemotron and Granite were refused when
started from the Studio. -
Switching Vision off switches it off. A vision file beside the model
is no longer loaded anyway, and saving the endpoint through the Advanced
tab keeps Vision off. -
The start form offers document and image features only to models that
chat.
macOS (pre-release)
- Kumo Tabular runs natively on Metal, with a native Tables workspace.
- Qwen 3.8 Flash-Next decodes and prefills faster on Metal.
Known
-
Laya reads text only.
-
Laya's confidence on choices with more than ten options is not
calibrated: the English checkpoint ships an out-of-range temperature for
that case, which Paddock clamps, as Laya's own server does. -
Kumo Tabular predicts up to ten classes.
-
On Qwen 3.8 Flash-Next's NVFP4 versions a greedy speculative reply can
part from plain decoding at a near tie - the same model with slightly
different numerics, not a wrong answer. -
On-demand loading covers Whisper only.
-
The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5).
Paddock 0.1.12, built from 6d725da3 on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
38cdc1ab5cdbe939ac071ecf9372bbfbc446697b6e3a83c45136a7a2cb53cab6 paddock-0.1.12-x86_64-windows.zip
7542bf06fdbf7de4537ea7261bdd430191e5499495a93c8270818e4655e5c6d2 paddock-0.1.12-x86_64-linux.tar.gz
aae665bd78ff47ebf6486c9525ab7d35708bc85874b4febc2a41b8ff95e3d398 paddock-0.1.12-aarch64-linux.tar.gz
Paddock 0.1.11
A memory release: the Qwen 3.5-3.8 models, Nemotron and Qwen 3.8 Flash-Next
fit more context on the same card, and long agent conversations resume from
the cache on every turn. Nemotron gains speculative decoding for agent
traffic and decodes long contexts much faster. Windows x64, Linux x64 and the
NVIDIA DGX Spark. NVIDIA GPUs, driver 580 or newer. The macOS pre-release for
Apple Silicon is built from the same commit.
New
-
More context on the same card. The resume points a long conversation
picks up from used to take a fixed reserve before the context got any
memory - on a 24 GB card that alone held a 27B model to about 21K tokens
at one request. They now live inside the context cache itself and hand
their memory back as a conversation grows. Measured on the same card and
budget: Qwen 3.5 9B went from 23.8K to 41.6K tokens. This covers the
Qwen 3.5-3.8 hybrids, Nemotron and Qwen 3.8 Flash-Next; a roomy card still
keeps just as many resume points as before. -
Qwen 3.8 Flash-Next's cache is paged. Its memory is now planned inside
the endpoint's budget (it used to take fixed amounts whatever the budget
said), and a cached prefix is shared instead of copied, so resumed turns
reach their first token 10-12% sooner. -
Nemotron speculates on agent traffic with its DFlash and DSpark
drafters, including sampled and tool-calling requests and any number of
concurrent requests. On the DGX Spark, tool-carrying sampled requests went
from 80 to 118 tokens/s. DSpark, NVIDIA's recommended Spark drafter, is in
the catalog.
Improved
-
Nemotron decodes long contexts about four times faster - 12.5 to 49.8
tokens/s at 259K. -
Long Nemotron agent conversations resume from the cache on every turn.
Past about 140K tokens every turn used to start over from the system
prompt. -
Bonsai runs faster on the DGX Spark, with its ternary weights on the
tensor cores at every batch size and an fp8 KV cache by default.
Fixed
-
A Qwen 3.8 Flash-Next reply generated while speculating keeps its resume
points, so the next turn picks up at the end of the reply instead of
reading the whole reply again. -
A vision or drafter file beside the weights is checked against the model
before it loads. A folder several models share no longer hands one model
another's file, and a mismatched file named in the configuration is
refused with the reason. -
A
vram_budgetwritten in GiB instead of MiB is refused with the value
it should have been. -
Prefill on the DGX Spark no longer fails for the Qwen 3.5-3.8
full-attention layers with an fp8 KV cache.
macOS (pre-release)
- Laya runs natively on Metal, with the native Reads page.
- Qwen 3.8 Flash-Next runs faster on Metal and reuses cached prefixes across
agent turns. - Qwen agent-turn resume points are kept on Metal.
Known
-
Laya reads text only.
-
Laya's confidence on choices with more than ten options is not
calibrated: the English checkpoint ships an out-of-range temperature for
that case, which Paddock clamps, as Laya's own server does. -
Switching Vision off does not unload the image tower when its file sits
beside the model: the model still answers images, and the memory estimate
does not count the tower (about 0.9 GB). -
On-demand loading covers Whisper only.
-
The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5).
Paddock 0.1.11, built from f99cacdf on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
07b866f3ad31aa916ef2e2604c5bf2f7fe046c73ac7473b99def550bb66e55bc paddock-0.1.11-x86_64-windows.zip
82411193a0f037d8f583a619c6b4edf6941660f027364117b96894714cb6d1d5 paddock-0.1.11-x86_64-linux.tar.gz
8b7ee3aaa2bc1d9d7273d0136544871da68b8042c6831046b5a9fd83057dcb59 paddock-0.1.11-aarch64-linux.tar.gz
Paddock 0.1.10
A feature release: Laya, an open model built only to answer questions about
a text, a much richer Reads page, image input for DiffusionGemma's reads, and
faster Qwen 3.8 Flash-Next. Windows x64, Linux x64 and the NVIDIA DGX Spark.
NVIDIA GPUs, driver 580 or newer. The macOS pre-release for Apple Silicon is
built from the same commit.
New
-
Laya, a decision model. Laya (ConvAI Innovations, Apache 2.0) answers
yes/no, multiple-choice and scored questions about a text with a calibrated
probability each, in one pass per question and without writing any text -
a few milliseconds a question. It serves the samePOST /v1/systemoneas
DiffusionGemma, so the Reads page works with it unchanged. The download
(2.4 GB) holds an English model, a multilingual one for 100+ languages and
a typed-decisions one; Paddock picks between the first two from the text
itself, as Laya's own server does. A text longer than one question's window
is read in overlapping windows instead of being cut. -
Images in reads. DiffusionGemma reads a question set over pictures as
well as text, with its own vision companion (1.2 GB) built from Google's
checkpoint. Attach pictures on the Reads page by button, drop or paste. -
Conditional questions. A question can be asked only when an earlier
answer is one of the values you pick, or read with an earlier question's
answer already known. A question that is skipped answersnulland says
why. -
More ways to read. Several denoising steps with the answer template
held in place, a thought written before the answers, and repeated reads
that report how much the answer moves between them.
Improved
-
The Reads page keeps its history in a side panel beside the questions,
stored by the manager like conversations, and stays in the sidebar with no
model running so earlier reads open any time. -
DiffusionGemma's compact build starts on NVIDIA GPUs.
-
Qwen 3.8 Flash-Next decodes faster, with and without speculative
decoding, and requests that carry tools now speculate too. -
Paddock's own app icon in the Studio, the Windows executables and the
menu bar.
Fixed
-
Flash-Next greedy decoding with speculation no longer drifts from the
answer it gives without speculation on a small share of requests. -
The Studio over the network opens at the key prompt again instead of
reporting that it could not open its saved data.
macOS (pre-release)
- DiffusionGemma reads images on Metal.
- The native Reads page takes pictures by paste and import, keeps its history
in the shared database the web Studio uses, and navigates it from the
sidebar. - Fixes to notification previews and benchmark forms.
Known
-
Laya runs on NVIDIA GPUs only for now, and reads text only.
-
Laya's confidence on choices with more than ten options is not
calibrated: the English checkpoint ships an out-of-range temperature for
that case, which Paddock clamps, as Laya's own server does. -
Switching Vision off does not unload the image tower when its file sits
beside the model: the model still answers images, and the memory estimate
does not count the tower (about 0.9 GB). -
On-demand loading covers Whisper only.
-
The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5).
Paddock 0.1.10, built from 610cd8f5 on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
3a0949a92bab20c15f37c8a5c29b1f8d80ecbeda23111f90ac4c4fa65208f97d paddock-0.1.10-x86_64-windows.zip
ca0cdce0c59520eb3a4f9e5930144222d9ad8a630f9909676d2bfb83ba90873b paddock-0.1.10-x86_64-linux.tar.gz
9f50baf2c14b22df39cd970792f02218eb92679f4c1dea98815d2168c3a2ab99 paddock-0.1.10-aarch64-linux.tar.gz
Paddock 0.1.9
A feature release: a diffusion language model with a structured-read API,
image generation and editing, a small MiniCPM5 model, Claude Code working
again with its current versions, and long-context fixes for Qwen 3.8
Flash-Next. Windows x64, Linux x64 and the NVIDIA DGX Spark. NVIDIA GPUs,
driver 580 or newer. The macOS pre-release for Apple Silicon includes the corrected Paddock
branding; its exact source is recorded below.
New
-
DiffusionGemma 26B A4B. Google's diffusion version of Gemma 4 26B
writes a whole block of text per step instead of one token at a time.
27 GB of weights at full quality, text only for now. -
Structured reads.
POST /v1/systemoneasks DiffusionGemma a fixed set
of yes/no, multiple-choice and scored questions about a text and returns
every answer with its confidence, in the shape of the Jev API. The
Studio's new Reads page builds the questions, runs them over a pasted or
attached text and saves question sets for reuse. -
Qwen-Image 2.1: pictures from a prompt, and edits of your own
pictures.POST /v1/images/generationsandPOST /v1/images/edits, with
previews streamed while the picture renders. In the Studio a picture is a
conversation of its own: ask for one, then ask for changes. Full quality
needs about 17 GB of VRAM and the compact build about 10 GB, plus room
that grows with the picture. The Qwen Research License allows research and
evaluation only, not commercial use. -
MiniCPM5 2B, a small model with tool calling.
-
Whisper endpoints can load their model on demand. An endpoint set to
load on first request stays up without holding GPU memory, loads the model
when a transcription arrives and unloads it after the idle time you
choose. Set it under Model loading in the endpoint's settings; the change
applies without a restart. Other models stay loaded as before.
Improved
-
Qwen 3.8 Flash-Next at long context. Past about 2,000 tokens the model
attends sparsely, and Paddock now serves it that way. Before, longer
prompts got full attention, which is not the model's own behaviour. The
full 262K context now loads, a long prompt no longer freezes the other
conversations while it is read, and a continued conversation resumes from
its cache at any depth. -
Long prompts no longer stall other conversations on any model: while
one conversation reads a long prompt, the others keep getting tokens at a
steady pace. -
Agent loops reuse more of their cache. The next turn of an agent
resumes at the tool call instead of reading the whole previous reply
again, and the model's earlier reasoning now reaches its template on every
model, as the model expects. -
Replies run to the context window when no limit is set. A request
withoutmax_tokensused to stop at 1,024 tokens, which cut reasoning
models off before they answered. -
Streams stay alive during long work. A long prompt or a long tool call
no longer leaves a stream silent long enough for a client to hang up. -
Gemma 4 serves 4-bit k-quant GGUF files such as Q4_K_M at 4 bits.
Fixed
-
Claude Code works again. Current Claude Code versions send system
messages in the middle of a conversation, and every request failed with
"invalid message role". They are accepted now. -
Claude Code at long context. Token usage counts the cached prefix
once, so Claude Code no longer compacts at half the window. A prompt past
the window returns Anthropic's own "prompt is too long" error, so Claude
Code compacts and carries on. Requests with tools keep speculating after
the first turn.
macOS (pre-release)
Apple Silicon, macOS 26 or later. The native Swift app is pre-release;
visual glitches and incomplete features remain. The CLI's shared web Studio
is currently more complete.
- DMG: drag Paddock to Applications for the native app with its bundled Metal runner.
- CLI PKG: installs
paddockandpaddock-runnerin/usr/local/bin.
Runpaddockand open http://localhost:11500 for the web Studio. - CLI tar.gz: the same command-line tools for manual installation.
The app and command-line tools are Developer ID signed. The app, DMG and
PKG are notarized; the app, DMG and PKG carry stapled tickets. Each download
has a separate .sha256 file.
macOS build 1902 comes from public commit
d4b45480:
v0.1.9 plus the corrected app/menu-bar/Studio branding and a sidebar layout
test correction. The shared v0.1.9 tag and existing Windows/Linux assets
are unchanged. Validated on an M5 Mac running macOS 26.6.1; M1–M4 and a clean
macOS 26.0 installation were not requalified for this build.
- Image generation and editing with Qwen-Image on Metal.
- DiffusionGemma on Metal, with the Reads page in the native app.
- MiniCPM5 on Metal, and faster Bonsai.
- Usage, activity and cache views, storage per model, saved settings
profiles and on-demand Whisper loading in Settings. - Fixes to documents, transcripts, viewers and scrolling.
Known
-
DiffusionGemma on Metal: structured-read labels pass the release smoke
tests, but confidence values are not bit-identical across batching or model
reloads. These checks do not establish full numerical parity. -
Switching Vision off does not unload the image tower when its file sits
beside the model: the model still answers images, and the memory estimate
does not count the tower (about 0.9 GB). -
On-demand loading covers Whisper only.
-
The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5).
Paddock 0.1.9, built from fd34f33a on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
55598602d6c6ea6991de1d7a2faf3947d906fb60638aedd590ce5bbf47b58f4f paddock-0.1.9-x86_64-windows.zip
96040231d2ebb8757f3b1a32233bc5868067017ef8624a1149f624680b33f151 paddock-0.1.9-x86_64-linux.tar.gz
e784a9c89778155d15ec9b77a888b160f9da72dcade510ba99223415b754d7ce paddock-0.1.9-aarch64-linux.tar.gz
Paddock 0.1.8
A feature release: a new model that puts 27B on a 16 GB card, a fix for image
requests under load, and a smoother Studio. Windows x64, Linux x64 and the
NVIDIA DGX Spark. NVIDIA GPUs, driver 580 or newer.
New
-
Bonsai 2 27B: a 27B model on a 16 GB card. Prism ML's ternary build of
Qwen 3.8 27B is 5.9 GB of weights and is served exactly as shipped. Image
input is a separate 0.9 GB download, on by default. A 16 GB card holds about
43,000 tokens of context with text only and about 27,000 with images; a
24 GB card reaches the full 256K with the 8-bit KV cache. Ternary weights
are close to the original model, not identical to it. -
Qwen 3.8 Flash-Next speculates for sampled requests. The drafter used to
help only requests that did not sample. Requests with a temperature, which
is the default, now speculate too. -
Qwen 3.8 Flash-Next has an NVFP4 build for Blackwell cards. The catalog
offers the distilled NVFP4 checkpoint for RTX 50-series class GPUs, and the
engine reads MX-quantized checkpoints as they ship. -
An image segmentation endpoint.
POST /v1/segmentationstakes image
chips and returns rasters from a DINOv3-based segmentation checkpoint served
by path. No catalog model uses it yet.
Improved
-
Long chats stay smooth in the Studio. Formatting a long conversation now
runs in the background and is spread across frames, so scrolling and
streaming no longer stall on large replies, tables and diagrams. -
Cloud models get reply lengths sized from the provider's own token
counts, not from an estimate. -
Conversation summaries hold up under pressure. Four fixes contributed by
@DivyamTalwar: a window with no room left is no longer sent the whole
transcript, a model switch mid-summary no longer saves the old model's
summary under the new one, a stalled summary request times out, and an
injected summary counts against the reply budget.
Fixed
-
Several image requests at once could all answer with garbage. On the
Qwen 3.5, 3.6 and 3.8 vision models, two or more image requests arriving
together, each with a prompt of 128 tokens or more, could all return a run
of "!" in place of an answer. A single request at a time was never affected. -
Qwen 3.8 Flash-Next tool calls with the IQ3 build. Tool calls came back
as plain text,tool_choice: "required"was refused and thinking budgets
were rejected. The engine now reads the tool-call format from the model's
own template when it has no entry of its own.
Known
-
Switching Vision off does not unload the image tower when its file sits
beside the model: the model still answers images, and the memory estimate
does not count the tower (about 0.9 GB). -
The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5).
Paddock 0.1.8, built from 1b321df5 on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: linked in (sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a), nvcc 13.3.73.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
b95fdc25cb83f237543592dc3ae99aba062c06e59863990dc7e96cd6a267db18 paddock-0.1.8-x86_64-windows.zip
1db2ec41a280ca8dc57f82749bed41521e70d02e126075d7744fc99877ddaddc paddock-0.1.8-x86_64-linux.tar.gz
6a3689efb185fe0c06f7944d018266cdf930cea02c1c9cfad78c34122fdd2624 paddock-0.1.8-aarch64-linux.tar.gz
Paddock 0.1.7
Paddock 0.1.7, built from 98e6703f on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: linked in (sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a), nvcc 13.3.73.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
f2edd3375864776c1fff1f14e487294fe9ec0da3c293bf4287fb5c84e468c292 paddock-0.1.7-x86_64-windows.zip
73b18b7db01e7420d8117e3f1825d7367809a2058b5c462f1609aece9698bdaa paddock-0.1.7-x86_64-linux.tar.gz
92b22246c9e64bea9d18658348f88ceaab818e3906aaf4cf14088563795e7b0e paddock-0.1.7-aarch64-linux.tar.gz
Paddock 0.1.6
A feature release. Windows x64, Linux x64 and, new in this release, the NVIDIA
DGX Spark. NVIDIA GPUs, driver 580 or newer.
New
-
The DGX Spark is supported. The GB10 in NVIDIA's DGX Spark joins the
supported GPUs, with its own Linux build for the Spark (DGX OS 7 / Ubuntu
24.04 or newer). Its kernels were tuned on the Spark and the model catalog
was tested on it. -
Gemma 4 reads small print. A new endpoint setting,
max_image_tokens,
sets how much detail an image carries. Gemma 4's default is too coarse for
fine print on a full page; raising it reads the figures correctly. More
detail costs prompt tokens and time per image, so the default is unchanged. -
Conversations resume from the last reply. Qwen 3.5, 3.6 and 3.8,
Nemotron 3.5 Lightning and Qwen 3.8 Flash-Next pick the next turn up from the
end of the previous reply instead of reading the conversation again, and
many conversations running at once each keep their place. -
Qwen 3.8 Flash-Next caches repeated prompts.
Improved
-
Gemma 4 keeps to its memory budget. With a tight
vram_budgetit now
makes room for the conversation instead of going past the budget. -
Nemotron 3.5 Lightning and 4-bit NVFP4 models are faster on the DGX
Spark, both reading prompts and generating text. -
The Studio opens quickly even with a very large chat history, and a
stopped model keeps its logo and name.
Fixed
- Opening a Studio chat before it had finished loading could save an empty
copy over its history.
Known
- The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5).
Paddock 0.1.5
A feature release. Windows x64 and Linux x64, NVIDIA GPUs, driver 580 or newer.
The theme is small cards: 1-, 2- and 3-bit files now serve, and the largest
mixture-of-experts models reach a 16 GB card through expert offload.
New
-
1-, 2- and 3-bit GGUF files serve. The ggml i-quant family (IQ1_S,
IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL) and the Q2_K and
Q3_K formats load and serve end to end, for dense weights and for MoE
experts alike, with decoders written here from the format specification.
Prompt processing on these files runs on the same tiled GEMM as the 4-bit
files, so a long prompt costs the same as it does at 4 bits. A file is
named by its real bit width everywhere it appears. Contributed by
NodeNestor (#17, #21). -
Qwen 3.8 Flash-Next serves from GGUF on RTX 50-series cards. The
mixture-of-experts lane loads unsloth's UD files, and with expert offload
the UD-IQ1_S file serves on a 16 GB card at llama.cpp parity. Prompt
processing streams the offloaded experts through the cache expert by
expert: a 427-token prompt went from 308 s to 37 s, and the first token of
a 1,500-word prompt from 31 s to 15 s, on an RTX 5060 Ti. Blackwell only in
this release; the family's kernels are built for sm_100 and sm_120.
Contributed by NodeNestor (#18). -
Qwen 3.8 27B has a Compact lane in the catalog: the UD-Q3_K_XL file,
for cards where the standard file does not fit.
Improved
-
Shared-prefix serving under load. Requests that arrive with the same
prefix, the shape of an agent fleet, are now recognised at admission and
warmed in one batched wave instead of each recomputing the prefix. With
eight concurrent sessions on a shared prompt the time to first token fell
from 1184 ms to 248 ms and the p99 gap between tokens from 88 ms to
31 ms. -
Small cards decode faster with many streams. The k-quant K-split path
now takes batched decode, so on an RTX 5060 Ti eight concurrent streams of
a Qwen 3.5/3.6 k-quant file step in 17.8 ms instead of 30.0 ms, twice
llama.cpp's aggregate on the same card. -
Memory planned by demand. Prompt-processing scratch is profiled at
load, and state checkpoints and cache retention are sized by what the
configuration actually needs, across Qwen 3.5/3.6, Gemma 4, Granite,
Laguna and Nemotron. Laguna serves batched. Agraph_scratch_mibsetting
in the Advanced tab reserves extra scratch by hand (ErikBPF, #14). -
A GPU the engine has not validated serves under a startup warning
instead of a refusal.
Fixed
- Blackwell kernels on Windows could fail: the Windows build laid out
tensor-map kernel parameters at 8-byte alignment where the hardware wants
128 (#6). - The multi-column weight kernel faulted on narrow planes at two concurrent
streams on sm_120 (#19, reported by L4GN). - A dense i-quant kernel's launch grid overflowed past 134 million outputs.
- Qwen 3.5/3.6: a failed attempt to enable batching is released before the
width ladder retries (ErikBPF, #15).
Known
- The fp8 KV cache's paged attention on RTX 50-series cards shows a small
numeric deviation in one split configuration (#5). Being fixed on an fp8
card.
Paddock 0.1.4
Paddock 0.1.4, built from ac58eded on the maintainers' machines (Windows natively, Linux in a Docker builder), each archive booted on a GPU and asked a question before it was uploaded.
Requirements
- An NVIDIA card the kernel pack covers: linked in (sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a), nvcc 13.3.73 - RTX 30/40/50 series, the RTX A / PRO line, B200. Other cards get an honest refusal, never a slow fallback.
- A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside
paddock-runner. - Windows 11 x64, or Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+).
Run
- Unzip or untar anywhere.
paddockstarts the manager and the Studio on http://localhost:11500.paddock-runner --model <file.gguf>serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.
sha256
40f7df7f17cf4b1f00cc51a6261ceda75c718d470727a84a9e9958178062221e paddock-0.1.4-x86_64-windows.zip
d2a62ddf032b0bc8c97c324c8c2f1f471ba323470dea44db0a02befac230a19c paddock-0.1.4-x86_64-linux.tar.gz