Skip to content

Releases: truespar/paddock

Paddock 0.1.12

Choose a tag to compare

@panterlo panterlo released this 01 Oct 11:32

A tables release: Kumo Tabular, NVIDIA's model that predicts a missing
column from labelled rows, runs natively on NVIDIA GPUs and on the Mac, with
a Tables page in the Studio and the Mac app. Requests with many images or
long audio use far less memory, Gemma 4 no longer fails on the DGX Spark and
RTX 50-series cards with an f16 KV cache, and long Qwen 3.8 Flash-Next
prompts prefill several times faster. Windows x64, Linux x64 and the NVIDIA
DGX Spark. NVIDIA GPUs, driver 580 or newer. The macOS pre-release for Apple
Silicon is built from the same commit.

New

  • Kumo Tabular. Give it labelled rows and it predicts the missing column
    of new ones - a class with its confidence, or a number with an 80% range -
    with no training: the labelled rows are read in context on every call.
    Small, medium and large, each for classification and for regression, in
    the F32 weights NVIDIA ships, behind the same /v1/tabular/predictions
    API on NVIDIA GPUs and on the Mac.

  • Tables in the Studio and the Mac app. Paste or open a CSV, pick the
    column to predict, and the rows with it empty come back filled in; copy
    the result as CSV, or see the equivalent curl. Every table and its runs
    are saved and shared between the web Studio and the Mac app, and each
    table has its own link.

Improved

  • Pictures cost far less memory. The images in a prompt are encoded as
    each prefill pass needs them, into memory the plan sets aside, instead of
    all at once and held twice. On the DGX Spark an 18K-token prompt with a
    picture peaks 1.3 GB above idle instead of 13.6 GB. This covers the Qwen
    3.5-3.8 models, Gemma 4, Granite Vision and PaddleOCR-VL.

  • A Qwen prompt with pictures no longer pauses other conversations while
    its pictures are read: on the DGX Spark the longest pause another session
    saw went from 15.4 s to 3.3 s.

  • Long audio clips use a fixed amount of memory in Qwen3-ASR and Granite
    Speech, with the same transcript as before.

  • Qwen 3.8 Flash-Next prefills long prompts much faster on its NVFP4
    versions - a 16.5K-token document question on the DGX Spark went from
    85 s to under 15 s - keeps its prefix cache across agent sessions, and an
    exact re-send of a long prompt resumes from its last walk instead of from
    the start.

  • An idle server hands the memory it freed back instead of holding it
    until the next request.

  • Quieter logs. A warning now means a client was held up or a real
    fallback happened; long prompts and normal startup no longer warn.

Fixed

  • A Qwen 3.8 server on the DGX Spark could run out of memory after about a
    day:
    startup set aside up to 20 GB it never used. Under memory pressure
    the system now stops the runner, not the manager.

  • Gemma 4 with an f16 KV cache failed every request on the DGX Spark and
    RTX 50-series cards.
    It now serves there.

  • A long Gemma 4 prompt with a picture could get a wrong answer when the
    picture fell across two prefill passes.

  • Safetensors checkpoints start from the Studio. The NVFP4 and FP8
    versions of Qwen 3.8 Flash-Next, Nemotron and Granite were refused when
    started from the Studio.

  • Switching Vision off switches it off. A vision file beside the model
    is no longer loaded anyway, and saving the endpoint through the Advanced
    tab keeps Vision off.

  • The start form offers document and image features only to models that
    chat.

macOS (pre-release)

  • Kumo Tabular runs natively on Metal, with a native Tables workspace.
  • Qwen 3.8 Flash-Next decodes and prefills faster on Metal.

Known

  • Laya reads text only.

  • Laya's confidence on choices with more than ten options is not
    calibrated:
    the English checkpoint ships an out-of-range temperature for
    that case, which Paddock clamps, as Laya's own server does.

  • Kumo Tabular predicts up to ten classes.

  • On Qwen 3.8 Flash-Next's NVFP4 versions a greedy speculative reply can
    part from plain decoding at a near tie
    - the same model with slightly
    different numerics, not a wrong answer.

  • On-demand loading covers Whisper only.

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).


Paddock 0.1.12, built from 6d725da3 on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

38cdc1ab5cdbe939ac071ecf9372bbfbc446697b6e3a83c45136a7a2cb53cab6  paddock-0.1.12-x86_64-windows.zip
7542bf06fdbf7de4537ea7261bdd430191e5499495a93c8270818e4655e5c6d2  paddock-0.1.12-x86_64-linux.tar.gz
aae665bd78ff47ebf6486c9525ab7d35708bc85874b4febc2a41b8ff95e3d398  paddock-0.1.12-aarch64-linux.tar.gz

Paddock 0.1.11

Choose a tag to compare

@panterlo panterlo released this 28 Sep 10:30

A memory release: the Qwen 3.5-3.8 models, Nemotron and Qwen 3.8 Flash-Next
fit more context on the same card, and long agent conversations resume from
the cache on every turn. Nemotron gains speculative decoding for agent
traffic and decodes long contexts much faster. Windows x64, Linux x64 and the
NVIDIA DGX Spark. NVIDIA GPUs, driver 580 or newer. The macOS pre-release for
Apple Silicon is built from the same commit.

New

  • More context on the same card. The resume points a long conversation
    picks up from used to take a fixed reserve before the context got any
    memory - on a 24 GB card that alone held a 27B model to about 21K tokens
    at one request. They now live inside the context cache itself and hand
    their memory back as a conversation grows. Measured on the same card and
    budget: Qwen 3.5 9B went from 23.8K to 41.6K tokens. This covers the
    Qwen 3.5-3.8 hybrids, Nemotron and Qwen 3.8 Flash-Next; a roomy card still
    keeps just as many resume points as before.

  • Qwen 3.8 Flash-Next's cache is paged. Its memory is now planned inside
    the endpoint's budget (it used to take fixed amounts whatever the budget
    said), and a cached prefix is shared instead of copied, so resumed turns
    reach their first token 10-12% sooner.

  • Nemotron speculates on agent traffic with its DFlash and DSpark
    drafters, including sampled and tool-calling requests and any number of
    concurrent requests. On the DGX Spark, tool-carrying sampled requests went
    from 80 to 118 tokens/s. DSpark, NVIDIA's recommended Spark drafter, is in
    the catalog.

Improved

  • Nemotron decodes long contexts about four times faster - 12.5 to 49.8
    tokens/s at 259K.

  • Long Nemotron agent conversations resume from the cache on every turn.
    Past about 140K tokens every turn used to start over from the system
    prompt.

  • Bonsai runs faster on the DGX Spark, with its ternary weights on the
    tensor cores at every batch size and an fp8 KV cache by default.

Fixed

  • A Qwen 3.8 Flash-Next reply generated while speculating keeps its resume
    points
    , so the next turn picks up at the end of the reply instead of
    reading the whole reply again.

  • A vision or drafter file beside the weights is checked against the model
    before it loads.
    A folder several models share no longer hands one model
    another's file, and a mismatched file named in the configuration is
    refused with the reason.

  • A vram_budget written in GiB instead of MiB is refused with the value
    it should have been.

  • Prefill on the DGX Spark no longer fails for the Qwen 3.5-3.8
    full-attention layers with an fp8 KV cache.

macOS (pre-release)

  • Laya runs natively on Metal, with the native Reads page.
  • Qwen 3.8 Flash-Next runs faster on Metal and reuses cached prefixes across
    agent turns.
  • Qwen agent-turn resume points are kept on Metal.

Known

  • Laya reads text only.

  • Laya's confidence on choices with more than ten options is not
    calibrated:
    the English checkpoint ships an out-of-range temperature for
    that case, which Paddock clamps, as Laya's own server does.

  • Switching Vision off does not unload the image tower when its file sits
    beside the model: the model still answers images, and the memory estimate
    does not count the tower (about 0.9 GB).

  • On-demand loading covers Whisper only.

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).


Paddock 0.1.11, built from f99cacdf on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

07b866f3ad31aa916ef2e2604c5bf2f7fe046c73ac7473b99def550bb66e55bc  paddock-0.1.11-x86_64-windows.zip
82411193a0f037d8f583a619c6b4edf6941660f027364117b96894714cb6d1d5  paddock-0.1.11-x86_64-linux.tar.gz
8b7ee3aaa2bc1d9d7273d0136544871da68b8042c6831046b5a9fd83057dcb59  paddock-0.1.11-aarch64-linux.tar.gz

Paddock 0.1.10

Choose a tag to compare

@panterlo panterlo released this 25 Sep 18:35

A feature release: Laya, an open model built only to answer questions about
a text, a much richer Reads page, image input for DiffusionGemma's reads, and
faster Qwen 3.8 Flash-Next. Windows x64, Linux x64 and the NVIDIA DGX Spark.
NVIDIA GPUs, driver 580 or newer. The macOS pre-release for Apple Silicon is
built from the same commit.

New

  • Laya, a decision model. Laya (ConvAI Innovations, Apache 2.0) answers
    yes/no, multiple-choice and scored questions about a text with a calibrated
    probability each, in one pass per question and without writing any text -
    a few milliseconds a question. It serves the same POST /v1/systemone as
    DiffusionGemma, so the Reads page works with it unchanged. The download
    (2.4 GB) holds an English model, a multilingual one for 100+ languages and
    a typed-decisions one; Paddock picks between the first two from the text
    itself, as Laya's own server does. A text longer than one question's window
    is read in overlapping windows instead of being cut.

  • Images in reads. DiffusionGemma reads a question set over pictures as
    well as text, with its own vision companion (1.2 GB) built from Google's
    checkpoint. Attach pictures on the Reads page by button, drop or paste.

  • Conditional questions. A question can be asked only when an earlier
    answer is one of the values you pick, or read with an earlier question's
    answer already known. A question that is skipped answers null and says
    why.

  • More ways to read. Several denoising steps with the answer template
    held in place, a thought written before the answers, and repeated reads
    that report how much the answer moves between them.

Improved

  • The Reads page keeps its history in a side panel beside the questions,
    stored by the manager like conversations, and stays in the sidebar with no
    model running so earlier reads open any time.

  • DiffusionGemma's compact build starts on NVIDIA GPUs.

  • Qwen 3.8 Flash-Next decodes faster, with and without speculative
    decoding, and requests that carry tools now speculate too.

  • Paddock's own app icon in the Studio, the Windows executables and the
    menu bar.

Fixed

  • Flash-Next greedy decoding with speculation no longer drifts from the
    answer it gives without speculation on a small share of requests.

  • The Studio over the network opens at the key prompt again instead of
    reporting that it could not open its saved data.

macOS (pre-release)

  • DiffusionGemma reads images on Metal.
  • The native Reads page takes pictures by paste and import, keeps its history
    in the shared database the web Studio uses, and navigates it from the
    sidebar.
  • Fixes to notification previews and benchmark forms.

Known

  • Laya runs on NVIDIA GPUs only for now, and reads text only.

  • Laya's confidence on choices with more than ten options is not
    calibrated:
    the English checkpoint ships an out-of-range temperature for
    that case, which Paddock clamps, as Laya's own server does.

  • Switching Vision off does not unload the image tower when its file sits
    beside the model: the model still answers images, and the memory estimate
    does not count the tower (about 0.9 GB).

  • On-demand loading covers Whisper only.

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).


Paddock 0.1.10, built from 610cd8f5 on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

3a0949a92bab20c15f37c8a5c29b1f8d80ecbeda23111f90ac4c4fa65208f97d  paddock-0.1.10-x86_64-windows.zip
ca0cdce0c59520eb3a4f9e5930144222d9ad8a630f9909676d2bfb83ba90873b  paddock-0.1.10-x86_64-linux.tar.gz
9f50baf2c14b22df39cd970792f02218eb92679f4c1dea98815d2168c3a2ab99  paddock-0.1.10-aarch64-linux.tar.gz

Paddock 0.1.9

Choose a tag to compare

@panterlo panterlo released this 24 Sep 18:29

A feature release: a diffusion language model with a structured-read API,
image generation and editing, a small MiniCPM5 model, Claude Code working
again with its current versions, and long-context fixes for Qwen 3.8
Flash-Next. Windows x64, Linux x64 and the NVIDIA DGX Spark. NVIDIA GPUs,
driver 580 or newer. The macOS pre-release for Apple Silicon includes the corrected Paddock
branding; its exact source is recorded below.

New

  • DiffusionGemma 26B A4B. Google's diffusion version of Gemma 4 26B
    writes a whole block of text per step instead of one token at a time.
    27 GB of weights at full quality, text only for now.

  • Structured reads. POST /v1/systemone asks DiffusionGemma a fixed set
    of yes/no, multiple-choice and scored questions about a text and returns
    every answer with its confidence, in the shape of the Jev API. The
    Studio's new Reads page builds the questions, runs them over a pasted or
    attached text and saves question sets for reuse.

  • Qwen-Image 2.1: pictures from a prompt, and edits of your own
    pictures.
    POST /v1/images/generations and POST /v1/images/edits, with
    previews streamed while the picture renders. In the Studio a picture is a
    conversation of its own: ask for one, then ask for changes. Full quality
    needs about 17 GB of VRAM and the compact build about 10 GB, plus room
    that grows with the picture. The Qwen Research License allows research and
    evaluation only, not commercial use.

  • MiniCPM5 2B, a small model with tool calling.

  • Whisper endpoints can load their model on demand. An endpoint set to
    load on first request stays up without holding GPU memory, loads the model
    when a transcription arrives and unloads it after the idle time you
    choose. Set it under Model loading in the endpoint's settings; the change
    applies without a restart. Other models stay loaded as before.

Improved

  • Qwen 3.8 Flash-Next at long context. Past about 2,000 tokens the model
    attends sparsely, and Paddock now serves it that way. Before, longer
    prompts got full attention, which is not the model's own behaviour. The
    full 262K context now loads, a long prompt no longer freezes the other
    conversations while it is read, and a continued conversation resumes from
    its cache at any depth.

  • Long prompts no longer stall other conversations on any model: while
    one conversation reads a long prompt, the others keep getting tokens at a
    steady pace.

  • Agent loops reuse more of their cache. The next turn of an agent
    resumes at the tool call instead of reading the whole previous reply
    again, and the model's earlier reasoning now reaches its template on every
    model, as the model expects.

  • Replies run to the context window when no limit is set. A request
    without max_tokens used to stop at 1,024 tokens, which cut reasoning
    models off before they answered.

  • Streams stay alive during long work. A long prompt or a long tool call
    no longer leaves a stream silent long enough for a client to hang up.

  • Gemma 4 serves 4-bit k-quant GGUF files such as Q4_K_M at 4 bits.

Fixed

  • Claude Code works again. Current Claude Code versions send system
    messages in the middle of a conversation, and every request failed with
    "invalid message role". They are accepted now.

  • Claude Code at long context. Token usage counts the cached prefix
    once, so Claude Code no longer compacts at half the window. A prompt past
    the window returns Anthropic's own "prompt is too long" error, so Claude
    Code compacts and carries on. Requests with tools keep speculating after
    the first turn.

macOS (pre-release)

Apple Silicon, macOS 26 or later. The native Swift app is pre-release;
visual glitches and incomplete features remain. The CLI's shared web Studio
is currently more complete.

  • DMG: drag Paddock to Applications for the native app with its bundled Metal runner.
  • CLI PKG: installs paddock and paddock-runner in /usr/local/bin.
    Run paddock and open http://localhost:11500 for the web Studio.
  • CLI tar.gz: the same command-line tools for manual installation.

The app and command-line tools are Developer ID signed. The app, DMG and
PKG are notarized; the app, DMG and PKG carry stapled tickets. Each download
has a separate .sha256 file.

macOS build 1902 comes from public commit
d4b45480:
v0.1.9 plus the corrected app/menu-bar/Studio branding and a sidebar layout
test correction. The shared v0.1.9 tag and existing Windows/Linux assets
are unchanged. Validated on an M5 Mac running macOS 26.6.1; M1–M4 and a clean
macOS 26.0 installation were not requalified for this build.

  • Image generation and editing with Qwen-Image on Metal.
  • DiffusionGemma on Metal, with the Reads page in the native app.
  • MiniCPM5 on Metal, and faster Bonsai.
  • Usage, activity and cache views, storage per model, saved settings
    profiles and on-demand Whisper loading in Settings.
  • Fixes to documents, transcripts, viewers and scrolling.

Known

  • DiffusionGemma on Metal: structured-read labels pass the release smoke
    tests, but confidence values are not bit-identical across batching or model
    reloads. These checks do not establish full numerical parity.

  • Switching Vision off does not unload the image tower when its file sits
    beside the model: the model still answers images, and the memory estimate
    does not count the tower (about 0.9 GB).

  • On-demand loading covers Whisper only.

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).


Paddock 0.1.9, built from fd34f33a on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

55598602d6c6ea6991de1d7a2faf3947d906fb60638aedd590ce5bbf47b58f4f  paddock-0.1.9-x86_64-windows.zip
96040231d2ebb8757f3b1a32233bc5868067017ef8624a1149f624680b33f151  paddock-0.1.9-x86_64-linux.tar.gz
e784a9c89778155d15ec9b77a888b160f9da72dcade510ba99223415b754d7ce  paddock-0.1.9-aarch64-linux.tar.gz

Paddock 0.1.8

Choose a tag to compare

@panterlo panterlo released this 21 Sep 17:51

A feature release: a new model that puts 27B on a 16 GB card, a fix for image
requests under load, and a smoother Studio. Windows x64, Linux x64 and the
NVIDIA DGX Spark. NVIDIA GPUs, driver 580 or newer.

New

  • Bonsai 2 27B: a 27B model on a 16 GB card. Prism ML's ternary build of
    Qwen 3.8 27B is 5.9 GB of weights and is served exactly as shipped. Image
    input is a separate 0.9 GB download, on by default. A 16 GB card holds about
    43,000 tokens of context with text only and about 27,000 with images; a
    24 GB card reaches the full 256K with the 8-bit KV cache. Ternary weights
    are close to the original model, not identical to it.

  • Qwen 3.8 Flash-Next speculates for sampled requests. The drafter used to
    help only requests that did not sample. Requests with a temperature, which
    is the default, now speculate too.

  • Qwen 3.8 Flash-Next has an NVFP4 build for Blackwell cards. The catalog
    offers the distilled NVFP4 checkpoint for RTX 50-series class GPUs, and the
    engine reads MX-quantized checkpoints as they ship.

  • An image segmentation endpoint. POST /v1/segmentations takes image
    chips and returns rasters from a DINOv3-based segmentation checkpoint served
    by path. No catalog model uses it yet.

Improved

  • Long chats stay smooth in the Studio. Formatting a long conversation now
    runs in the background and is spread across frames, so scrolling and
    streaming no longer stall on large replies, tables and diagrams.

  • Cloud models get reply lengths sized from the provider's own token
    counts
    , not from an estimate.

  • Conversation summaries hold up under pressure. Four fixes contributed by
    @DivyamTalwar: a window with no room left is no longer sent the whole
    transcript, a model switch mid-summary no longer saves the old model's
    summary under the new one, a stalled summary request times out, and an
    injected summary counts against the reply budget.

Fixed

  • Several image requests at once could all answer with garbage. On the
    Qwen 3.5, 3.6 and 3.8 vision models, two or more image requests arriving
    together, each with a prompt of 128 tokens or more, could all return a run
    of "!" in place of an answer. A single request at a time was never affected.

  • Qwen 3.8 Flash-Next tool calls with the IQ3 build. Tool calls came back
    as plain text, tool_choice: "required" was refused and thinking budgets
    were rejected. The engine now reads the tool-call format from the model's
    own template when it has no entry of its own.

Known

  • Switching Vision off does not unload the image tower when its file sits
    beside the model: the model still answers images, and the memory estimate
    does not count the tower (about 0.9 GB).

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).


Paddock 0.1.8, built from 1b321df5 on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: linked in (sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a), nvcc 13.3.73.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

b95fdc25cb83f237543592dc3ae99aba062c06e59863990dc7e96cd6a267db18  paddock-0.1.8-x86_64-windows.zip
1db2ec41a280ca8dc57f82749bed41521e70d02e126075d7744fc99877ddaddc  paddock-0.1.8-x86_64-linux.tar.gz
6a3689efb185fe0c06f7944d018266cdf930cea02c1c9cfad78c34122fdd2624  paddock-0.1.8-aarch64-linux.tar.gz

Paddock 0.1.7

Choose a tag to compare

@panterlo panterlo released this 16 Sep 17:15

Paddock 0.1.7, built from 98e6703f on the maintainers' machines (Windows natively, Linux x64 in a Docker builder, the DGX Spark natively on a Spark), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers - x64: RTX 30/40/50 series, the RTX A / PRO line, B200; aarch64: the DGX Spark (GB10). Other cards get an honest refusal, never a slow fallback. This build: linked in (sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a), nvcc 13.3.73.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64; Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+); or a DGX Spark on DGX OS 7 (Ubuntu 24.04, glibc 2.39).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

f2edd3375864776c1fff1f14e487294fe9ec0da3c293bf4287fb5c84e468c292  paddock-0.1.7-x86_64-windows.zip
73b18b7db01e7420d8117e3f1825d7367809a2058b5c462f1609aece9698bdaa  paddock-0.1.7-x86_64-linux.tar.gz
92b22246c9e64bea9d18658348f88ceaab818e3906aaf4cf14088563795e7b0e  paddock-0.1.7-aarch64-linux.tar.gz

Paddock 0.1.6

Choose a tag to compare

@panterlo panterlo released this 13 Sep 10:10

A feature release. Windows x64, Linux x64 and, new in this release, the NVIDIA
DGX Spark. NVIDIA GPUs, driver 580 or newer.

New

  • The DGX Spark is supported. The GB10 in NVIDIA's DGX Spark joins the
    supported GPUs, with its own Linux build for the Spark (DGX OS 7 / Ubuntu
    24.04 or newer). Its kernels were tuned on the Spark and the model catalog
    was tested on it.

  • Gemma 4 reads small print. A new endpoint setting, max_image_tokens,
    sets how much detail an image carries. Gemma 4's default is too coarse for
    fine print on a full page; raising it reads the figures correctly. More
    detail costs prompt tokens and time per image, so the default is unchanged.

  • Conversations resume from the last reply. Qwen 3.5, 3.6 and 3.8,
    Nemotron 3.5 Lightning and Qwen 3.8 Flash-Next pick the next turn up from the
    end of the previous reply instead of reading the conversation again, and
    many conversations running at once each keep their place.

  • Qwen 3.8 Flash-Next caches repeated prompts.

Improved

  • Gemma 4 keeps to its memory budget. With a tight vram_budget it now
    makes room for the conversation instead of going past the budget.

  • Nemotron 3.5 Lightning and 4-bit NVFP4 models are faster on the DGX
    Spark
    , both reading prompts and generating text.

  • The Studio opens quickly even with a very large chat history, and a
    stopped model keeps its logo and name.

Fixed

  • Opening a Studio chat before it had finished loading could save an empty
    copy over its history.

Known

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5).

Paddock 0.1.5

Choose a tag to compare

@panterlo panterlo released this 07 Sep 08:27

A feature release. Windows x64 and Linux x64, NVIDIA GPUs, driver 580 or newer.

The theme is small cards: 1-, 2- and 3-bit files now serve, and the largest
mixture-of-experts models reach a 16 GB card through expert offload.

New

  • 1-, 2- and 3-bit GGUF files serve. The ggml i-quant family (IQ1_S,
    IQ1_M, IQ2_XXS, IQ2_XS, IQ2_S, IQ3_XXS, IQ3_S, IQ4_NL) and the Q2_K and
    Q3_K formats load and serve end to end, for dense weights and for MoE
    experts alike, with decoders written here from the format specification.
    Prompt processing on these files runs on the same tiled GEMM as the 4-bit
    files, so a long prompt costs the same as it does at 4 bits. A file is
    named by its real bit width everywhere it appears. Contributed by
    NodeNestor (#17, #21).

  • Qwen 3.8 Flash-Next serves from GGUF on RTX 50-series cards. The
    mixture-of-experts lane loads unsloth's UD files, and with expert offload
    the UD-IQ1_S file serves on a 16 GB card at llama.cpp parity. Prompt
    processing streams the offloaded experts through the cache expert by
    expert: a 427-token prompt went from 308 s to 37 s, and the first token of
    a 1,500-word prompt from 31 s to 15 s, on an RTX 5060 Ti. Blackwell only in
    this release; the family's kernels are built for sm_100 and sm_120.
    Contributed by NodeNestor (#18).

  • Qwen 3.8 27B has a Compact lane in the catalog: the UD-Q3_K_XL file,
    for cards where the standard file does not fit.

Improved

  • Shared-prefix serving under load. Requests that arrive with the same
    prefix, the shape of an agent fleet, are now recognised at admission and
    warmed in one batched wave instead of each recomputing the prefix. With
    eight concurrent sessions on a shared prompt the time to first token fell
    from 1184 ms to 248 ms and the p99 gap between tokens from 88 ms to
    31 ms.

  • Small cards decode faster with many streams. The k-quant K-split path
    now takes batched decode, so on an RTX 5060 Ti eight concurrent streams of
    a Qwen 3.5/3.6 k-quant file step in 17.8 ms instead of 30.0 ms, twice
    llama.cpp's aggregate on the same card.

  • Memory planned by demand. Prompt-processing scratch is profiled at
    load, and state checkpoints and cache retention are sized by what the
    configuration actually needs, across Qwen 3.5/3.6, Gemma 4, Granite,
    Laguna and Nemotron. Laguna serves batched. A graph_scratch_mib setting
    in the Advanced tab reserves extra scratch by hand (ErikBPF, #14).

  • A GPU the engine has not validated serves under a startup warning
    instead of a refusal.

Fixed

  • Blackwell kernels on Windows could fail: the Windows build laid out
    tensor-map kernel parameters at 8-byte alignment where the hardware wants
    128 (#6).
  • The multi-column weight kernel faulted on narrow planes at two concurrent
    streams on sm_120 (#19, reported by L4GN).
  • A dense i-quant kernel's launch grid overflowed past 134 million outputs.
  • Qwen 3.5/3.6: a failed attempt to enable batching is released before the
    width ladder retries (ErikBPF, #15).

Known

  • The fp8 KV cache's paged attention on RTX 50-series cards shows a small
    numeric deviation in one split configuration (#5). Being fixed on an fp8
    card.

Paddock 0.1.4

Choose a tag to compare

@panterlo panterlo released this 05 Sep 19:12

Paddock 0.1.4, built from ac58eded on the maintainers' machines (Windows natively, Linux in a Docker builder), each archive booted on a GPU and asked a question before it was uploaded.

Requirements

  • An NVIDIA card the kernel pack covers: linked in (sm_86 sm_89 sm_100 sm_100a sm_120 sm_120a), nvcc 13.3.73 - RTX 30/40/50 series, the RTX A / PRO line, B200. Other cards get an honest refusal, never a slow fallback.
  • A driver from the r580 branch or newer. No CUDA toolkit: the kernels and the runtime are inside paddock-runner.
  • Windows 11 x64, or Linux x64 with glibc 2.35 or newer (Ubuntu 22.04+, Debian 12+).

Run

  • Unzip or untar anywhere. paddock starts the manager and the Studio on http://localhost:11500. paddock-runner --model <file.gguf> serves a bare OpenAI/Anthropic-compatible endpoint. README.txt inside says the rest.

sha256

40f7df7f17cf4b1f00cc51a6261ceda75c718d470727a84a9e9958178062221e  paddock-0.1.4-x86_64-windows.zip
d2a62ddf032b0bc8c97c324c8c2f1f471ba323470dea44db0a02befac230a19c  paddock-0.1.4-x86_64-linux.tar.gz