Releases: Eliovp-BV/paiton-vllm-plugin
Release list
Paiton v0.3.4 — native model serving
Run supported language models through your existing vLLM environment with paiton serve NAME. This release provides a lightweight Python wheel and seven native runtime bundles; the CLI downloads and verifies the matching bundle automatically.
python -m pip install https://github.com/Eliovp-BV/paiton-vllm-plugin/releases/download/v0.3.4/paiton_vllm_plugin-0.3.4-py3-none-any.whl
paiton serve minicpm5Activate the model's supported vLLM environment first. Paiton preserves installed vLLM, Torch and ROCm versions and reuses existing checkpoint files. The qualified hardware is one Radeon AI PRO R9700, 32 GB, on Linux.
- Presets:
minicpm5,qwen38-qronos,qwen38-nvfp4,qwen38-neo,ornith,gpt-oss-20b,qwen3-coder. - Setup, supported environments and model guides.
- Offline preparation, local checkpoints and lockfiles.
- Existing model launchers, console commands and container workflows remain supported. FLUX, MiniMax H3, Wan and FastWan keep their image/video launchers.
The native Qwen NVFP4 and Ornith presets are non-speculative; published DFlash/DFlash2 performance figures belong to their separate launch profiles. Quantization and input support are documented per model.
SHA256SUMS covers the wheel and runtime archives. Model weights are downloaded separately from their pinned publishers. Compiler source and private diagnostics are excluded; third-party components retain their applicable notices and licenses.
Qwen3.8 MXFP4 + DFlash2 — native R9700 runtime v1.0.0
Serve the pinned AMD Qwen3.8-27B MXFP4 checkpoint with DFlash2 on one Radeon AI PRO R9700 using Paiton's native runtime on the official vLLM 0.28 ROCm image.
The matched full 188-request comparison against vLLM-Radiance + DFlash2 measures 21.9% higher weighted decode throughput, 57.0% higher C8 throughput, 12.5–17.3% faster prefill, and 2.89× estimated token-cache capacity within the same 5 GiB pool. The model page includes the three-engine graphs, full tables, workload settings and low-concurrency latency limits.
- One-command launch and model page
- Benchmark evidence
- Pinned native runtime companion
- Container:
ghcr.io/eliovp/paiton-vllm-plugin:qwen38-mxfp4-dflash2-rdna4-v1.0.0 - Immutable image:
ghcr.io/eliovp/paiton-vllm-plugin@sha256:9b2dae214076d35de785e073b31294b033a376b16e6bc1ec1fdada4e54d96c59
The image contains the public serving adapter and allowlisted native runtime files. Original target/draft weights download from pinned publisher snapshots and are hash-verified. Compiler and generated implementation source remain private. The installed official vLLM files and benchmarked native binaries are unchanged; the packaged offline launch passed streaming, eight concurrent requests and the configured 8K context boundary.
Native GGUF through vLLM on AMD RDNA4 — NEO v1.1.0
GGUF weights. vLLM serving. Native Paiton execution on AMD RDNA4.
The pinned DavidAU Qwen3.8 NEO CODER MAX mixed Q4_K_M fine-tune now has a qualified native Paiton profile inside vLLM, including text and single-image API requests. This is the actual vLLM model path; llama.cpp is the independent comparison engine.
| Input / output tokens | Paiton + vLLM median | llama.cpp median |
|---|---|---|
| 128 / 128 | 4.925 s | 5.264 s |
| 1,024 / 128 | 5.631 s | 5.933 s |
| 4,096 / 128 | 9.023 s | 9.099 s |
Fixed token counts, greedy sampling, one active sequence, identical weights and matched runtime settings; five measured repetitions after warmup. Some prefill-only cases still favor llama.cpp, and the long-request lead is small. Different engine activation arithmetic is disclosed in the report.
Native GGUF article · Median/p95 results and quality · Launch and deployment requirements
Image: ghcr.io/eliovp/paiton-vllm-plugin:qwen38-neo-coder-max-q4km-rdna4-v1.1.0
Immutable digest: sha256:534287969135f581744ae481b578599468b0bf7ac9a4051b0941500e4c18da4d
./models/Qwen3.8-NEO-CODER-MAX/serve-docker.shQualified for R9700/gfx1201, 8K total context, one active sequence with queued clients, text and one PNG/JPEG. MTP, video and prefix caching are disabled. Original quantized weight values are preserved; activation precision and held-out quality differences are documented. The compiler and native artifacts add no Torch/Triton dependencies; the existing external vLLM stack remains separate.
The image passed empty-cache download/startup, layer and artifact inspection, immutable pull and offline read-only-cache API qualification without compiler or plugin checkouts. The archive contains only allowlisted runtime artifacts and metadata. It is the byte-identical qualified rc4 payload promoted under the stable asset name; internal bundle identity is retained. The signature is a local integrity check, not independent publisher identity attestation. Weights download separately from pinned upstream revisions. Other model tags remain unchanged.
Post-publication verification of the pulled immutable image repeated the 128/128 workload at 4.931 s median / 4.932 s p95, after warmup with five measured requests and no concurrent CPU build. The two-client queued pair measured 9.908 / 9.919 s. These repeats are separate from the matched comparison table above.
Qwen3-Coder 30B for R9700 — local coding chat and API
Run Qwen3-Coder-30B-A3B-Instruct locally for coding on one AMD Radeon AI PRO R9700 (32 GB).
Download paiton-qwen3-coder-r9700-v1.0.0.tar.gz, extract it, and run:
./run.sh --chatThe launcher pulls the prebuilt Paiton container and downloads the pinned 18.1 GB INT4 checkpoint on first use. Model and runtime caches persist. No compiler or local build is required.
Coding client settings:
- OpenAI-compatible base URL:
http://127.0.0.1:8010/v1 - Model:
qwen3-coder - API key, if the client requires a value:
local - Context budget: 4096 tokens including output
- Streaming and automatic function tool calls are enabled.
Container: ghcr.io/eliovp/paiton-vllm-plugin:qwen3-coder-30b-awq-rdna4-v1.0.0
Model guide · Benchmarks and quality limits
Matched fixed-workload measurements, two runs of 16 requests per engine/setting:
| Concurrency | Stock output tokens/s | Paiton output tokens/s | Gain |
|---|---|---|---|
| 1 | 104.42 | 126.68 | 21.3% |
| 2 | 101.61 | 172.82 | 70.1% |
All 128 measured requests completed with matching token counts. Mean request latency decreased at both matched settings. Both engines passed all four executable coding checks and scored 7/8 on the small quality suite. Numerical and generated-text differences remain. These short-workload results do not establish comprehensive quality parity, long-context performance or a direct win against the external Hyperloom report.
Requirements: Linux, Docker, working ROCm GPU devices, one R9700, 16 GB host RAM with swap, and at least 30 GB free disk. The qualified context cap is 4096; use focused coding requests. No LoRA, speculation or higher-concurrency performance is qualified. GPU clocks and power settings are unchanged.
The bundle contains the public runtime source, launch and benchmark scripts, compiled artifact/manifest, reproducible container recipe and license notices. Model weights are downloaded from the pinned public Hugging Face checkpoint into the user's cache. Private compiler source and unpublished assets are excluded. Verify the archive with its .sha256 file.
Hugging Face runtime package: EliovpAI/Qwen3-Coder-30B-A3B-Instruct-AWQ-4bit-Paiton-RDNA4, pinned tag v1.0.0. Includes compiled artifacts, manifests, checksums and the coding launcher. No weights are mirrored; the launcher downloads the pinned checkpoint directly from its original publisher into the local cache.
FLUX.2 klein v1.0.1: ComfyUI, 1.05 s images, 12.9 GiB peak allocation
FLUX.2 klein 4B now has a free Paiton profile for local image generation on the AMD Radeon AI PRO R9700. Start a connected ComfyUI workflow, select Paiton or Stock (Diffusers), enter a prompt and click Run. Downloads, compilation caches, UI settings and outputs persist.
At identical 1024 × 1024, four-step settings, the qualified Paiton pipeline averages 1.054 seconds per image versus 1.258 seconds stock, with 12.9 GiB peak Torch allocation versus 19.3 GiB. That is 16.2% lower latency, 19.4% higher projected throughput and 33.4% less peak allocation. Maximum sampled driver VRAM is 14.6 GiB; allocation is not total GPU memory use. The complete pipeline stays on the GPU without CPU offload.
One-command ComfyUI
git clone --depth 1 --branch paiton-flux2-klein-gfx1201-v1.0.1 https://github.com/Eliovp-BV/paiton-vllm-plugin.git && cd paiton-vllm-plugin && ./models/FLUX.2-klein/launch.shOpen http://127.0.0.1:8188/?paiton=1. Requires Linux, Docker with Compose and an R9700 with working GPU device access. The first image includes loading and compilation. No registry login, model API or paid service is required.
Warm generation timings include text encoding, denoising, VAE decode and PIL image creation. They exclude startup, PNG writing and UI overhead. Separate warm ComfyUI workflow measurements average 1.406 seconds including image handling and saving. Two measured runs follow two warmups for each of three fixed prompts. Sub-second generation and a 20% latency reduction are not claimed.
- Complete model guide, requirements and quality limits
- Benchmark protocol and raw results
- Compiled artifacts on Hugging Face
- Published container digests
Paiton also covers CDNA accelerators for larger inference workloads. Learn more about Paiton.
Source weights and compilation caches persist across container removal with named volumes or a host directory. Nine runtime/cache tests pass, including real Docker persistence checks. Existing prepared tensors are retained when missing source-cache files need restoring.
FLUX.2 klein on R9700: ComfyUI, 1.05 s images, 12.9 GiB peak allocation
Superseded by v1.0.1, which fixes source and compiler cache persistence across container removal. Use that release for the one-command setup.
FLUX.2 klein 4B now has a free Paiton profile for local image generation on the AMD Radeon AI PRO R9700. Start a connected ComfyUI workflow, select Paiton or Stock (Diffusers), enter a prompt and click Run. Downloads, compilation caches, UI settings and outputs persist.
At identical 1024 × 1024, four-step settings, the qualified Paiton pipeline averages 1.054 seconds per image versus 1.258 seconds stock, with 12.9 GiB peak Torch allocation versus 19.3 GiB. That is 16.2% lower latency, 19.4% higher projected throughput and 33.4% less peak allocation. Maximum sampled driver VRAM is 14.6 GiB; allocation is not total GPU memory use. The complete pipeline stays on the GPU without CPU offload.
One-command ComfyUI
git clone --depth 1 --branch paiton-flux2-klein-gfx1201-v1.0.0 https://github.com/Eliovp-BV/paiton-vllm-plugin.git && cd paiton-vllm-plugin && ./models/FLUX.2-klein/launch.shOpen http://127.0.0.1:8188/?paiton=1. Requires Linux, Docker with Compose and an R9700 with working GPU device access. The first image includes loading and compilation. No registry login, model API or paid service is required.
Warm generation timings include text encoding, denoising, VAE decode and PIL image creation. They exclude startup, PNG writing and UI overhead. Separate warm ComfyUI workflow measurements average 1.406 seconds including image handling and saving. Two measured runs follow two warmups for each of three fixed prompts. Sub-second generation and a 20% latency reduction are not claimed.
- Complete model guide, requirements and quality limits
- Benchmark protocol and raw results
- Compiled artifacts on Hugging Face
- Published container digests
Paiton also covers CDNA accelerators for larger inference workloads. Learn more about Paiton.
Paiton Qwen3.8 RDNA4 v1.3.0
Public runtime package for AMD Qwen3.8 27B Qronos on the Radeon AI PRO R9700.
Scope:
- ROCm 7.14
- gfx1201, TP1, batch size 1
- 8,192-token context
- O2 full-and-piecewise graph capture
- fitted MLP decode shadows with the source W4 path retained for prefill
Container:
ghcr.io/eliovp/paiton-vllm-plugin:qwen38-qronos-rdna4-v1.3.0
Digest:
sha256:c56baf54aca1ad229829c1de26e8792806608e65ee9d06ee210b79cd49f70bc9
The original model weights are downloaded directly from the public AMD checkpoint on Hugging Face.