Repository navigation
Releases: localai-org/vllm.cpp
Releases · localai-org/vllm.cpp
Release list
v0.0.2
vllm.cpp v0.0.2 binary index
Source: 7020de93652ca920424a10ac5255b34810dd2f24
| Artifact | Channel | Host ABI | Backend | CPU tiers / CUDA SMs | Boundary and limitations |
|---|---|---|---|---|---|
| linux-aarch64-glibc-cpu (sha256) | stable | linux/aarch64/glibc 2.39 | cpu | portable-neon, dotprod, i8mm | static-core; not-applicable |
| linux-aarch64-glibc-cuda-fat (sha256) | preview | linux/aarch64/glibc 2.39 | cuda | 80, 86, 87, 89, 90a, 100a, 103a, 110, 120a, 121a | static-core; external-host-never-bundled; preview: see per-gate evidence in the embedded manifest; external runtime: nvidia-driver |
| linux-x86_64-glibc-cpu (sha256) | stable | linux/x86_64/glibc 2.39 | cpu | portable-sse2, sse2-f16c, avx2-f16c, avx512f | static-core; not-applicable |
| linux-x86_64-glibc-cuda-fat (sha256) | preview | linux/x86_64/glibc 2.39 | cuda | 80, 86, 87, 89, 90a, 100a, 103a, 110, 120a, 121a | static-core; external-host-never-bundled; preview: see per-gate evidence in the embedded manifest; external runtime: nvidia-driver |
| linux-x86_64-glibc-vulkan (sha256) | preview | linux/x86_64/glibc 2.39 | vulkan | platform-native | static-core; external-host-never-bundled; preview: see per-gate evidence in the embedded manifest; external runtime: vulkan-loader, vulkan-icd, vulkan-driver |
| linux-x86_64-musl-cpu-static (sha256) | experimental-preview | linux/x86_64/musl 1.2.5 | cpu | portable-sse2, sse2-f16c, avx2-f16c, avx512f | literal-static; not-applicable; experimental-preview: see per-gate evidence in the embedded manifest; experimental CPU-only literal-static feasibility lane |
| macos-arm64-metal-mlx (sha256) | preview | macos/aarch64/macos 15.7.7 | mlx | platform-native | static-core; external-host-never-bundled; preview: see per-gate evidence in the embedded manifest; external runtime: Accelerate.framework, CoreFoundation.framework, Foundation.framework, Metal.framework, QuartzCore.framework |
| macos-arm64-metal (sha256) | stable | macos/aarch64/macos 15.7.7 | metal | platform-native | static-core; external-host-never-bundled; external runtime: CoreFoundation.framework, Foundation.framework, Metal.framework |
CI handoff artifacts are retained for 7 days. Published GitHub release assets follow the maintainer-deletion-only policy.
v0.0.2-alpha1-ci-test
What's Changed
- row/SERVE-INTAKE-CADENCE: intake lever RESOLVED NEGATIVE (attribution boundary) + VT_LOOP_TRACE probe by @localai-bot in #34
- row/SERVE-ASYNC-EXECUTOR: decode-graph slot double-buffer, gated OFF (Option-A groundwork) by @localai-bot in #36
- row/MODEL-KIMI-LINEAR-GPU: device compute GPU-VERIFIED 12/12·614 on GB10; e2e disk-blocked by @localai-bot in #37
- perf(qwen3_5): SERVE-ASYNC-OPTION-A — decode-graph input H2D staged out of capture (faithful vLLM, gated OFF) by @localai-bot in #39
- row/MODEL-KIMI-LINEAR-E2E: §8 oracle golden CAPTURED (STRICT) — e2e blocked on bf16-resident loader by @localai-bot in #40
- row/MODEL-KIMI-LINEAR-BF16: bf16-resident loader/forward — full 48.9B runs e2e on one GB10, 106/128 near-tie by @localai-bot in #42
- docs(benchmarks,features): DFlash rows catch up with D14 (drift repair) by @localai-bot in #43
- fix(cuda-arch): arch-gate the WMMA selectors — pre-Ampere guards were a live trap by @localai-bot in #56
- Session onboarding: ask what the work is, derive the role — and the role mechanism was broken by @localai-bot in #73
- The orchestration harness: a reviewer that mutates, and a gate command that can fail by @localai-bot in #76
- docs(roadmap,status): pre-Ampere rows catch up with W1b/W1c (drift repair) by @localai-bot in #78
- Vulkan: a model runs end to end, token-exact — 16 native kernels, variant pipeline, two CI gates by @localai-bot in #80
- protocol(doc-gate): preflight checks the committed RANGE — --staged was vacuous by @localai-bot in #91
- feat(videos): /v1/videos speaks OpenAI's Sora wire shape, and serves the MP4 by @localai-bot in #71
- refactor(platforms): gate the platform-SELECTION walk, the one non-additive site (#41) by @localai-bot in #87
- feat(backend): BACKEND-ROCM W0 — the AMD skeleton, landing UNBUILT and saying so (#41) by @localai-bot in #88
- feat(videos): reference conditioning over /v1/videos, image and clip and voice by @localai-bot in #92
- fix(minimax-h3): half of every video was discarded, plus the missing user docs by @localai-bot in #68
- VK-C: cooperative-matrix GEMM tactic, verified executing on GB10 by @localai-bot in #97
- VK-C bench: coopmat GEMM is 11.1x-32.9x over our scalar kernel on Thor by @localai-bot in #101
- feat(minimax-h3): the ORIGINAL bf16 DiT release is indexable — 13 shards, 66.3 GB by @localai-bot in #98
- VK-E: llama.cpp-Vulkan denominator measured; our arm blocked on a shared model by @localai-bot in #106
- feat(cpu): wide-x86 elementwise GEMM tiers, AVX2 and AVX-512, byte-identical by @localai-bot in #108
- feat(minimax-h3): stream the ORIGINAL bf16 DiT to the device, one tensor at a time by @localai-bot in #105
- feat(minimax-h3): load the bf16 TEXT ENCODER (14 shards, 63 GB) and measure what Q4_K_M costs the conditioning by @localai-bot in #100
- P0: the CPU grouped keep-quant GEMM was f32-ONLY, so CPU MoE decode emitted token-0 garbage since b4f5610 by @localai-bot in #103
- feat(cpu): M-blocked transpose-free [K,N] elementwise GEMM and an opt-in load-time repack by @localai-bot in #109
- MODEL-AUDIO-PARAKEET-ENCODER: Parakeet/FastConformer ASR runs natively on vt:: (CPU), real transcript verified by @localai-bot in #89
- fix(vulkan): the e2e run was hanging, not slow — OOB coopmat load, plus VK-E vs llama.cpp by @localai-bot in #112
- docs(usage): how to actually prompt MiniMax-H3 for speech and references by @localai-bot in #96
- docs(protocol): welcome contributors into the agent workflow by @localai-bot in #119
- perf(vulkan): VK-A2 command-buffer batching, default ON — decode 8.59 → 59.3 tok/s (6.9x) by @localai-bot in #124
- record(vulkan): last two levers closed by measurement; fix the VK-A2 profiler blindness by @localai-bot in #126
- fix(qwen3_5): host pointers to device kernels off-CUDA (#125); GPU timestamp profiling by @localai-bot in #130
- perf(vulkan): descriptor ring was capping batches — 1.51x, decode 87.3 tok/s by @localai-bot in #131
- perf(vulkan): re-profile — GPU-bound again, and the GEMV unroll reverses to 7/8 by @localai-bot in #133
- fix(h3): route video device selection through the backend seam by @localai-bot in #135
- fix(surface): remove shared CUDA literals from device selection by @localai-bot in #139
- docs(release): spike downloadable server binary matrix by @localai-bot in #129
- feat(vulkan): six native GDN/conv1d kernels, and the sub-word buffer bug they found by @localai-bot in #145
- record(vulkan): 27B re-run measures GDN coverage 11->5; corrects the 1.75x shape error by @localai-bot in #151
- fix(cuda): non-Triton build was red; paged_attn batching refuted; 27B five-way plan by @localai-bot in #142
- feat(rocm): gfx1201 hipBLAS ops + Gemma-4-26B-A4B MoE (BF16/FP8) by @bakon11 in #140
- feat(mem): ROAD-V1-MEM M1+M2 - KV-sizing knobs + group-aware bytes_per_block by @localai-bot in #153
- feat(tp): TP-W1 - rank-layout group table + per-rank TensorParallel on LoadedModel by @localai-bot in #156
- policy: optimize the agent protocol for strict, executable compliance by @localai-bot in #128
- fix(vt,hooks): async readback is a Backend capability + unbreak the pre-push sandbox by @localai-bot in #159
- docs(readme): News carries MXFP4 vLLM parity and the Vulkan e2e milestone by @localai-bot in #160
- perf(vulkan): ragged M reaches coopmat — 27B prefill 21.5x by @localai-bot in #162
- docs(readme): restore the landing-page budget after the MANIFESTO link by @localai-bot in #161
- feat(protocol): add universal agent session entrypoint by @localai-bot in #167
- Merge #127 (richiejp): perf(gdn) dispatch exact causal-conv chunks by @localai-bot in #165
- fix(qwen3_5): load quantized lm_head, not just BF16 (#164) by @localai-bot in #169
- fix(minimax-h3): REF canvas renders COHERENT + drain the scratch pool at denoise→decode by @localai-bot in #157
- record: track published GHCR container images as ENG-RELEASE-CONTAINERS by @localai-bot in #172
- perf(vulkan): drain the batch only on a real dependency — decode 2.74 → 2.99 by @localai-bot in #174
- diag(minimax-h3): dump the AUDIO latent too, and measure it — the arm is NOT white by @localai-bot in #175
- docs(contributing): squash-merge, and why --merge reds main by @localai-bot in #176
- perf(gemma4): dual-GPU FP8 resident without host OOM + peer mix by @bakon11 in #154
- record(minimax-h3): CALIBRATE the audio arm — VAE cleared by round-trip, DiT latent is OUT OF DISTRIBUTION by @localai-bot in #179
- diag(minimax-h3): PRE-denormalize audio dump — denormalize CLEARED, the DiT arm regresses to the corpus mean by @localai-bot in #181
- fix(minimax-h3): the metallic voice — the audio latent was denormalized ACROSS THE WRONG AXIS by @localai-bot in #182
- perf(vulkan): fused attn preamble native -- 16 host round trips/token gone by @localai-bot in #183
- perf(vulkan): decode GEMV reads four elements per load, and the row lever is dead by @localai-bot in #184
- perf(vulkan): decode RmsNorm ran on ONE workgroup -- 7.85 -> 1.59 ms/token by @localai-bot in #185
- perf(vulkan): lm_head column blocking -1.07ms/tok; the 9.3ms floor I briefed does not exist by @localai-bot in #186
- record: the three Vulkan levers measured TOGETHER -- 4.285 tok/s BINDING by @localai-bot in #187
- docs(roadmap): add C10 Qwen3.5 throughput levers and C11 Krea 2 (user-directed) by @localai-bot in https://githu...
v0.0.2-alpha
What's Changed
- chore: retest and rebench on RTX 5070 by @richiejp in #9
- docs(capi): define stable ABI contract by @localai-org-maint-bot in #10
- fix(build): clear GCC 12 production warnings by @localai-org-maint-bot in #11
- fix(records): repair unreachable DONE owner commits by @localai-org-maint-bot in #12
- docs(cli): spike chat and complete parity by @localai-org-maint-bot in #13
- fix(build): demote MLX header folding warning by @localai-org-maint-bot in #23
- fix(build): suppress MLX folding warning after flags by @localai-org-maint-bot in #24
- fix(build): guard Laguna Marlin-only helper by @localai-org-maint-bot in #25
- fix(build): isolate MLX headers from project warnings by @localai-org-maint-bot in #27
- docs(readme): stop claiming faster than vLLM, the curve is a tie by @localai-bot in #29
- fix(metal): scope MLX AppleClang warning suppression by @localai-org-maint-bot in #30
- row/MODEL-FACTORY-registry: repair registry-surface test for KimiLinearForCausalLM by @localai-bot in #33
- row/SERVE-ASYNC-LLM: async batch-1 P0 root-caused + gated + fixed (device-mirror default ON) by @localai-bot in #31
- row/BENCH-35B-TTFT: TTFT deficit attributed to serving INTAKE (diag instrument, no lever yet) by @localai-bot in #32
New Contributors
- @richiejp made their first contribution in #9
- @localai-org-maint-bot made their first contribution in #10
Full Changelog: v0.0.1-m03-parity-27b-35b...v0.0.2-alpha
v0.0.1-m03-parity-27b-35b
perf(kernel): fold 35B GDN out_proj fp8-quant into gated-RMSNorm — by…