Skip to content

v0.0.2-alpha1-ci-test

Choose a tag to compare

@mudler mudler released this 09 Aug 08:58
· 4160 commits to main since this release
bd20da3

What's Changed

  • row/SERVE-INTAKE-CADENCE: intake lever RESOLVED NEGATIVE (attribution boundary) + VT_LOOP_TRACE probe by @localai-bot in #34
  • row/SERVE-ASYNC-EXECUTOR: decode-graph slot double-buffer, gated OFF (Option-A groundwork) by @localai-bot in #36
  • row/MODEL-KIMI-LINEAR-GPU: device compute GPU-VERIFIED 12/12·614 on GB10; e2e disk-blocked by @localai-bot in #37
  • perf(qwen3_5): SERVE-ASYNC-OPTION-A — decode-graph input H2D staged out of capture (faithful vLLM, gated OFF) by @localai-bot in #39
  • row/MODEL-KIMI-LINEAR-E2E: §8 oracle golden CAPTURED (STRICT) — e2e blocked on bf16-resident loader by @localai-bot in #40
  • row/MODEL-KIMI-LINEAR-BF16: bf16-resident loader/forward — full 48.9B runs e2e on one GB10, 106/128 near-tie by @localai-bot in #42
  • docs(benchmarks,features): DFlash rows catch up with D14 (drift repair) by @localai-bot in #43
  • fix(cuda-arch): arch-gate the WMMA selectors — pre-Ampere guards were a live trap by @localai-bot in #56
  • Session onboarding: ask what the work is, derive the role — and the role mechanism was broken by @localai-bot in #73
  • The orchestration harness: a reviewer that mutates, and a gate command that can fail by @localai-bot in #76
  • docs(roadmap,status): pre-Ampere rows catch up with W1b/W1c (drift repair) by @localai-bot in #78
  • Vulkan: a model runs end to end, token-exact — 16 native kernels, variant pipeline, two CI gates by @localai-bot in #80
  • protocol(doc-gate): preflight checks the committed RANGE — --staged was vacuous by @localai-bot in #91
  • feat(videos): /v1/videos speaks OpenAI's Sora wire shape, and serves the MP4 by @localai-bot in #71
  • refactor(platforms): gate the platform-SELECTION walk, the one non-additive site (#41) by @localai-bot in #87
  • feat(backend): BACKEND-ROCM W0 — the AMD skeleton, landing UNBUILT and saying so (#41) by @localai-bot in #88
  • feat(videos): reference conditioning over /v1/videos, image and clip and voice by @localai-bot in #92
  • fix(minimax-h3): half of every video was discarded, plus the missing user docs by @localai-bot in #68
  • VK-C: cooperative-matrix GEMM tactic, verified executing on GB10 by @localai-bot in #97
  • VK-C bench: coopmat GEMM is 11.1x-32.9x over our scalar kernel on Thor by @localai-bot in #101
  • feat(minimax-h3): the ORIGINAL bf16 DiT release is indexable — 13 shards, 66.3 GB by @localai-bot in #98
  • VK-E: llama.cpp-Vulkan denominator measured; our arm blocked on a shared model by @localai-bot in #106
  • feat(cpu): wide-x86 elementwise GEMM tiers, AVX2 and AVX-512, byte-identical by @localai-bot in #108
  • feat(minimax-h3): stream the ORIGINAL bf16 DiT to the device, one tensor at a time by @localai-bot in #105
  • feat(minimax-h3): load the bf16 TEXT ENCODER (14 shards, 63 GB) and measure what Q4_K_M costs the conditioning by @localai-bot in #100
  • P0: the CPU grouped keep-quant GEMM was f32-ONLY, so CPU MoE decode emitted token-0 garbage since b4f5610 by @localai-bot in #103
  • feat(cpu): M-blocked transpose-free [K,N] elementwise GEMM and an opt-in load-time repack by @localai-bot in #109
  • MODEL-AUDIO-PARAKEET-ENCODER: Parakeet/FastConformer ASR runs natively on vt:: (CPU), real transcript verified by @localai-bot in #89
  • fix(vulkan): the e2e run was hanging, not slow — OOB coopmat load, plus VK-E vs llama.cpp by @localai-bot in #112
  • docs(usage): how to actually prompt MiniMax-H3 for speech and references by @localai-bot in #96
  • docs(protocol): welcome contributors into the agent workflow by @localai-bot in #119
  • perf(vulkan): VK-A2 command-buffer batching, default ON — decode 8.59 → 59.3 tok/s (6.9x) by @localai-bot in #124
  • record(vulkan): last two levers closed by measurement; fix the VK-A2 profiler blindness by @localai-bot in #126
  • fix(qwen3_5): host pointers to device kernels off-CUDA (#125); GPU timestamp profiling by @localai-bot in #130
  • perf(vulkan): descriptor ring was capping batches — 1.51x, decode 87.3 tok/s by @localai-bot in #131
  • perf(vulkan): re-profile — GPU-bound again, and the GEMV unroll reverses to 7/8 by @localai-bot in #133
  • fix(h3): route video device selection through the backend seam by @localai-bot in #135
  • fix(surface): remove shared CUDA literals from device selection by @localai-bot in #139
  • docs(release): spike downloadable server binary matrix by @localai-bot in #129
  • feat(vulkan): six native GDN/conv1d kernels, and the sub-word buffer bug they found by @localai-bot in #145
  • record(vulkan): 27B re-run measures GDN coverage 11->5; corrects the 1.75x shape error by @localai-bot in #151
  • fix(cuda): non-Triton build was red; paged_attn batching refuted; 27B five-way plan by @localai-bot in #142
  • feat(rocm): gfx1201 hipBLAS ops + Gemma-4-26B-A4B MoE (BF16/FP8) by @bakon11 in #140
  • feat(mem): ROAD-V1-MEM M1+M2 - KV-sizing knobs + group-aware bytes_per_block by @localai-bot in #153
  • feat(tp): TP-W1 - rank-layout group table + per-rank TensorParallel on LoadedModel by @localai-bot in #156
  • policy: optimize the agent protocol for strict, executable compliance by @localai-bot in #128
  • fix(vt,hooks): async readback is a Backend capability + unbreak the pre-push sandbox by @localai-bot in #159
  • docs(readme): News carries MXFP4 vLLM parity and the Vulkan e2e milestone by @localai-bot in #160
  • perf(vulkan): ragged M reaches coopmat — 27B prefill 21.5x by @localai-bot in #162
  • docs(readme): restore the landing-page budget after the MANIFESTO link by @localai-bot in #161
  • feat(protocol): add universal agent session entrypoint by @localai-bot in #167
  • Merge #127 (richiejp): perf(gdn) dispatch exact causal-conv chunks by @localai-bot in #165
  • fix(qwen3_5): load quantized lm_head, not just BF16 (#164) by @localai-bot in #169
  • fix(minimax-h3): REF canvas renders COHERENT + drain the scratch pool at denoise→decode by @localai-bot in #157
  • record: track published GHCR container images as ENG-RELEASE-CONTAINERS by @localai-bot in #172
  • perf(vulkan): drain the batch only on a real dependency — decode 2.74 → 2.99 by @localai-bot in #174
  • diag(minimax-h3): dump the AUDIO latent too, and measure it — the arm is NOT white by @localai-bot in #175
  • docs(contributing): squash-merge, and why --merge reds main by @localai-bot in #176
  • perf(gemma4): dual-GPU FP8 resident without host OOM + peer mix by @bakon11 in #154
  • record(minimax-h3): CALIBRATE the audio arm — VAE cleared by round-trip, DiT latent is OUT OF DISTRIBUTION by @localai-bot in #179
  • diag(minimax-h3): PRE-denormalize audio dump — denormalize CLEARED, the DiT arm regresses to the corpus mean by @localai-bot in #181
  • fix(minimax-h3): the metallic voice — the audio latent was denormalized ACROSS THE WRONG AXIS by @localai-bot in #182
  • perf(vulkan): fused attn preamble native -- 16 host round trips/token gone by @localai-bot in #183
  • perf(vulkan): decode GEMV reads four elements per load, and the row lever is dead by @localai-bot in #184
  • perf(vulkan): decode RmsNorm ran on ONE workgroup -- 7.85 -> 1.59 ms/token by @localai-bot in #185
  • perf(vulkan): lm_head column blocking -1.07ms/tok; the 9.3ms floor I briefed does not exist by @localai-bot in #186
  • record: the three Vulkan levers measured TOGETHER -- 4.285 tok/s BINDING by @localai-bot in #187
  • docs(roadmap): add C10 Qwen3.5 throughput levers and C11 Krea 2 (user-directed) by @localai-bot in #180
  • fix(policy): repair the five gates that were blocking CORRECT changes by @localai-bot in #190
  • perf(vulkan): pipelined submission, and llama.cpp's fence spin REJECTED for wrong numbers by @localai-bot in #191
  • arch(ARCH-ONE-SURFACE): the OpenAI server becomes a THIN ABI CLIENT — vllm_server_main, ABI v17 by @localai-bot in #189
  • release: add versioned binary manifest (W5) by @localai-bot in #141

New Contributors

Full Changelog: v0.0.2-alpha...v0.0.2-alpha1-ci-test