Repository navigation
1Cat-vLLM 1.5.1 发布说明
1.5.1 聚焦 V100 长文 prefill/decode、TP1/TP2 兼容布局、并发、Flash-Next 与显存复用和默认编译缓存,并随 wheel 交付启动配置与加速自检。
1.5.0 → 1.5.1:27B 单请求性能
| 输入长度 | 1.5.0 pure decode tok/s | 1.5.1 pure decode tok/s | 提速 |
|---|---|---|---|
| 4K | 103.87 | 229.34 | +120.80% |
| 32K | 49.04 | 201.18 | +310.27% |
| 128K | 未取得完成结果 | 171.43 | 不计算 |
完整投机轮耗时:4K 40.917 → 19.855 ms,32K 81.253 → 21.125 ms;1.5.1 的 128K 为 25.212 ms。1.5.0 的 128K 首次请求在 180 秒内未完成,停止该项测试,因此没有旧版速度或提速比例。
**测试条件:**使用正式 1.5.0 wheel 与本次 1.5.1 发布 wheel、同一组 4×V100-SXM2-32GB、TP4、QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 和 incoai/Qwen3.8-27B-DFlash2。Python 3.12.3、Torch 2.10.0+cu128、CUDA 12.8.93、驱动 580.173.02;FP16 compute/draft KV、E4M3 target KV、7 个 probabilistic draft tokens。两版采用相同的 1.5.1 发布参数:maxlen 262144、prefill budget 8192、maxseq 4、memory .80、KV block 2048、Mamba block 8192、前缀缓存开启、异步调度、FULL_AND_PIECEWISE CUDA Graphs。采样 T=1、top-p=.95、top-k=20、seed=0,输出上限 256,允许正常 EOS;不是两版各自最佳调参的比较。
每个长度 1 次冷请求 + 3 次暖请求,表中为 3 次暖测的中位数;每版来自 一次成功启动。pure decode = (输出 tokens−1) / 服务端 decode 时间,排除 TTFT/prefill;完整投机轮 = 服务端 decode 时间 / 投机轮数,包含 draft、target、采样与调度。32K 暖请求中,1.5.0 命中 28,672 个前缀 tokens,1.5.1 重算完整提示;两版的实际输入上下文均为 32,768 tokens。这里比较的是发布配置下的 decode,不是缓存命中率或 prefill 提速。
1.5.0 基线是正式 v1.5.0 wheel,SHA256:2a4d6bee4e19d315b142f2c563059f3064ddeeca563a6bdc828c33e1073c825b。
并发 decode 优化
并发是1.5.1的重点:优化多请求的目标验证、草稿 attention、批量权重供给与采样,同时减少工作区重复分配。兼容算子按SM70硬件、dtype、layout和shape自动选择,推荐参数无需额外导出性能开关。
- **FP4/FP8批量GEMM与预热:**共享activation和解码权重供给,覆盖M32/M48/M64及尾部;修复已准备batch layout未进入协同预热的问题。#691、#706、#708。
- **联合草稿attention和GDN:**跨请求合并draft attention launch,复用packed GDN、metadata及状态工作区;保持请求各自的页表、长度与接受状态。#688、#699。
- **长文并发验证与采样:**将已有长文q8 attention扩展到B2–B16,复用Graph工作区和logits,减少采样同步并优化兼容TP4归约。#697、#706、#708。
**纳入本版PR的同条件开发对照:**下面各PR使用自己的前后基线,数值不能累加,也不是正式1.5.0→1.5.1 wheel的并发提升比例。最终wheel的实测另列于下方。
| 优化 PR | 并发 | 对照 → 优化后 pure decode tok/s | PR阶段收益 |
|---|---|---|---|
| #697:长文batch attention | C4 | 326.832 → 429.134 | +31.30% |
| 同上 | C8 | 492.025 → 624.863 | +27.00% |
| #706:batch预热、供给与采样 | C2 | 236.175 → 252.465 | +6.90% |
| 同上 | C4 | 377.689 → 436.004 | +15.44% |
| 同上 | C8 | 612.982 → 709.714 | +15.78% |
条件为27B NVFP4开发checkpoint、4×V100-SXM2-32GB/TP4、FP16 compute/draft KV、E4M3 target KV、DFlash2 q7(target验证q8)、maxlen262144、memory.80、prefix caching和CUDA Graphs。#706使用32K/256、T=.7/p=.8/k20、匹配seed和固定顺序入队,1组预热后3组正式波取中位数,完整输出与接受计数匹配。#697为32K/256同时入队、3次暖测共同decode窗口,C4/C8接受率有变化,不能把全部收益归给单算子;该PR的具体采样与artifact条件见原始报告。两者均排除初始prefill,单启动记录不代表跨机器稳定收益。
最终1.5.1 wheel的27B并发实测:
QUASAR-QAT 发布配置的 32K C4 pure decode 为 445.45 tok/s。与上面的单请求使用相同硬件、模型和默认启动配置;本项使用固定 32,768-token synthetic 输入、每请求 256 强制输出 tokens、T=.7/p=.8/k20、seed=20260923+请求序号,共享前缀预热后按固定顺序入队、同时恢复。一次启动,1 组预热 + 3 组正式测量,取中位数;三次为 445.48 / 435.35 / 445.45 tok/s。
按实际返回 token IDs 计算最后一个请求首 token 到第一个请求末 token 的共同 decode 窗口,排除 TTFT 和新 prefill。测试仅为同步入队额外开启 VLLM_SERVER_DEV_MODE=1,不改变加速策略;强制输出用于性能计量,不作为质量证据。该指标与含 prefill 的整体吞吐不同。
本版亮点
- **长文 prefill/decode:**原生 Q8000/Q8192 attention、E4M3 KV 桥接、Mamba grid解耦和 FP32 敏感状态;兼容大分块和长文/尾部 CUDA Graph 默认选路。
- **TP1/TP2:**按本地头组和投影几何拓宽兼容路径,修复 TP2 GDN 对齐;PR 已验证 target-only TP1 64K、TP2 256K,条件和 DFlash 限制见下文。
- **并发:**FP4/FP8 batch GEMM、联合 draft attention、GDN metadata/copy融合和采样复用,保留原归约顺序与正常 fallback。
- **Flash-Next:**FP16 GEMV/HC/GDN 按算子条件选路;前缀缓存减少不必要 prefill切分,支持 native MTP 的共享 metadata和 draft graphs。
- **显存与容量:**兼容权重/scales/scores/scratch默认共享,AWQ紧凑元数据、并发规模 scratch、相同 MTP head共享;省下的空间可用于 KV池。
- **启动与缓存:**编译缓存默认开启、稳态 KV预算、跳过未用权重读取、减少加载转换/重排峰值、修复图重载与 PLE自动放置。
- **其他工作流:**分组 RAM/FS KV缓存、DCP同步移除、混合设备选路及原生 H3/Z-Image相关计算与 host内存优化;按需配置和验证范围见完整变更表。
- **安装与自检:**两套随包启动脚本、本地
--draft和配置加速自检;无需开发者私有.so或性能环境变量,标准 Toolkit/JIT要求见安装说明。
优化效果与验证范围
以下记录说明 1.5.0 之后纳入的改动各自解决了什么、在其 PR 测试中取得什么收益。它们来自不同的开发阶段和同条件对照,不是最终 wheel 的新增复测,也不是统一的 1.5.0→1.5.1 提升比例;不能将百分比或显存节省相加。最终 wheel 实测另列于下一节。原始 PR 与集成 PR 不重复计功;已在1.5.0的旧优化、未纳入产物的 PR 和未交付私有候选不计入新收益。
长文 prefill、算子与缓存调度
原生 sm70_d256_gqa_architecture_fwd / sm70_d256_gqa_architecture_q8192_fwd 把兼容长文 QK/PV 放到 Tensor Core;E4M3 paged KV 通过 sm70_v37_e4m3_bridge 进入共享 FP16 workspace,保留 QK/PV FP32 累加。支持 Q8000/Q8192、多个请求与 TP 本地 GQA6/D256 头组。#548、#628、#635、#638、#666。桥接没有独立匹配耗时对照,不能单独计算其加速倍数。
保留的完整 FP32 算子测试:单张 V100-SXM2-32GB、Torch2.10+cu128/CUDA12.8,Q8192/KV262144,CUDA Graph replay并包含各头组复制:
| 本地布局 | attention耗时 | 有效算子吞吐 |
|---|---|---|
| TP4,6 query heads / 1 KV head | 182.645 ms | 71.111 TFLOP/s |
| TP2,12 query heads / 2 KV heads | 369.916 ms | 70.221 TFLOP/s |
| TP1,24 query heads / 4 KV heads | 742.363 ms | 69.982 TFLOP/s |
这是 #666 阶段的单 GPU 本地布局测试,非多 GPU 整模型吞吐或最终 wheel 复测。早期75–77T记录尚未包含完整 QK FP32 修正,不能套用到当前精度合同;功耗/频率也影响数值,同PR的185W记录为53.816T。因此不承诺用户设备固定达到71T。
| 改动及 PR | 同条件结果 | 测试条件与含义 |
|---|---|---|
| Mamba grid解耦 #645 | 100K prefill 2954→3336 tok/s,+12.93% | QUASAR 27B NVFP4,4×V100 TP4,E4M3/DFlash7,budget8192/maxseq4/prefix ON,unique-salt无命中100K提示;固定KV block2048,grid4096→8192,实际chunk进入快速attention。为PR探测值;相比原KV4096部署保持其prefill速度带、KV容量增加约23–29%。随包脚本已采用KV2048/Mamba8192。 |
| Flash-Next prefix prefill #754 | 8K 3060→5356;32K 3046→5068;128K 2818→4666 tok/s,分别 +75.03% / +66.38% / +65.58% | Flash-Next NVFP4,4×V100 TP4/MTP4,FP16 compute/KV,maxlen131072/budget8192/seq1/memory.90/sync/Graphs,原测试显式PLE host12GiB/rank;每arm一次启动。8K取第二次无命中样本,32K/128K取两次无命中合计;T0/seed0/max16/正常EOS。减少跨丢弃状态边界的forward次数,保留相同缓存复用;不是新单算子倍数。 |
| AWQ indexed-A专家链 #464 | W13→SiLU→W2 6.64→4.97 ms,耗时−25.15%;避免 400 MiB展开输入 | Flash-Next AWQ native-g32,V100 TP4,FP16 activation,8192 tokens/top-k10/hidden2560,rank0/3同权重与metadata,输出bitwise相同。整模型同服务C1、16K–128K prefill几何平均 +5.83%,每cell1预热+3计量;不可移作NVFP4 27B收益。 |
| 通用加速组合 #666 | 262128输入冷TTFT 174.343→132.334秒,−24.10%;输入/TTFT +31.74% | QUASAR 27B NVFP4,4×V100 TP4,E4M3目标KV/FP16 draft KV,DFlash7,maxlen262144/budget8192/maxseq32/memory.85,Graphs、无前缀命中;自然EOS检索均正确。为PR所记的一组冷请求,包含首token及多项路径/修复,不能全部归给attention或当pure-prefill算子测速。 |
GDN既有路线恢复和早期缺失扩展诊断没有被计作新增生产内核收益;未交付的私有 GDN候选也未计入。
TP1/TP2 与显存容量
#666 按本地头组和投影几何选路,兼容算子不再统一要求TP4;#517 修复TP2 GDN物理宽度8240→8256。27B NVFP4 target-only 的默认路径在PR中通过 TP1 64K、TP2 256K 冷检索及C1–C32客户端负载测试:FP16/E4M3、chunk8192/maxseq32、Graphs;TP1单张32GB/maxlen65536/memory.92,TP2两张32GB/maxlen262144/memory.85,DFlash关闭。TP1最多9个驻留请求、较高客户端并发排队;TP2测试达到32个驻留请求。它们是能力验证,不是最终wheel的旧版缩放比。
TP2 DFlash的 #566 记录完整B1/q8轮 31.885 / 29.280 ms(release1k / MBPP28):两张V100、CUDA12.8/Torch2.10、E4M3/FP32、T1/p.95/k20/正常EOS,3次独立启动、每fixture预热后5对交替测量。该完整组合包含显式开启项和专用测试条件,不是零配置最终wheel速度承诺。TP1 27B+DFlash2的大分块容量限制仍保留。
下表单位为每卡;原验证主要使用Torch2.10+cu128/CUDA12.8,模型与模式分别独立。自动KV分配可使用省下空间,因此总NVML占用不必按相同比例下降。
| 优化 | PR同条件证据 | 范围与开启条件 |
|---|---|---|
| 权重/codes/scales共享 #561、#671 | 模型加载 10.07→6.47 GiB;可用KV预算 10.57→12.68 GiB | #671,QUASAR NVFP4 TP4/DFlash7/E4M3目标与FP16 draft KV,V10032GB/185W,maxlen262144/chunk8192/seq32/memory.85/Graphs。最终兼容共享权重/scales默认开启,旧PR的opt-in状态已被后续默认推广覆盖。 |
| score/graph scratch/draft snapshots #673、#677 | 非KV活跃显存:TP2 17.117→15.146 GiB;TP4 10.225→8.794 GiB | #677,QUASAR NVFP4/DFlash7/E4M3与FP16 draft KV,V10032GB,maxlen262144/chunk8192/seq32/memory.90/Graphs,default ON/OFF对照。当前score block8192;部分prefill延迟有约1–2%取舍,节省量不与上一行累加。 |
| AWQ紧凑metadata #473、#522 | resident metadata减少约 0.879 GiB | Flash-Next AWQ TP4/E512/native-g32,兼容3-byte layout默认;其他形状fallback。算子测试保持解码权重/输出一致。 |
| AWQ resident scratch #465 | 加载 21.24→21.11 GiB,KV容量 771736→789443 | V100 TP4、MTP0、seq8;按并发和MTP verifier宽度自动分配,保持已有上限。该容量不是任意seq/模型的保证。 |
| 多模态目标/草稿共享head #594 | 请求后idle显存 30712→30412 MiB | Flash-Next AWQ-g32,4×V100 TP4,FP16 head、MTP3/C2/Graphs,固定5GiB KV,4条确定性请求;相同head自动共享。不是加载峰值节省。 |
| FP8-resident MTP专家 #595 | 原生FP8 checkpoint对照:AWQ 1010 MiB、NVFP4 892 MiB节省 | 原开发环境的TP4/MTP3/C2、FP16 compute/KV、固定5GiB KV/Graphs,greedy/stochastic请求后idle读数;checkpoint原始FP8支持与显式online转换分开。不是最终MTP4 profile复测或全部来自expert payload。 |
补齐的通信、启动与工作流优化
- QSA稀疏prefill算子:#469 调整pre-Ampere launch profile。单V100、TP2本地布局(Hq12/Hkv1/D256、top-k2048、page16、BF16 cache、KV16K),3次预热+30次CUDA-event样本:64 rows 2.40→1.11 ms,256 rows 9.13→3.54 ms;相对之前可运行的SM70 retune,非整模型倍数。SM75另有共享内存资源修复;数值按原BF16 split归约容差核对。
- GDN copy-only投影融合:#549 合并四次现有输出拷贝,保持两个GEMM与算术边界不变;V100、实际层权重/合成activation、M2/4/8/16/32/64的成对完整投影Graph耗时改善 11–19%。兼容copy路径默认开启;没有独立整模型吞吐声明。
- DCP消除主机同步:#694 将rank当标量处理,移除每步pageable H2D同步;原PR的4×V100/TP4/DCP2、Qwen4Exp QSA/MTP3/E4M3/Graphs短提示decode 57.46→69.15 tok/s,输出与接受率匹配。未给多启动统计;不作为最终wheel或DCP1的版本倍率。
- 跳过无用权重读取:#669 在读取tensor前过滤DSpark不加载的权重;原PR DeepSeek-V4-Flash-NVFP4、DSpark5、PP5(3×V100+2×RTX8000)的一组冷启动,draft加载 255→17秒、整体 10分25秒→6分03秒。这是该模型/部署的历史记录,非27B启动数字。
- **更多显存/启动修复:**共享QSA top-k workspace、减少未用LM-head packing、AWQ/NVFP4逐层释放闲置转换缓存、MXFP4重排峰值、冷编译与稳态KV预算分离、host/device PLE预算及mmap恢复。
- **缓存与通信:**SWA投机lookahead命中、Mamba状态保留、cache-dead块优先回收、分组RAM/FS KV恢复与异步store排序;兼容TP4 push消息路径与DCP开销减少。额外offload配置及真实互联仍必要。
- **其他模型/工作流:**GLM DFlash2、KDA/mHC/grouped experts;原生H3 D128 attention、operand/MLP复用、可选残差分片和ConvRot/weight cache、低RAM mmap masters;H3/Z-Image任务与真实进度、显式NVENC。各模型/硬件/质量条件独立,最终wheel未重测完整工作流。
完整分组和开启条件见变更附件;全部211条PR的CSV已附优化归属。
1.5.1 补充实测
以下项目没有同条件的 1.5.0 数据,只报告新版结果,不计算版本提速。
27B 并发与加速路径
32K C4结果及其完整计量条件见前面的“并发 decode 优化”。
27B 发布 profile strict 配置自检 8/8,Flash-Next 3/3,算子普查 10/10,全新 venv 单 wheel 安装通过。自检使用 VLLM_SM70_REQUIRE_PROFILE_ACCELERATION=1,其余加速配置由启动脚本自动选择。实际通信受 GPU 互联限制:测试卡组使用 PYNCCL;配置自检不代表 CUSTOM/push 通信已经启用。
另一次 27B 显存采样为 26,849 MiB/卡(32K C4、每请求 512 输出上限、T=1/p=.95/k20、memory .80、0.25 秒采样);它不是上面 256-token pure decode 测试的显存读数。
Flash-Next 前缀缓存
RadixArk/Qwen3.8-Flash-Next-NVFP4,4×V100-SXM2-32GB、TP4、FP16 compute/KV、MTP4、maxlen 131072、prefill 8192、maxseq 1、memory .90、同步调度、CUDA Graphs、自动 PLE host 预算;软件版本同上。开/关各一次启动,每长度两次无命中请求;8K 报第二次稳态值,32K 报总计算 tokens / 总 native prefill 时间,包含首 token 工作。固定输入、T=0、seed=0、16 输出上限、正常 EOS。
| 输入 | 开前缀缓存 prefill tok/s | 关前缀缓存 prefill tok/s | 暖请求复用 tokens | 暖请求 prefill 秒 |
|---|---|---|---|---|
| 8K | 5248 | 5978 | 7344 | 0.243 |
| 32K | 4868 | 5191 | 31824 | 0.280 |
这些是 1.5.1 内开/关缓存的对比。无命中时开缓存仍低约 12.22% / 6.22%;重复请求可复用前缀,不能宣称前缀缓存零开销。
编译缓存与重启
27B 首次启动 271.26 秒,同配置第二次启动 95.30 秒。rank 0 首次编译 3 个图,重启加载 3 个图、新编译 0 个。这是 1.5.1 的两次单次观测,说明缓存复用生效;不是相对 1.5.0 的启动提速,也不代表所有设备的启动耗时。
安装与启动
Linux x86_64、CPython 3.12、SM70/V100。objdump -T 实测最高 GLIBC 2.38、GLIBCXX 3.4.21。
这是原生 linux wheel,manylinux 容器门禁未通过。低于声明的 GLIBC 版本不能用此二进制包。
驱动实测 580.173.02;建议采用 CUDA 12.8 Update 1 对应的 570.124.06 或更新驱动,较旧驱动的兼容模式未测。NVIDIA 官方版本表。
仍需正常安装的 CUDA 12.8 Toolkit、C++ 编译器及 ninja,完成剩余 TileLang JIT;标准 Toolkit 的 nvcc 应可从 PATH 找到。不能宣称 Toolkit-free。
依赖由安装器解析,模型权重另行下载;不需要复制开发者 .so、源码 overlay、LD_PRELOAD 或私有库路径。
sha256sum -c SHA256SUMS
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python ./1cat_vllm-1.5.1-cp312-cp312-linux_x86_64.whl
source .venv/bin/activate以下是可直接复制的 vllm serve 参数,适用于互联正常的4×V100-SXM2-32GB。替换本地模型路径,两套命令分别运行。
27B(E4M3 KV / DFlash2):
CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve /models/Qwen3.8-27B-QUASAR-NVFP4 \
--host 127.0.0.1 \
--port 8000 \
--dtype half \
--tensor-parallel-size 4 \
--attention-backend FLASH_ATTN_V100 \
--kv-cache-dtype fp8_e4m3 \
--max-model-len 262144 \
--gpu-memory-utilization 0.8 \
--max-num-batched-tokens 8192 \
--max-num-seqs 4 \
--block-size 2048 \
--mamba-block-size 8192 \
--enable-prefix-caching \
--mamba-cache-mode align \
--served-model-name qwen3.8-27b-dflash2 \
--limit-mm-per-prompt '{"image":0,"video":0}' \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--default-chat-template-kwargs '{"enable_thinking":false}' \
--seed 0 \
--speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","revision":"dedf8df68adfb1afeaf7b7480c0a0243108177b4","num_speculative_tokens":7,"kv_cache_dtype":"auto","attention_backend":"FLASH_ATTN_V100","draft_sample_method":"probabilistic","enforce_eager":false}' \
--trust-remote-codeFlash-Next(FP16 KV / MTP4):
CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve /models/Qwen3.8-Flash-Next-NVFP4 \
--tensor-parallel-size 4 \
--dtype half \
--kv-cache-dtype auto \
--language-model-only \
--max-model-len 131072 \
--max-num-batched-tokens 8192 \
--max-num-seqs 1 \
--gpu-memory-utilization 0.90 \
--no-async-scheduling \
--mamba-cache-mode align \
--enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":4}'27B:TP4、E4M3 目标 KV、FP16 draft KV、7 draft tokens、262144 context、8192 prefill、4 sequences、memory .80、prefix caching、CUDA Graphs。
Flash-Next:TP4、FP16 compute/KV、MTP4、131072 context、8192 prefill、1 sequence、memory .90、prefix caching、同步调度;参数可按需调整。
上述命令无需性能环境变量;编译缓存默认开启,27B按默认策略启用异步调度和CUDA Graphs。加速自检接口为 /v1/sm70/acceleration,状态报告中的 capabilities 需结合 worker 选路日志理解。
Flash-Next 测试环境有充足 host RAM,PLE 表合计约 47.68 GiB,另需模型 mmap、缓存与进程内存;这不是低 RAM 的验证。
27B命令中的草稿模型使用固定revision,首次启动在线下载约 3.85 GB 权重。可提前下载到本地:
huggingface-cli download incoai/Qwen3.8-27B-DFlash2 \
--revision dedf8df68adfb1afeaf7b7480c0a0243108177b4 \
--local-dir /models/Qwen3.8-27B-DFlash2
# Optional downloader installation / 可选下载器安装
uv pip install --python .venv/bin/python modelscope
modelscope download --model incoai/Qwen3.8-27B-DFlash2 \
--revision 5cfe71fd56ed83d5ca27f04c1345cc7d862bbe4f \
--local_dir /models/Qwen3.8-27B-DFlash2使用已下载的草稿模型时,将27B命令中的整个 --speculative-config 参数替换为下面这一项(本地目录不传远端revision):
--speculative-config '{"method":"dflash","model":"/models/Qwen3.8-27B-DFlash2","num_speculative_tokens":7,"kv_cache_dtype":"auto","attention_backend":"FLASH_ATTN_V100","draft_sample_method":"probabilistic","enforce_eager":false}'新版本 Hugging Face CLI 可将 huggingface-cli download 换为 hf download。
下载命令参考:Hugging Face CLI、ModelScope CLI。
驱动建议参考 CUDA 12.8 Update 1 官方说明。
已知问题与验证范围
- 部分并发配置仍有未达到性能验收目标的情况,投机接受率变化的原因尚未完全定位;不能宣称所有并发场景均达到最佳速度。GPU/NVLink 拓扑影响通信路径,但不能将全部性能差异归因于互联。
- manylinux 容器门禁未通过;此 wheel 要求 GLIBC≥2.38,剩余 TileLang JIT 仍需标准 CUDA 12.8 Toolkit、C++ 编译器与 ninja。
- 首次启动需要下载权重和编译;默认编译缓存可减少后续启动开销。Flash-Next 的 PLE host 表约 47.68 GiB,另需其他进程内存;低 RAM 环境未验证。
- 输出仅抽查正常 EOS、乱码与基础回答,不是全面质量评测;一条 27B 前缀缓存解释不准确。
- 未完成同条件旧版 Flash-Next / C4 对比;QUASAR C8/C16、256K 并发、35B AWQ/FP8 完整速度矩阵,以及全部 TP/KV/量化/LoRA/EP/DP/PP 组合未完成最终验证。
- H3 视频/音频、Z-Image、GLM DFlash2、Qwen 多模态及 Bailing 的相关代码变更已纳入;未据此宣称所有工作流均在最终 wheel 完成性能与质量验收。
版本变更与 PR 清单
正式 v1.5.0 (d8f42b39) 到本版源码 (589a5d1e) 共纳入 211 个 merged PR 记录:171 个 main 直接 PR 更新、40 个经集成带入的原始/嵌套 PR。原始 PR 与集成 PR 保留归属,不重复计算优化收益;未合入或晚于产物提交的 PR 不计入。
| 主目的 | PR 数 |
|---|---|
| 性能与内存 | 62 |
| 正确性、安全与兼容 | 88 |
| 模型能力、缓存功能与默认策略 | 33 |
| 构建与发布 | 15 |
| 文档、测试与基准工具 | 13 |
完整变更与逐项链接见附件 CHANGELOG_1.5.0_to_1.5.1.zh-CN.md、CHANGELOG_1.5.0_to_1.5.1.en.md、PR_INDEX.csv;CSV 保留 merge SHA、合并时间、分类与纳入方式。
贡献者致谢
感谢以下贡献者为1.5.1提交代码、修复、测试和文档。按本版纳入的已合并PR提交数量降序排列;数量相同按GitHub账号名排序。范围为正式v1.5.0至本版源码589a5d1e,包含211个不同PR编号(171个直接PR、40个经集成带入的PR),每个PR只计一次。完整归属可查附件 PR_INDEX.csv 的 author 列。
| 贡献者 | 已合并 PR 数 |
|---|---|
| @yangzhuxinyzx | 131 |
| @Peuqui | 36 |
| @Leonccaa | 26 |
| @areslp | 3 |
| @carrey-feng | 3 |
| @hdq66666 | 2 |
| @wfhe | 2 |
| @b89703001 | 1 |
| @dnv2003 | 1 |
| @fiesh | 1 |
| @ga-it | 1 |
| @MacroModel | 1 |
| @ProprietaryLegal | 1 |
| @publee | 1 |
| @zhaochengggg | 1 |
源码:589a5d1e1a9ea7c42d533eed724a7d67c5f1779d,release/1.5.1。
wheel:1cat_vllm-1.5.1-cp312-cp312-linux_x86_64.whl。
SHA256:757f2200ba6bd20f7e3e5c8fc68c7e05461df6d284baf676ac0c16ce19434959。
1Cat-vLLM 1.5.1 release notes
1.5.1 focuses on V100 long-context prefill/decode, compatible TP1/TP2 layouts, concurrency, Flash-Next and memory reuse and default compile caching, with launch profiles and acceleration self-checks included in the wheel.
1.5.0 → 1.5.1: single-request 27B performance
| Input length | 1.5.0 pure decode tok/s | 1.5.1 pure decode tok/s | Speedup |
|---|---|---|---|
| 4K | 103.87 | 229.34 | +120.80% |
| 32K | 49.04 | 201.18 | +310.27% |
| 128K | No completed measurement | 171.43 | Not calculated |
Full speculative-round latency: 40.917 → 19.855 ms at 4K and 81.253 → 21.125 ms at 32K; 1.5.1 measures 25.212 ms at 128K. The first 1.5.0 128K request did not finish within 180 seconds and that test was stopped. No old-version speed or speedup is reported for it.
Conditions: the published 1.5.0 wheel and the final 1.5.1 release wheel, the same four V100-SXM2-32GB GPUs, TP4, QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 and incoai/Qwen3.8-27B-DFlash2. Python 3.12.3, Torch 2.10.0+cu128, CUDA 12.8.93, driver 580.173.02; FP16 compute/draft KV, E4M3 target KV and seven probabilistic draft tokens. Both use the same 1.5.1 release parameters: max length 262144, prefill budget 8192, max sequences 4, memory fraction .80, KV blocks 2048, Mamba blocks 8192, prefix caching, async scheduling and FULL_AND_PIECEWISE CUDA Graphs. Sampling T=1, top-p=.95, top-k=20, seed=0, 256-token cap and normal EOS allowed. This is not a comparison of separately tuned best configurations.
One cold request plus three warm requests per length; the table reports warm medians from one successful startup per version. Pure decode = (output tokens−1) / server decode time, excluding TTFT/prefill. Full speculative round = server decode time / round count, including draft, target, sampling and scheduling. For warm 32K requests, 1.5.0 reuses 28,672 prefix tokens while 1.5.1 recomputes the prompt; both decode with an actual 32,768-token input context. These figures compare decode under the release parameters, not cache hit rates or prefill speedups.
The baseline is the official v1.5.0 wheel, SHA256: 2a4d6bee4e19d315b142f2c563059f3064ddeeca563a6bdc828c33e1073c825b.
Concurrent decode optimization
Concurrency is a focus of1.5.1: improve target verification, draft attention, batched weight supply and sampling while reducing duplicate workspace allocations. Compatible operators select automatically bySM70 hardware, dtype, layout and shape; the recommended arguments need no extra acceleration switches.
- FP4/FP8 batchGEMM and warmup: share activation/decoded-weight supply, coverM32/M48/M64 and tails, and warm prepared batch layouts before importing captured plans. #691, #706, #708.
- Joint draft attention andGDN: combine draft-attention launches across requests and reuse packedGDN, metadata and state workspaces while preserving request-local page tables, lengths and acceptance state. #688, #699.
- Long-context batch verification and sampling: extend existing q8 attention toB2–B16, reuse graph workspaces/logits, reduce sampler synchronization and tune compatibleTP4 reductions. #697, #706, #708.
Matched development comparisons from includedPRs: eachPR has its own before/after baseline. Gains are not additive and are not a formal1.5.0→1.5.1 wheel concurrency ratio. Final-wheel measurements follow below.
| Optimization PR | Concurrency | Control → optimized pure decode tok/s | PR-stage gain |
|---|---|---|---|
| #697: long-context batch attention | C4 | 326.832 → 429.134 | +31.30% |
| Same PR | C8 | 492.025 → 624.863 | +27.00% |
| #706: batch warmup, supply and sampling | C2 | 236.175 → 252.465 | +6.90% |
| Same PR | C4 | 377.689 → 436.004 | +15.44% |
| Same PR | C8 | 612.982 → 709.714 | +15.78% |
Conditions: a27B NVFP4 development checkpoint, fourV100-SXM2-32GB GPUs/TP4, FP16 compute/draftKV, E4M3 targetKV, DFlash2 q7(q8 target verification), max length262144, memory fraction.80, prefix caching andCUDA Graphs. #706 uses32K/256, T=.7/p=.8/k20, matched seeds and ordered admission: one warmup plus three retained waves, reporting medians with identical full output arrays and acceptance counters. #697 uses simultaneous32K/256 waves and three warm all-live decode windows; C4/C8 acceptance changes, so its entire gain cannot be attributed to the attention operator. See its primary report for exact sampling/artifact details. Both exclude initial prefill; single-startup observations do not establish stable gains across machines.
27B concurrency measured with the final1.5.1 wheel:
The QUASAR-QAT release profile measures 445.45 tok/s pure decode at 32K C4. Hardware, model and default launch settings match the single-request test above. This test uses a fixed 32,768-token synthetic input, 256 forced output tokens per request, T=.7/p=.8/k20 and seed=20260923+request index. Shared prefixes are primed; requests enter in fixed order and resume together. One startup, one warmup wave and three measured waves; the median is reported. The three results are 445.48 / 435.35 / 445.45 tok/s.
The metric counts actual returned token IDs in the common decode window between the latest first token and earliest last token, excluding TTFT and new prefill. VLLM_SERVER_DEV_MODE=1 is added only for synchronized admission; it does not change acceleration settings. Forced output is for performance measurement, not quality evidence. This metric differs from overall throughput including prefill.
Highlights
- Long prefill/decode: native Q8000/Q8192 attention, E4M3 KV bridging, decoupled Mamba grids and sensitive FP32 state; compatible large chunks and long/tail CUDA Graph paths dispatch automatically.
- TP1/TP2: broaden local head/projection geometry and repair TP2 GDN alignment. PR tests cover target-only TP1 64K and TP2 256K; conditions and DFlash limits are below.
- Concurrency: FP4/FP8 batch GEMM, joint draft attention, fused GDN metadata/copies and sampling reuse preserve reduction order and normal fallback.
- Flash-Next: FP16 GEMV/HC/GDN follow operator requirements; prefix caching avoids unnecessary prefill splits, with shared native-MTP metadata and draft graphs.
- Memory/capacity: compatible weights/scales/scores/scratch share by default; compact AWQ metadata, concurrency-sized scratch and identical MTP head sharing free space for KV pools.
- Startup/caching: default compile caching, steady KV budgeting, skipping unused weight reads, lower conversion/repacking peaks, safe graph reload and automatic PLE placement.
- Other workflows: grouped RAM/FS KV caching, fewer DCP synchronizations, mixed-device dispatch and native H3/Z-Image compute/host-memory improvements. Optional setup and validation scope are in the complete changelog.
- Installation/self-check: two bundled launchers, local
--draftand configured acceleration checks; no private developer.soor performance environment variables. Standard Toolkit/JIT requirements are below.
Optimization effects and validation scope
These records explain what individual changes included after 1.5.0 solve and what their PR tests measured. They come from different development stages and matched controls, not new final-wheel reruns or a unified1.5.0→1.5.1 speedup. Percentages and memory savings must not be added. Final-wheel measurements follow separately. Originals and integration PRs are not double-counted; already1.5.0 work, excluded PRs and unshipped private candidates are not new gains.
Long prefill, operators and cache scheduling
Native sm70_d256_gqa_architecture_fwd / sm70_d256_gqa_architecture_q8192_fwd execute compatible long-prefix QK/PV on Tensor Cores. sm70_v37_e4m3_bridge brings E4M3 paged KV into shared FP16 workspace while QK/PV accumulate in FP32. Coverage includes Q8000/Q8192, multiple requests and TP-local GQA6/D256 groups. #548, #628, #635, #638, #666. No isolated matched bridge latency is available, so no bridge-only speedup is calculated.
Retained full-FP32 operator tests: one V100-SXM2-32GB, Torch2.10+cu128/CUDA12.8, Q8192/KV262144, CUDA Graph replay including head-group copies:
| Local layout | Attention latency | Useful operator throughput |
|---|---|---|
| TP4,6 query heads /1 KV head | 182.645 ms | 71.111 TFLOP/s |
| TP2,12 query heads /2 KV heads | 369.916 ms | 70.221 TFLOP/s |
| TP1,24 query heads /4 KV heads | 742.363 ms | 69.982 TFLOP/s |
These are PR #666-stage single-GPU local-layout tests, not multi-GPU whole-model throughput or final-wheel reruns. Earlier75–77T samples predate full QK FP32 correction and cannot describe current precision. Power/clocks also matter: the same PR records53.816T at185W. A fixed71T is not promised on users' devices.
| Change / PR | Matched result | Conditions and meaning |
|---|---|---|
| Decoupled Mamba grid #645 | 100K prefill 2954→3336 tok/s,+12.93% | QUASAR27B NVFP4,four V100 GPUs,TP4,E4M3/DFlash7,budget8192/maxseq4/prefix ON,unique-salt100K zero-hit prompts. Fixed KVblock2048,grid4096→8192 enables native attention. PR probes; versus the earlier KV4096 deployment,prefill stays in its measured band while KV capacity grows about23–29%. Launcher adopts KV2048/Mamba8192. |
| Flash-Next prefix prefill #754 | 8K 3060→5356,32K 3046→5068,128K 2818→4666 tok/s; +75.03% /+66.38% /+65.58% | Flash-Next NVFP4,four V100 GPUs,TP4/MTP4,FP16 compute/KV,maxlen131072/budget8192/seq1/memory.90/sync/Graphs,original explicit PLE host12GiB/rank. One startup per arm;8K second zero-hit sample,32K/128K two-request zero-hit totals;T0/seed0/max16/normal EOS. Fewer model forwards across discarded states,preserving reuse; not a new single-kernel multiplier. |
| AWQ indexed-A expert chain #464 | W13→SiLU→W2 6.64→4.97 ms,−25.15% time,avoiding 400 MiB expanded input | Flash-Next AWQ native-g32,V100 TP4,FP16 activations,8192tokens/top-k10/hidden2560,ranks0/3,identical weights/metadata and bitwise outputs. Same-service C1 16K–128K geometric-mean whole-model prefill +5.83%,one warmup plus three scored requests per cell; not NVFP4 27B evidence. |
| General acceleration package #666 | Cold262128-input TTFT 174.343→132.334s,−24.10%;input/TTFT +31.74% | QUASAR27B NVFP4,four V100 GPUs,TP4,E4M3 target/FP16 draft KV,DFlash7,maxlen262144/budget8192/maxseq32/memory.85,Graphs,zero prefix hits and correct natural-EOS retrieval. One recorded cold pair,including first-token work and several paths/fixes; not attention-only or pure-prefill operator timing. |
Restoration of existing GDN routes and missing-extension diagnostics are not counted as new production kernels. Unshipped private GDN candidates are excluded.
TP1/TP2 and memory capacity
#666 dispatches by local heads/projection geometry rather than globally requiringTP4; #517 repairs TP2 GDN physical width8240→8256. Default27B NVFP4 target-only PR tests pass cold TP1 64K /TP2 256K retrieval and C1–C32 offered loads:FP16/E4M3,chunk8192/maxseq32,Graphs;TP1 one32GB GPU,maxlen65536/memory.92;TP2 two32GB GPUs,maxlen262144/memory.85;DFlash off. TP1 reaches9 resident requests and queues higher offered concurrency;TP2 reaches32 resident requests. These are capability checks,not final-wheel old-version scaling ratios.
TP2 DFlash #566 records full B1/q8 rounds 31.885 /29.280 ms for release1k /MBPP28:two V100 GPUs,CUDA12.8/Torch2.10,E4M3/FP32,T1/p.95/k20/normal EOS,three independent startups and five alternating pairs after warmup per fixture. Its full measured combination includes opt-ins and dedicated conditions,not a zero-configuration final-wheel promise. TP1 27B+DFlash2 large-chunk admission limits remain.
The following memory values are per GPU. Original tests primarily use Torch2.10+cu128/CUDA12.8;models/modes are separate. Automatic KV sizing can reuse savings,so total NVML usage need not fall equally.
| Optimization | Matched PR evidence | Scope and activation |
|---|---|---|
| Shared weights/codes/scales #561,#671 | Model loading 10.07→6.47 GiB;available KV 10.57→12.68 GiB | PR671,QUASAR NVFP4,TP4/DFlash7,E4M3 target/FP16 draft KV,V10032GB/185W,maxlen262144/chunk8192/seq32/memory.85/Graphs. Compatible sharing is now default;later promotion supersedes the earlier opt-in status. |
| Score/graph scratch/draft snapshots #673,#677 | Non-KV active memory:TP2 17.117→15.146 GiB,TP4 10.225→8.794 GiB | PR677,QUASAR NVFP4/DFlash7,E4M3 and FP16 draft KV,V10032GB,maxlen262144/chunk8192/seq32/memory.90/Graphs,default ON/OFF comparison. Current score block8192;some prefill latency trades about1–2%. Do not add savings to the previous row. |
| Compact AWQ metadata #473,#522 | About 0.879 GiB less resident metadata | Flash-Next AWQ TP4/E512/native-g32,compatible3-byte layout default;other shapes fall back. Operator tests preserve decoded weights/outputs. |
| AWQ resident scratch #465 | Loading 21.24→21.11 GiB;KV capacity 771736→789443 | V100 TP4,MTP0,seq8;automatic sizing follows concurrency/MTP verifier width and retains the cap. Not a universal capacity guarantee. |
| Multimodal target/draft head sharing #594 | Post-request idle 30712→30412 MiB | Flash-Next AWQ-g32,four V100 GPUs,TP4,FP16 head,MTP3/C2/Graphs,fixed5GiB KV,four deterministic requests. Identical heads share automatically;not a loading-peak reduction. |
| FP8-resident MTP experts #595 | Native FP8 checkpoint controls save 1010 MiB AWQ /892 MiB NVFP4 | Original development environment,TP4/MTP3/C2,FP16 compute/KV,fixed5GiB KV/Graphs,idle after greedy/stochastic campaigns. Native-checkpoint support and explicit online conversion are separate. Not a final-MTP4-profile rerun or entirely expert-payload savings. |
Additional communication, startup and workflow optimizations
- QSA sparse-prefill operator: #469 retunes pre-Ampere launch profiles. One V100,TP2-local Hq12/Hkv1/D256,top-k2048,page16,BF16 cache,KV16K,three warmups and30 CUDA-event samples:64 rows 2.40→1.11 ms,256 rows 9.13→3.54 ms versus the previously runnable SM70 retune,not whole-model multipliers. SM75 shared-memory resource failures are also repaired;outputs use the recorded BF16 split-reduction tolerance.
- GDN copy-only projection fusion: #549 combines four existing output copies while retaining both GEMMs/arithmetic boundaries. V100,real layer weights/synthetic activations,M2/4/8/16/32/64,paired complete-projection Graph latency improves 11–19%. Compatible copies default on;no independent whole-model throughput is claimed.
- Remove DCP host sync: #694 treats rank as a scalar,avoiding per-step pageable H2D sync. Its four-V100/TP4/DCP2,Qwen4Exp QSA/MTP3/E4M3/Graphs short-prompt decode is 57.46→69.15 tok/s,with matching outputs/acceptance. No multi-startup statistics are provided;not final-wheel or DCP1 version speedups.
- Skip unused weight reads: #669 filters unused DSpark tensors before reading. Its recorded DeepSeek-V4-Flash-NVFP4/DSpark5/PP5 cold pair (three V100 +two RTX8000) reduces draft loading 255→17s and total boot 10m25s→6m03s. Historical model/deployment observations,not27B boot measurements.
- More memory/startup changes: shared QSA top-k workspace,skipping unused LM-head packing,per-layer idle AWQ/NVFP4 conversion-cache release,lower MXFP4 repacking peaks,separating cold compile from steady KV profiling,and host/device PLE budgeting/mmap restore.
- Caching/communication: SWA speculative lookahead hits,Mamba retention,cache-dead-first eviction,grouped RAM/FS KV restore and ordered asynchronous stores;compatible TP4 push message routes and fewer DCP stalls. Extra offload setup and actual interconnect still matter.
- Other models/workflows: GLM DFlash2,KDA/mHC/grouped experts;native H3 D128 attention,operand/MLP reuse,optional residual sharding and ConvRot/weight cache,low-RAM mmap masters;H3/Z-Image jobs/real progress and explicit NVENC. Model/hardware/quality conditions remain separate;complete final-wheel workflows were not rerun.
Full group/activation scope is in the changelog attachments;CSV maps all 211 PR records to optimization areas.
Additional 1.5.1 measurements
No matched 1.5.0 measurements are available for these items. Only new-version results are reported; no version speedup is calculated.
27B concurrency and acceleration paths
The 32K C4 result and its complete measurement conditions are listed in the earlier concurrent-decode section.
Strict configured-capability checks pass 8/8 for the 27B release profile and 3/3 for Flash-Next. Operator census passes 10/10, and installation from one wheel into a fresh venv passes. Self-checks use VLLM_SM70_REQUIRE_PROFILE_ACCELERATION=1; launchers choose the acceleration configuration automatically. Actual communication depends on GPU interconnect: the test group uses PYNCCL. A configuration check does not prove CUSTOM/push communication is active.
A separate 27B memory sample measured 26,849 MiB/GPU (32K C4, 512-token output cap per request, T=1/p=.95/k20, memory .80, sampled every .25 seconds). This is not the memory reading for the 256-token pure-decode test above.
Flash-Next prefix caching
RadixArk/Qwen3.8-Flash-Next-NVFP4, four V100-SXM2-32GB GPUs, TP4, FP16 compute/KV, MTP4, max length 131072, prefill budget 8192, one sequence, memory .90, synchronous scheduling, CUDA Graphs and automatic host PLE budgeting; software versions match those above. One startup per side, two zero-hit requests per length. 8K reports the second steady sample; 32K reports total computed tokens / total native prefill time, including first-token work. Fixed inputs, T=0, seed=0, 16-token cap and normal EOS.
| Input | Prefix ON prefill tok/s | Prefix OFF prefill tok/s | Warm reused tokens | Warm prefill seconds |
|---|---|---|---|---|
| 8K | 5248 | 5978 | 7344 | 0.243 |
| 32K | 4868 | 5191 | 31824 | 0.280 |
This compares cache settings within 1.5.1. With no hit, enabling prefix caching remains approximately 12.22% / 6.22% slower. Repeated requests can reuse prefixes; zero cache overhead is not claimed.
Compile cache and restart
The first 27B startup took 271.26 seconds; the second with the same configuration took 95.30 seconds. Rank 0 compiled three graphs initially, then loaded three graphs with zero new compilations after restart. These are two single observations within 1.5.1 demonstrating cache reuse, not a startup speedup over 1.5.0 or a prediction for every device.
Installation and launch
Linux x86_64, CPython 3.12, SM70/V100. objdump -T measured a maximum requirement of GLIBC 2.38 and GLIBCXX 3.4.21.
This native Linux wheel has not passed the manylinux container gate. Systems below the stated GLIBC requirement cannot use this binary.
Tested driver: 580.173.02. Recommended: 570.124.06 or newer; older-driver compatibility modes were not tested.
A standard CUDA 12.8 Toolkit, C++ compiler and ninja are still needed for remaining TileLang JIT, with nvcc discoverable on PATH. This is not a Toolkit-free artifact.
The installer resolves declared Python/Torch dependencies. Download model weights separately. No private developer DSO, source overlay, LD_PRELOAD or private library path is needed.
sha256sum -c SHA256SUMS
uv venv --python 3.12 .venv
uv pip install --python .venv/bin/python ./1cat_vllm-1.5.1-cp312-cp312-linux_x86_64.whl
source .venv/bin/activateCopy these explicit vllm serve arguments for four peer-connected V100-SXM2-32GB GPUs. Replace the local model directories and run the two commands separately.
27B (E4M3 KV / DFlash2):
CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve /models/Qwen3.8-27B-QUASAR-NVFP4 \
--host 127.0.0.1 \
--port 8000 \
--dtype half \
--tensor-parallel-size 4 \
--attention-backend FLASH_ATTN_V100 \
--kv-cache-dtype fp8_e4m3 \
--max-model-len 262144 \
--gpu-memory-utilization 0.8 \
--max-num-batched-tokens 8192 \
--max-num-seqs 4 \
--block-size 2048 \
--mamba-block-size 8192 \
--enable-prefix-caching \
--mamba-cache-mode align \
--served-model-name qwen3.8-27b-dflash2 \
--limit-mm-per-prompt '{"image":0,"video":0}' \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--default-chat-template-kwargs '{"enable_thinking":false}' \
--seed 0 \
--speculative-config '{"method":"dflash","model":"incoai/Qwen3.8-27B-DFlash2","revision":"dedf8df68adfb1afeaf7b7480c0a0243108177b4","num_speculative_tokens":7,"kv_cache_dtype":"auto","attention_backend":"FLASH_ATTN_V100","draft_sample_method":"probabilistic","enforce_eager":false}' \
--trust-remote-codeFlash-Next (FP16 KV / MTP4):
CUDA_VISIBLE_DEVICES=0,1,2,3 \
vllm serve /models/Qwen3.8-Flash-Next-NVFP4 \
--tensor-parallel-size 4 \
--dtype half \
--kv-cache-dtype auto \
--language-model-only \
--max-model-len 131072 \
--max-num-batched-tokens 8192 \
--max-num-seqs 1 \
--gpu-memory-utilization 0.90 \
--no-async-scheduling \
--mamba-cache-mode align \
--enable-prefix-caching \
--speculative-config '{"method":"mtp","num_speculative_tokens":4}'27B defaults: TP4, E4M3 target KV, FP16 draft KV, seven drafts, 262144 context, 8192 prefill, four sequences, memory .80, prefix caching and CUDA Graphs.
Flash-Next defaults: TP4, FP16 compute/KV, MTP4, 131072 context, 8192 prefill, one sequence, memory .90, prefix caching and synchronous scheduling. Adjust the arguments as needed.
These commands need no performance environment variables. Compile caching is enabled by default; 27B selects asynchronous scheduling and CUDA Graphs through the default policy. Inspect /v1/sm70/acceleration together with worker dispatch logs; its capability report is not a per-kernel execution trace.
The Flash-Next test machine has ample host RAM: PLE tables total approximately 47.68 GiB, plus model mmap, caches and process memory. Low-RAM operation was not qualified.
The 27B command pins the draft revision and downloads approximately 3.85 GB of weights on first startup. Download ahead of time to a local directory:
huggingface-cli download incoai/Qwen3.8-27B-DFlash2 \
--revision dedf8df68adfb1afeaf7b7480c0a0243108177b4 \
--local-dir /models/Qwen3.8-27B-DFlash2
# Optional downloader installation / 可选下载器安装
uv pip install --python .venv/bin/python modelscope
modelscope download --model incoai/Qwen3.8-27B-DFlash2 \
--revision 5cfe71fd56ed83d5ca27f04c1345cc7d862bbe4f \
--local_dir /models/Qwen3.8-27B-DFlash2To use the downloaded draft, replace the entire --speculative-config argument in the 27B command with the following (omit the remote revision for a local directory):
--speculative-config '{"method":"dflash","model":"/models/Qwen3.8-27B-DFlash2","num_speculative_tokens":7,"kv_cache_dtype":"auto","attention_backend":"FLASH_ATTN_V100","draft_sample_method":"probabilistic","enforce_eager":false}'With the newer Hugging Face CLI, replace huggingface-cli download with hf download.
References: Hugging Face CLI, ModelScope CLI, and CUDA 12.8 Update 1 driver notes.
Known issues and validation scope
- Some concurrency configurations still miss performance acceptance targets, and changes in speculative acceptance are not fully attributed. Best speed across all concurrency settings is not claimed. GPU/NVLink topology affects communication selection but is not an established explanation for every performance difference.
- The manylinux container gate has not passed. This wheel requires GLIBC≥2.38; remaining TileLang JIT requires a standard CUDA 12.8 Toolkit, C++ compiler and ninja.
- First startup includes weight downloads and compilation. Default compile caching reduces subsequent compilation overhead. Flash-Next PLE host tables total approximately 47.68 GiB, with additional process memory required; low-RAM operation was not validated.
- Outputs were spot-checked for normal EOS, corruption and basic answers, not broad quality. One 27B explanation of prefix caching was inaccurate.
- Matched old-version Flash-Next / C4 comparisons are unavailable. QUASAR C8/C16, 256K concurrency, the full 35B AWQ/FP8 speed matrix and every TP/KV/quantization/LoRA/EP/DP/PP combination have not completed final validation.
- Changes for H3 video/audio, Z-Image, GLM DFlash2, Qwen multimodal and Bailing are included; complete final-wheel performance and quality qualification is not claimed for every workflow.
Version changes and PR inventory
The official v1.5.0 (d8f42b39) to this artifact's source (589a5d1e) includes 211 merged PR records: 171 direct main PR updates and 40 original/nested PRs brought in through integration. Original and integration records preserve provenance; optimization gains are not counted twice. Unmerged or later PRs are excluded.
| Principal purpose | PRs |
|---|---|
| Performance and memory | 62 |
| Correctness, security and compatibility | 88 |
| Model/caching features and defaults | 33 |
| Build and release | 15 |
| Documentation, tests and benchmark tools | 13 |
Full changes and individual links are attached in CHANGELOG_1.5.0_to_1.5.1.zh-CN.md, CHANGELOG_1.5.0_to_1.5.1.en.md and PR_INDEX.csv. CSV retains merge SHA, merge date, classification and integration provenance.
Contributor thanks
Thank you to the contributors who submitted code, fixes, tests and documentation for1.5.1. Ranked by the number of merged PR submissions included in this release, descending; ties use GitHub login order. The scope is officialv1.5.0 to artifact source589a5d1e:211 distinct PR numbers (171 direct and40 brought in through integration), each counted once. Full attribution is available in the author column of the attached PR_INDEX.csv.
| Contributor | Merged PRs |
|---|---|
| @yangzhuxinyzx | 131 |
| @Peuqui | 36 |
| @Leonccaa | 26 |
| @areslp | 3 |
| @carrey-feng | 3 |
| @hdq66666 | 2 |
| @wfhe | 2 |
| @b89703001 | 1 |
| @dnv2003 | 1 |
| @fiesh | 1 |
| @ga-it | 1 |
| @MacroModel | 1 |
| @ProprietaryLegal | 1 |
| @publee | 1 |
| @zhaochengggg | 1 |
Source: 589a5d1e1a9ea7c42d533eed724a7d67c5f1779d, release/1.5.1.
Wheel: 1cat_vllm-1.5.1-cp312-cp312-linux_x86_64.whl.
SHA256: 757f2200ba6bd20f7e3e5c8fc68c7e05461df6d284baf676ac0c16ce19434959.