Skip to content

[Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.19-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 - #3012

Open
Klaud-Cold wants to merge 3 commits into
mainfrom
klaud/auto-0c1068466c38565f-013203e0f89888cf
Open

[Klaud Cold] Update dsv4-fp4-b200-sglang-agentic-hicache-mtp SGLang image to v0.5.19-cu130 / 将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像更新至 v0.5.19-cu130#3012
Klaud-Cold wants to merge 3 commits into
mainfrom
klaud/auto-0c1068466c38565f-013203e0f89888cf

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator

Update the dsv4-fp4-b200-sglang-agentic-hicache-mtp image from the 2026-08-27 dev nightly lmsysorg/sglang:nightly-dev-20260827-20621aa1 to the SGLang v0.5.19 release image lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, pushed 2026-09-04). Model, DSpark speculative decoding, HiCache profiles, search space and runner are unchanged.

Baseline

  • Published date: 2026-09-08 (public dashboard API: workflow-info?date=2026-09-08&benchmarkType=agentic_traces, benchmarks?model=DeepSeek-V4-Pro&date=2026-09-08&exact=true&sequence=agentic-traces, evaluations?model=DeepSeek-V4-Pro&date=2026-09-08)
  • Old image: lmsysorg/sglang:nightly-dev-20260827-20621aa1 (sglang commit 20621aa1, 2026-08-27)
  • New image source: v0.5.19 at 0bcd8223; the release is 317 commits ahead of and 0 behind the pinned nightly (compare)
  • Workload / topology: AgentX agentic-coding traces, DeepSeek-V4-Pro-0813 NVFP4, DSpark block 6 (1 step, 7 draft tokens), single node cluster:b200-nscale; TP8 no offload (c1-5), TP8 HiCache DRAM (c8-16), TP8/EP8/DP-attention HiCache DRAM with sglang-router 0.3.2 (c64-160)
  • Producer: run 33830419034 attempt 2, head 874e8a66297b79ff38fcb9b608396c35430e24cd (Use DSpark6 for B200 DSV4 SGLang AgentX / B200 DSV4 SGLang AgentX 使用 DSpark6 #2821)
conc KV offload out tok/s/GPU median TPOT ms median TTFT ms median E2E s
1 none 17.6 3.36 830 3.74
2 none 20.6 3.36 509 2.29
3 none 25.7 3.49 523 2.24
4 none 27.4 3.58 518 2.34
5 none 30.7 3.70 474 3.07
8 hicache 48.9 4.39 491 2.76
10 hicache 59.0 4.80 531 3.00
16 hicache 81.3 5.95 552 3.24
64 hicache DEP8 270.8 13.79 1770 9.04
96 hicache DEP8 367.0 16.72 2512 13.11
128 hicache DEP8 398.4 17.88 4234 17.54
160 hicache DEP8 388.2 20.40 12649 27.17
  • Published eval: gsm8k at c160 (DEP8 HiCache), em_strict 0.9666 / em_flexible 0.9659, n=1319 (run 33830419034). The published eval row labels the deployment disagg: true, which mismatches the single-node aggregated recipe; the run itself is the aggregated job.

dsv4-fp4-b200-sglang-agentic-hicache-mtp 的镜像从 2026-08-27 的开发 nightly lmsysorg/sglang:nightly-dev-20260827-20621aa1 更新至 SGLang v0.5.19 正式版镜像 lmsysorg/sglang:v0.5.19-cu130(Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,2026-09-04 推送)。模型、DSpark 投机解码、HiCache 配置、搜索空间与 runner 均保持不变。

基线

  • 发布日期: 2026-09-08(公开仪表盘 API:workflow-info?date=2026-09-08&benchmarkType=agentic_tracesbenchmarks?model=DeepSeek-V4-Pro&date=2026-09-08&exact=true&sequence=agentic-tracesevaluations?model=DeepSeek-V4-Pro&date=2026-09-08
  • 旧镜像: lmsysorg/sglang:nightly-dev-20260827-20621aa1(sglang 提交 20621aa1,2026-08-27)
  • 新镜像来源: v0.5.19,提交 0bcd8223;发布版本领先所固定 nightly 317 个提交、落后 0 个(对比
  • 负载 / 拓扑: AgentX agentic-coding 轨迹回放,DeepSeek-V4-Pro-0813 NVFP4,DSpark block 6(1 步、7 个草稿 token),单节点 cluster:b200-nscale;TP8 无卸载(c1-5)、TP8 HiCache DRAM(c8-16)、TP8/EP8/DP-attention HiCache DRAM 配 sglang-router 0.3.2(c64-160)
  • 生产运行: run 33830419034 第 2 次尝试,head 874e8a66297b79ff38fcb9b608396c35430e24cdUse DSpark6 for B200 DSV4 SGLang AgentX / B200 DSV4 SGLang AgentX 使用 DSpark6 #2821
并发 KV 卸载 输出 tok/s/GPU 中位 TPOT ms 中位 TTFT ms 中位 E2E s
1 17.6 3.36 830 3.74
2 20.6 3.36 509 2.29
3 25.7 3.49 523 2.24
4 27.4 3.58 518 2.34
5 30.7 3.70 474 3.07
8 hicache 48.9 4.39 491 2.76
10 hicache 59.0 4.80 531 3.00
16 hicache 81.3 5.95 552 3.24
64 hicache DEP8 270.8 13.79 1770 9.04
96 hicache DEP8 367.0 16.72 2512 13.11
128 hicache DEP8 398.4 17.88 4234 17.54
160 hicache DEP8 388.2 20.40 12649 27.17
  • 已发布评测: c160(DEP8 HiCache)gsm8k,em_strict 0.9666 / em_flexible 0.9659,n=1319(run 33830419034)。已发布评测行将部署标记为 disagg: true,与单节点聚合配方不符;该运行本身是聚合作业。

🤖 Generated with Claude Code


Note

Low Risk
Config-only Docker pin plus changelog; no application code or benchmark script edits in the diff.

Overview
Pins dsv4-fp4-b200-sglang-agentic-hicache-mtp to the release image lmsysorg/sglang:v0.5.19-cu130, replacing the nightly-dev-20260827-20621aa1 dev build. Model, agentic-coding search space, HiCache/MTP/DSpark benchmark recipe, and cluster:b200-nscale runner are unchanged.

Adds a perf-changelog.yaml entry for this config key noting the image swap and that the benchmark script flags stay the same while the container bumps DeepGEMM (0.1.5.post2 → 0.1.7), FlashInfer (0.6.17 → 0.6.18), and routes W4A4 MXFP4 MegaMoE through --enable-w4a4-mxfp4-megamoe instead of DG_USE_FP4_ACTS / DG_USE_MXF4_KIND env vars.

Reviewed by Cursor Bugbot for commit da239a6. Bugbot is set up for automated code reviews on this repo. Configure here.

…to v0.5.19-cu130

Replace the 2026-08-27 dev nightly lmsysorg/sglang:nightly-dev-20260827-20621aa1
with the v0.5.19 release image lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9). The
release branch contains every commit of the pinned nightly; model, DSpark
speculative settings, HiCache profiles, search space and runner are unchanged.

将 dsv4-fp4-b200-sglang-agentic-hicache-mtp 的 SGLang 镜像从 2026-08-27 的开发
nightly lmsysorg/sglang:nightly-dev-20260827-20621aa1 更新至 v0.5.19 正式版镜像
lmsysorg/sglang:v0.5.19-cu130(Docker Hub digest
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9)。
发布分支包含该 nightly 的全部提交;模型、DSpark 投机解码设置、HiCache 配置、
搜索空间与 runner 均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / head: lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9) at 82ded8908a7ee201a5fb36ade101384c6c514b4c
  • Targeted run: 34593186565e2e-tests.yml on main, test-config --config-files configs/nvidia-master.yaml --config-keys dsv4-fp4-b200-sglang-agentic-hicache-mtp --trim-conc, fail-fast, Klaud background priority
  • Smoke points (lowest concurrency per deployment shape): TP8 no offload c1; TP8 HiCache DRAM c8; TP8/EP8/DP-attention HiCache DRAM c64 (router). This is startup/compatibility evidence only, not a full curve.
  • Change: one line in configs/nvidia-master.yaml (family image). Script, DSpark settings, HiCache ratios, search space and runner are untouched.
  • Upstream source review (old 20621aa1 → new 0bcd8223, compare, 317 commits, release contains the nightly):
    • Provenance: nightly-dev-20260827-20621aa1 and nightly-dev-cu13-20260827-20621aa1 share one Docker Hub digest, so the old pin is already a CUDA 13.0 build; v0.5.19 and v0.5.19-cu130 share digest d6e72886…, built by release-docker.yml from the tag. Both Dockerfiles use nvidia/cuda:13.0.3, Python 3.12, torch 2.13.0.
    • Every server flag used by benchmarks/single_node/agentic/dsv4_fp4_b200_sglang_mtp.sh exists in both server_args.py revisions with unchanged defaults/choices (--speculative-algorithm DSPARK, --speculative-dspark-block-size, HiCache ratio/write-policy/io-backend/mem-layout, --enable-dp-lm-head, --enable-dp-attention-local-control-broadcast, --load-balance-method, --load-snapshot-publish-interval, --prefill-decode-interval, --weight-loader-prefetch-checkpoints, --swa-full-tokens-ratio, parsers deepseekv4 / deepseek-v4). --moe-a2a-backend choices only gained deepep_v2/ascend_tp; megamoe is retained.
    • --enable-w4a4-mxfp4-megamoe implementation changed from exporting DG_USE_FP4_ACTS/DG_USE_MXF4_KIND (arg_groups/mega_moe_hook.py) to selecting the DeepGEMM mxf4xmxf4 MMA type directly in layers/moe/mega_moe.py; the script already passes the flag, so the W4A4 path stays selected. DeepGEMM 0.1.5.post2 → 0.1.7, FlashInfer 0.6.17 → 0.6.18 (#36954).
    • All SGLANG_* env vars the script sets keep identical definitions in environ.py (SGLANG_SIMULATE_ACC_*, SGLANG_OPT_*, SGLANG_ENABLE_UNIFIED_RADIX_TREE, SGLANG_TIMEOUT_KEEP_ALIVE); SGLANG_OPT_USE_JIT_NORM is undefined in both revisions.
    • Bundled sglang-router is 0.3.2 at both commits and router_args.py/launch_router.py are byte-identical, matching the config's router.version.
    • Relevant fixes in range: DSpark duplicated draft sample_block (#36934), DSpark/DFlash TP-rank state divergence (#33614), DP-attention decode→extend prefix off-by-one (#37505), SWA predicate convergence (#37550), MegaMoE SM reservation on Blackwell (#36657), HiCache host-pool mmap/registration fixes (#36705, #36798, #37026).
    • v0.5.19-cu130 already produced published results on cluster:b200-nscale (qwen3.5-fp8/fp4-b200-sglang, 2026-09-08/10), so CUDA/driver compatibility on the target cluster is established.
  • Result: targeted smoke passed — run 34593186565 completed success (11:16–12:46 UTC); all three agentic jobs green, aggregates present, power valid, no eval selected by the trimmed matrix. Server logs: enable_w4a4_mxfp4_megamoe on the DEP8 path, DSpark real-draft-token simulation active, no server errors (only psutil NoSuchProcess noise from the cpu_monitor thread at shutdown). Warn-only notices: --cuda-graph-max-bs is a deprecated alias of --cuda-graph-max-bs-decode (same alias in the old image) and SGLANG_ENABLE_UNIFIED_RADIX_TREE is now the default and ignored.
  • Smoke deltas vs published 2026-09-08 baseline (same topology, concurrency and dataset per point; smoke-only, not a curve):
point out tok/s/GPU new / base Δ p50 TPOT ms new / base Δ p50 TTFT ms new / base Δ p50 E2E s new / base Δ requests ok
c1 TP8 none 17.60 / 17.59 +0.0% 3.33 / 3.36 −0.9% 869 / 830 +4.7% 3.82 / 3.74 +1.9% 261/272
c8 TP8 HiCache 49.57 / 48.94 +1.3% 4.35 / 4.39 −0.9% 470 / 491 −4.3% 2.71 / 2.76 −1.7% 1288/1375
c64 DEP8 HiCache 285.97 / 270.80 +5.6% 12.64 / 13.79 −8.3% 1664 / 1770 −6.0% 8.54 / 9.04 −5.5% 8394/9101

Dropped request counts are the warmup records excluded by AIPerf accounting (no error drops). GPU prefix-cache hit rate: 0.978 / 0.973 / 0.863 vs baseline 0.964 / 0.959 / 0.876.

  • Next step: append the perf-changelog.yaml entry, recheck capacity and start the final full sweep (full-sweep-enabled, PR kept draft).

初始尝试

  • 镜像 / head: lmsysorg/sglang:v0.5.19-cu130(digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9),提交 82ded8908a7ee201a5fb36ade101384c6c514b4c
  • 定向运行: 34593186565——在 main 上的 e2e-tests.ymltest-config --config-files configs/nvidia-master.yaml --config-keys dsv4-fp4-b200-sglang-agentic-hicache-mtp --trim-conc,fail-fast,Klaud 后台优先级
  • 冒烟点(每种部署形态的最低并发): TP8 无卸载 c1;TP8 HiCache DRAM c8;TP8/EP8/DP-attention HiCache DRAM c64(router)。仅作为启动/兼容性证据,不是完整曲线。
  • 改动: configs/nvidia-master.yaml 中一行(该族镜像)。脚本、DSpark 设置、HiCache 比例、搜索空间与 runner 未改动。
  • 上游源码审查(旧 20621aa1 → 新 0bcd8223对比,317 个提交,发布版本包含该 nightly):
    • 来源确认:nightly-dev-20260827-20621aa1nightly-dev-cu13-20260827-20621aa1 在 Docker Hub 上共用同一 digest,因此旧镜像本身就是 CUDA 13.0 构建;v0.5.19v0.5.19-cu130 共用 digest d6e72886…,由 release-docker.yml 从 tag 构建。两版 Dockerfile 均基于 nvidia/cuda:13.0.3、Python 3.12、torch 2.13.0。
    • benchmarks/single_node/agentic/dsv4_fp4_b200_sglang_mtp.sh 使用的所有服务端参数在两版 server_args.py 中均存在且默认值/可选项未变(--speculative-algorithm DSPARK--speculative-dspark-block-size、HiCache 的 ratio/write-policy/io-backend/mem-layout、--enable-dp-lm-head--enable-dp-attention-local-control-broadcast--load-balance-method--load-snapshot-publish-interval--prefill-decode-interval--weight-loader-prefetch-checkpoints--swa-full-tokens-ratio、解析器 deepseekv4 / deepseek-v4)。--moe-a2a-backend 仅新增 deepep_v2/ascend_tp 选项,megamoe 保留。
    • --enable-w4a4-mxfp4-megamoe 的实现从导出 DG_USE_FP4_ACTS/DG_USE_MXF4_KINDarg_groups/mega_moe_hook.py)改为在 layers/moe/mega_moe.py 中直接选择 DeepGEMM 的 mxf4xmxf4 MMA 类型;脚本已传入该 flag,W4A4 路径保持生效。DeepGEMM 0.1.5.post2 → 0.1.7,FlashInfer 0.6.17 → 0.6.18(#36954)。
    • 脚本设置的所有 SGLANG_* 环境变量在 environ.py 中定义一致(SGLANG_SIMULATE_ACC_*SGLANG_OPT_*SGLANG_ENABLE_UNIFIED_RADIX_TREESGLANG_TIMEOUT_KEEP_ALIVE);SGLANG_OPT_USE_JIT_NORM 在两版中均未定义。
    • 两个提交内置的 sglang-router 均为 0.3.2,router_args.py/launch_router.py 逐字节相同,与配置中的 router.version 一致。
    • 范围内的相关修复:DSpark 重复草稿 sample_block#36934)、DSpark/DFlash TP rank 状态分歧(#33614)、DP-attention decode→extend 前缀差一错误(#37505)、SWA 判定统一(#37550)、Blackwell 上 MegaMoE SM 预留(#36657)、HiCache 主机池 mmap/注册修复(#36705#36798#37026)。
    • v0.5.19-cu130 已在 cluster:b200-nscale 上产出已发布结果(qwen3.5-fp8/fp4-b200-sglang,2026-09-08/10),目标集群的 CUDA/驱动兼容性已确认。
  • 结果:定向冒烟通过——运行 34593186565success 结束(UTC 11:16–12:46);三个 agentic 作业全部通过,聚合结果齐全,功耗数据有效,裁剪后的矩阵未选中评测。服务端日志:DEP8 路径 enable_w4a4_mxfp4_megamoe 生效,DSpark real-draft-token 模拟启用,无服务端错误(仅有 cpu_monitor 线程在关闭时的 psutil NoSuchProcess 噪音)。仅告警项:--cuda-graph-max-bs--cuda-graph-max-bs-decode 的废弃别名(旧镜像中同为别名),SGLANG_ENABLE_UNIFIED_RADIX_TREE 现为默认行为并被忽略。
  • 冒烟结果相对 2026-09-08 已发布基线的差异(逐点同拓扑、同并发、同数据集;仅冒烟,非完整曲线):
点位 输出 tok/s/GPU 新 / 基线 Δ p50 TPOT ms 新 / 基线 Δ p50 TTFT ms 新 / 基线 Δ p50 E2E s 新 / 基线 Δ 成功请求
c1 TP8 无卸载 17.60 / 17.59 +0.0% 3.33 / 3.36 −0.9% 869 / 830 +4.7% 3.82 / 3.74 +1.9% 261/272
c8 TP8 HiCache 49.57 / 48.94 +1.3% 4.35 / 4.39 −0.9% 470 / 491 −4.3% 2.71 / 2.76 −1.7% 1288/1375
c64 DEP8 HiCache 285.97 / 270.80 +5.6% 12.64 / 13.79 −8.3% 1664 / 1770 −6.0% 8.54 / 9.04 −5.5% 8394/9101

被剔除的请求为 AIPerf 统计中排除的预热记录(无错误剔除)。GPU 前缀缓存命中率:0.978 / 0.973 / 0.863,基线为 0.964 / 0.959 / 0.876。

  • 下一步: 追加 perf-changelog.yaml 条目,重新检查容量后启动最终完整扫描(加 full-sweep-enabled,PR 保持草稿)。

…te to v0.5.19-cu130

Append the perf-changelog entry for the B200 DeepSeek-V4-Pro-0813 SGLang AgentX
image update in PR #3012 after the targeted smoke run passed.

在定向冒烟运行通过后,为 PR #3012 中 B200 DeepSeek-V4-Pro-0813 SGLang AgentX
的镜像更新追加 perf-changelog 条目。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…mage bump branch

Resolve the perf-changelog.yaml append conflict by restoring main's file and
re-appending only this branch's entry, so the labeled PR can trigger run-sweep.

通过恢复 main 的 perf-changelog.yaml 并仅重新追加本分支条目来解决追加冲突,
使带标签的 PR 能够触发 run-sweep。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Milestone — final sweep unblocked by merging main. After the Initial attempt smoke passed, the changelog entry was pushed as 82091849 and full-sweep-enabled applied at 12:49 UTC, but run-sweep.yml created no run: the PR was CONFLICTING on the append-only perf-changelog.yaml because main had gained new tail entries since the base, and GitHub does not fire path-filtered pull_request workflows for a PR whose merge commit cannot be built. Merged main (e5f3e41c) into the branch, restoring main's changelog bytes and re-appending only this PR's entry; the new head da239a65 still changes exactly the family image line plus the one changelog entry. This is branch maintenance, not an image repair (repairs used: 0/5).


里程碑——通过合并 main 解除最终扫描阻塞。 初始尝试冒烟通过后,changelog 条目以 82091849 推送并于 UTC 12:49 添加 full-sweep-enabled,但 run-sweep.yml 未创建任何运行:由于 main 在基线之后新增了尾部条目,PR 在仅追加的 perf-changelog.yaml 上处于 CONFLICTING 状态,而 GitHub 不会为无法构建合并提交的 PR 触发带路径过滤的 pull_request 工作流。已将 maine5f3e41c)合并进分支,恢复 main 的 changelog 字节并仅重新追加本 PR 条目;新 head da239a65 仍只改动该族镜像行与一条 changelog 条目。这是分支维护,不是镜像修复(已用修复次数:0/5)。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Image / head: lmsysorg/sglang:v0.5.19-cu130 at da239a6585566fec56766c7185ecd22e147616af (main e5f3e41c merged; PR diff is the family image line plus one changelog entry)
  • Run: 34603861241run-sweep.yml on the PR head via full-sweep-enabled (PR kept draft), Klaud background priority
  • Scope: the complete dsv4-fp4-b200-sglang-agentic-hicache-mtp family selected by the changelog — 12 AgentX points (TP8 none c1–5; TP8 HiCache c8/10/16; TP8/EP8/DP-attention HiCache c64/96/128/160 with sglang-router 0.3.2) plus the default gsm8k eval at c160, all on cluster:b200-nscale. Nothing trimmed.
  • Status: running; results, per-point deltas against the published 2026-09-08 baseline and the eval score will be added here.

最终完整扫描

  • 镜像 / head: lmsysorg/sglang:v0.5.19-cu130,提交 da239a6585566fec56766c7185ecd22e147616af(已合并 main e5f3e41c;PR 差异为该族镜像行加一条 changelog 条目)
  • 运行: 34603861241——通过 full-sweep-enabled 在 PR head 上触发的 run-sweep.yml(PR 保持草稿),Klaud 后台优先级
  • 范围: changelog 选中的完整 dsv4-fp4-b200-sglang-agentic-hicache-mtp 族——12 个 AgentX 点位(TP8 无卸载 c1–5;TP8 HiCache c8/10/16;TP8/EP8/DP-attention HiCache c64/96/128/160 配 sglang-router 0.3.2)加 c160 的默认 gsm8k 评测,全部在 cluster:b200-nscale 上。未做任何裁剪。
  • 状态: 运行中;结果、相对 2026-09-08 已发布基线的逐点差异及评测分数将补充于此。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Final full sweep · Passed · Run 34603861241 / attempt 1 · 2026-09-11 20:34 UTC
lmsysorg/sglang:v0.5.19-cu130 · da239a658556 · Mean latency
Change: Validate the complete updated-image family at this head.

Point Output tok/s/GPU ↑ TTFT ms ↓ TPOT ms ↓
AgentX c1 TP8 EP1 17a240 17.59 (N/A) 1,012.58 (N/A) 3.4 (N/A)
AgentX c2 TP8 EP1 aef2ed 20.94 (N/A) 801.37 (N/A) 4.69 (N/A)
AgentX c3 TP8 EP1 61e4f8 25.29 (N/A) 670.71 (N/A) 3.71 (N/A)
AgentX c4 TP8 EP1 9268fa 25.8 (N/A) 728.41 (N/A) 3.81 (N/A)
AgentX c5 TP8 EP1 bc37f0 32.92 (N/A) 709.54 (N/A) 4.06 (N/A)
AgentX c8 TP8 EP1 569356 49.47 (N/A) 699.78 (N/A) 4.69 (N/A)
AgentX c10 TP8 EP1 10adf0 59.01 (N/A) 711.7 (N/A) 4.97 (N/A)
AgentX c16 TP8 EP1 0b1396 81.33 (N/A) 881.78 (N/A) 7.26 (N/A)
AgentX c64 TP8 EP8 797cc6 283.17 (N/A) 3,355.35 (N/A) 12.52 (N/A)
AgentX c96 TP8 EP8 3c4858 390.15 (N/A) 6,694 (N/A) 14.69 (N/A)
AgentX c128 TP8 EP8 4f7ef3 425.95 (N/A) 10,976.66 (N/A) 15.86 (N/A)
AgentX c160 TP8 EP8 310f39 423.34 (N/A) 20,709.21 (N/A) 18.31 (N/A)

Note: All rows: Δ N/A: no matched baseline.
Note: AgentX c3 TP8 EP1 61e4f8: 1 request error.
Note: AgentX c4 TP8 EP1 9268fa, AgentX c8 TP8 EP1 569356, AgentX c10 TP8 EP1 10adf0: 3 request errors.
Note: AgentX c16 TP8 EP1 0b1396: 2 request errors.
Note: AgentX c64 TP8 EP8 797cc6: 11 request errors.
Note: AgentX c96 TP8 EP8 3c4858: 28 request errors.
Note: AgentX c128 TP8 EP8 4f7ef3, AgentX c160 TP8 EP8 310f39: 32 request errors.

Eval Score ↑ Samples
gsm8k/em_strict · c160 96.36% (N/A) N/A/1,319 (old/new)

Note: All rows: Δ N/A: no matched eval baseline.

Next: Mark ready for maintainer review.

中文

最终完整 sweep · 已通过 · Run 34603861241 / attempt 1 · 2026-09-11 20:34 UTC
lmsysorg/sglang:v0.5.19-cu130 · da239a658556 · 平均延迟
**变更:**验证此提交更新镜像后的完整配置族。实测数值及异常说明见上表。
**下一步:**标记为就绪,等待维护者审查。

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 11, 2026 20:34
@Klaud-Cold
Klaud-Cold requested a review from a team September 11, 2026 20:34
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

validated · Repairs: unknown · Runs: 34593186565, 34603861241
All owned runs ended. Full sweep verified; ready for review.

中文

validated · 修复次数:unknown · 运行:34593186565, 34603861241
所有自有运行均已结束。完整 sweep 已验证;已就绪,等待审查。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward version-bump config change. Verified: the image bump in configs/nvidia-master.yaml (dsv4-fp4-b200-sglang-agentic-hicache-mtp) is single-node with no model.container field to keep in sync; the perf-changelog.yaml entry is appended after the last existing entry with no prior bytes touched. Checked whether the new tag should be digest-pinned like dsv4-fp4-gb300-dynamo-sglang-agentic-agg — the same bare lmsysorg/sglang:v0.5.19-cu130 tag (no digest) is already used by six other recipes in this file, so the bare tag here is consistent with the repo's dominant pattern rather than an outlier.

Extended reasoning...

Overview

The diff touches two files: a one-line image tag bump for the dsv4-fp4-b200-sglang-agentic-hicache-mtp recipe in configs/nvidia-master.yaml, and a corresponding perf-changelog.yaml entry appended at the tail. No scripts, Slurm labels, or model/search-space fields changed.

Security risks

None. This is a Docker image tag reference change and a changelog append; no code execution paths, credentials, or auth logic are touched.

Level of scrutiny

This warrants light scrutiny consistent with a mechanical version bump: confirm the changelog is append-only (verified — new entry follows the last pre-existing entry with no byte changes to prior content), confirm no model.container desync (not applicable — recipe is single-node), and check the one candidate issue already surfaced (bare vs. digest-pinned tag) against the rest of the file.

Other factors

I independently confirmed the bare-tag candidate issue that was already ruled out: lmsysorg/sglang:v0.5.19-cu130 without a digest appears in six other recipe blocks in nvidia-master.yaml, and only dsv4-fp4-gb300-dynamo-sglang-agentic-agg pins the digest — so the bare tag in this diff matches the dominant existing convention, not an anomaly specific to this PR. No third-party CHANGES_REQUESTED or unresolved objection is visible in the timeline metadata, and the bug hunter reported zero findings with a dry_streak exit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant