Skip to content

Merge vng-benchmark into main#29

Merged
aistackdev merged 25 commits into
mainfrom
vng-benchmark
Jul 24, 2026
Merged

Merge vng-benchmark into main#29
aistackdev merged 25 commits into
mainfrom
vng-benchmark

Conversation

@aistackdev

Copy link
Copy Markdown

Summary

Brings 25 commits from vng-benchmark into main, most recently PR #27 (remote-bench / BYO-endpoint support). This is required so remote-bench.yml's workflow_dispatch trigger becomes usable — GH Actions only allows dispatching a workflow that exists on the repo's default branch (main), and vng-benchmark has never been merged there before.

Fast-forwardable: 25 ahead, 0 behind main.

Test plan

Thắng. Lý Quang (5) and others added 25 commits July 22, 2026 13:17
Prepare FP8 and W4AFP8 single-node AgentX smoke configurations with server/DCGM telemetry. The initial dispatch selects only the locally cached FP8 model.\n\n中文:添加 GreenNode GLM-5.2 FP8 与 W4AFP8 单节点 AgentX 冒烟基准配置,并接入服务端指标与 DCGM 遥测。首次调度仅选择已在本地缓存的 FP8 模型。
Configure the GreenNode H200 GLM-5.2 FP8 agentic benchmark to run only at CCU 8.\n\n中文:将 GreenNode H200 GLM-5.2 FP8 智能体基准测试配置为仅运行 CCU 8。
Pass KV_OFFLOAD_BACKEND_METADATA into the benchmark container so agentic result aggregation can validate the HiCache backend. Add a regression check for the launcher environment.\n\n中文:将 KV_OFFLOAD_BACKEND_METADATA 传入基准测试容器,使智能体结果聚合能够校验 HiCache 后端;并新增启动器环境变量回归测试。
Add a guarded workflow for configuring, validating, previewing, and dispatching single-node agentic benchmarks.\n\n中文:新增项目级 agentic 基准测试调度 skill,支持配置收集、预检、预览确认及安全调度。
Add a runner-specific GPU-resident KV smoke configuration while preserving sglang-router session affinity. Enable DCGM and server metrics, and validate local weights inside the benchmark container.

中文:新增面向 B300 netperf 节点的 GLM-5.2 GPU 驻留 KV 冒烟测试配置,并保留 sglang-router 会话亲和路由。启用 DCGM 和服务端指标,并在基准测试容器内校验本地权重。
Configure the dedicated B300 GLM-5.2 NVFP4 agentic benchmark with HiCache ratio 1.0 and CCU 8/32/48/64. Use exact config-key generation to prevent unrelated matrix rows from being dispatched.\n\n中文:为专用 B300 GLM-5.2 NVFP4 智能体基准测试配置 HiCache ratio 1.0 及 CCU 8/32/48/64,并使用精确配置键生成,避免误调度无关矩阵任务。
Add a CCU16 agentic smoke configuration with DP8 attention, DeepEP, consistent-hash routing, GPU-resident KV, and mandatory telemetry.

中文:新增 GLM-5.2 H200 DP-attention 冒烟基准测试,使用 CCU16、DP8 attention、DeepEP、一致性哈希路由、GPU 驻留 KV 和完整遥测。
Add 20-minute CCU 8/12/16/32 agentic sweeps for HiCache ratios 0.5, 0.75, 1.5, and 2.0. Propagate the ratio through the generated matrix and reusable benchmark workflow into the runner recipe.\n\n中文:新增 B300 GLM-5.2 FP8 EAGLE 智能体基准测试,针对 HiCache ratio 0.5、0.75、1.5 和 2.0,按 CCU 8/12/16/32 各运行 20 分钟;并将 ratio 从生成矩阵经可复用工作流传递至运行器脚本。
Use the default MoE path for the H200 DP-attention smoke because DeepEP intranode combine rejects the W4AFP8 activation dtype.

中文:GLM-5.2 W4AFP8 H200 DP-attention 冒烟测试改用默认 MOE 路径,避免 DeepEP 节点内 combine 不支持该激活数据类型而崩溃。
Configure the 256K agentic sweep at CCU 8, 16, 32, 48, and 64 with HiCache ratio 2.0, SGLang DP-aware routing, server metrics, and DCGM telemetry. Forward HICACHE_RATIO through the GreenNode launcher while preserving the existing 128 GB fallback.\n\n中文:新增 H200 GLM-5.2 DPA HiCache 基准测试扫描,使用 256K Agentic 数据集、CCU 8/16/32/48/64、HiCache ratio 2.0、SGLang DP 感知路由、服务端指标和 DCGM 遥测。GreenNode 启动器新增 HICACHE_RATIO 透传,并保留现有的 128 GB 回退配置。
Extend the H200 GreenNode GLM-5.2 FP8 agentic full sweep to CCU 4 and 8 on the 256K-capped dataset.\n\n中文:将 H200 GreenNode GLM-5.2 FP8 Agentic 完整基准测试扩展为 CCU 4 和 8,并使用 256K 上下文上限数据集。
Switch the GLM-5.2 H200 DP-aware router from consistent hashing to cache-aware routing, remove the AIPerf session routing key, and narrow the matrix to CCU 32 for an apples-to-apples comparison.

中文:将 GLM-5.2 H200 的 DP 感知路由从一致性哈希切换为缓存感知路由,移除 AIPerf 会话路由键,并将基准测试矩阵收窄至 CCU 32,以进行同条件对比。
Add the H200 TP8/DP8 DeepEP agentic configuration at CCU 16 and allow an explicit user-authorized model verification override in the dispatch preflight.\n\n中文:新增 H200 TP8/DP8 DeepEP 的 GLM-5.2 FP8 智能体基准测试配置,使用 CCU 16;同时在调度预检中支持用户明确授权后跳过模型验证。
Run the H200 FP8 DP-attention smoke configuration with TP8, DP4, and DeepEP while keeping CCU 16 and all other settings unchanged.\n\n中文:将 H200 FP8 DP Attention 冒烟配置调整为 TP8、DP4 和 DeepEP,保持 CCU 16 及其他配置不变。
Enable EAGLE speculative decoding for the GLM-5.2 W4AFP8 DP-attention cache-aware agentic benchmark and run the full 256k trace at CCU 16 and 48.\n\n中文:为 GLM-5.2 W4AFP8 的 DP Attention 缓存感知 Agentic 基准测试启用 EAGLE 投机解码,并在 256k 数据集上以 CCU 16 和 48 运行完整测试。
Add the SGLang cache-aware DP router, set the DP4 ladder to CCU 8/12/20/32, and use 32k chunked prefill.

中文:添加 SGLang 缓存感知 DP 路由,将 DP4 并发扫描设置为 CCU 8/12/20/32,并使用 32k 分块预填充。
…ndpoint)

Adds a REMOTE_BASE_URL override to the three hardcoded http://0.0.0.0:$port
paths in benchmark_lib.sh (run_benchmark_serving, run_lm_eval,
build_replay_cmd), a new hw-agnostic *-remote-bench.sh recipe convention
that benchmarks an externally-managed endpoint instead of launching a local
server, and the runner/workflow plumbing to dispatch it (cluster:remote-bench
label, launch_bench-client.sh, remote-bench.yml workflow_dispatch entrypoint).
All additive; no existing recipe, workflow input, or config schema changes
behavior when unset.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Smoke-tested against a live sglang-vanilla endpoint and hit a
reproducible silent 100%-GPU hang, always on the same trace, always
after a partial chunked-prefill chunk. Root cause: without a client-side
context cap, aiperf replayed trace turns up to 255,999 tokens against a
model deployed with only a 131,072-token context, relying entirely on
server-side --allow-auto-truncate - which triggers the hang.

Capping via --max-context-length (aiperf already supports this; local
recipes never needed it since they hardcode --context-length for the
server they launch themselves) avoids the pathological case entirely -
confirmed with a full successful end-to-end run afterward. Required,
not optional, since every remote target has a real context limit an
operator can self-report, same as the other required remote-bench
inputs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Define configurable coding, short-chat, and RAG token mixes with pinned sources, timing semantics, cache measurement, and bilingual delivery constraints.

中文:设计 AgentX 混合工作负载,定义可配置的编程、短对话与 RAG token 比例,并固定数据源、时间语义、缓存测量方式和中英文交付约束。
Retarget the h200-greennode DPA8 router HiCache r2 configs for full 3600s
agentic runs: existing mtp/EAGLE key -> CCU [1,8,32], add a no-mtp sibling
key -> CCU [1,4]. Note the h200-greennode SSH target in the dispatch skill
for model verification.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Allow failed-only CCU subsets in agentic retry preflight and floor max-running-requests at the DP size.

中文:修复数据并行注意力模式下请求槽位为零的问题,并允许 Agentic 重试预检仅选择失败的 CCU 子集。
Documents the full workflow for bringing up a new remote-bench recipe:
required env-var contract, why config behind the endpoint still needs
self-reporting even though InferenceX doesn't control it, what file to
write (one per model+precision+framework, not per hardware - the body
is server-agnostic), the debug loop for pre-merge validation, and what
a real ingest-able run produces.

Also fixes remote-bench.yml: `image` was hardcoded to a placeholder
string that flows straight into the ingested artifact's `image` field;
`dp-attn` was hardcoded false. Both are now real, self-reported inputs.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
feat: support benchmarking remote/existing inference endpoints (BYO endpoint)
@github-actions

Copy link
Copy Markdown

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@aistackdev
aistackdev merged commit 305a626 into main Jul 24, 2026
3 checks passed
aistackdev pushed a commit that referenced this pull request Jul 24, 2026
…cident

Concrete rule: cherry-pick exact remote-bench commits onto main, never
merge the whole vng-benchmark branch (PR #29 did this by accident,
pulling in 21 unrelated commits, and had to be cleaned up after the fact).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant