Merge vng-benchmark into main#29
Conversation
Prepare FP8 and W4AFP8 single-node AgentX smoke configurations with server/DCGM telemetry. The initial dispatch selects only the locally cached FP8 model.\n\n中文:添加 GreenNode GLM-5.2 FP8 与 W4AFP8 单节点 AgentX 冒烟基准配置,并接入服务端指标与 DCGM 遥测。首次调度仅选择已在本地缓存的 FP8 模型。
Configure the GreenNode H200 GLM-5.2 FP8 agentic benchmark to run only at CCU 8.\n\n中文:将 GreenNode H200 GLM-5.2 FP8 智能体基准测试配置为仅运行 CCU 8。
Pass KV_OFFLOAD_BACKEND_METADATA into the benchmark container so agentic result aggregation can validate the HiCache backend. Add a regression check for the launcher environment.\n\n中文:将 KV_OFFLOAD_BACKEND_METADATA 传入基准测试容器,使智能体结果聚合能够校验 HiCache 后端;并新增启动器环境变量回归测试。
Add a guarded workflow for configuring, validating, previewing, and dispatching single-node agentic benchmarks.\n\n中文:新增项目级 agentic 基准测试调度 skill,支持配置收集、预检、预览确认及安全调度。
Add a runner-specific GPU-resident KV smoke configuration while preserving sglang-router session affinity. Enable DCGM and server metrics, and validate local weights inside the benchmark container. 中文:新增面向 B300 netperf 节点的 GLM-5.2 GPU 驻留 KV 冒烟测试配置,并保留 sglang-router 会话亲和路由。启用 DCGM 和服务端指标,并在基准测试容器内校验本地权重。
Configure the dedicated B300 GLM-5.2 NVFP4 agentic benchmark with HiCache ratio 1.0 and CCU 8/32/48/64. Use exact config-key generation to prevent unrelated matrix rows from being dispatched.\n\n中文:为专用 B300 GLM-5.2 NVFP4 智能体基准测试配置 HiCache ratio 1.0 及 CCU 8/32/48/64,并使用精确配置键生成,避免误调度无关矩阵任务。
Add a CCU16 agentic smoke configuration with DP8 attention, DeepEP, consistent-hash routing, GPU-resident KV, and mandatory telemetry. 中文:新增 GLM-5.2 H200 DP-attention 冒烟基准测试,使用 CCU16、DP8 attention、DeepEP、一致性哈希路由、GPU 驻留 KV 和完整遥测。
Add 20-minute CCU 8/12/16/32 agentic sweeps for HiCache ratios 0.5, 0.75, 1.5, and 2.0. Propagate the ratio through the generated matrix and reusable benchmark workflow into the runner recipe.\n\n中文:新增 B300 GLM-5.2 FP8 EAGLE 智能体基准测试,针对 HiCache ratio 0.5、0.75、1.5 和 2.0,按 CCU 8/12/16/32 各运行 20 分钟;并将 ratio 从生成矩阵经可复用工作流传递至运行器脚本。
Use the default MoE path for the H200 DP-attention smoke because DeepEP intranode combine rejects the W4AFP8 activation dtype. 中文:GLM-5.2 W4AFP8 H200 DP-attention 冒烟测试改用默认 MOE 路径,避免 DeepEP 节点内 combine 不支持该激活数据类型而崩溃。
Configure the 256K agentic sweep at CCU 8, 16, 32, 48, and 64 with HiCache ratio 2.0, SGLang DP-aware routing, server metrics, and DCGM telemetry. Forward HICACHE_RATIO through the GreenNode launcher while preserving the existing 128 GB fallback.\n\n中文:新增 H200 GLM-5.2 DPA HiCache 基准测试扫描,使用 256K Agentic 数据集、CCU 8/16/32/48/64、HiCache ratio 2.0、SGLang DP 感知路由、服务端指标和 DCGM 遥测。GreenNode 启动器新增 HICACHE_RATIO 透传,并保留现有的 128 GB 回退配置。
Extend the H200 GreenNode GLM-5.2 FP8 agentic full sweep to CCU 4 and 8 on the 256K-capped dataset.\n\n中文:将 H200 GreenNode GLM-5.2 FP8 Agentic 完整基准测试扩展为 CCU 4 和 8,并使用 256K 上下文上限数据集。
Switch the GLM-5.2 H200 DP-aware router from consistent hashing to cache-aware routing, remove the AIPerf session routing key, and narrow the matrix to CCU 32 for an apples-to-apples comparison. 中文:将 GLM-5.2 H200 的 DP 感知路由从一致性哈希切换为缓存感知路由,移除 AIPerf 会话路由键,并将基准测试矩阵收窄至 CCU 32,以进行同条件对比。
Add the H200 TP8/DP8 DeepEP agentic configuration at CCU 16 and allow an explicit user-authorized model verification override in the dispatch preflight.\n\n中文:新增 H200 TP8/DP8 DeepEP 的 GLM-5.2 FP8 智能体基准测试配置,使用 CCU 16;同时在调度预检中支持用户明确授权后跳过模型验证。
Run the H200 FP8 DP-attention smoke configuration with TP8, DP4, and DeepEP while keeping CCU 16 and all other settings unchanged.\n\n中文:将 H200 FP8 DP Attention 冒烟配置调整为 TP8、DP4 和 DeepEP,保持 CCU 16 及其他配置不变。
Enable EAGLE speculative decoding for the GLM-5.2 W4AFP8 DP-attention cache-aware agentic benchmark and run the full 256k trace at CCU 16 and 48.\n\n中文:为 GLM-5.2 W4AFP8 的 DP Attention 缓存感知 Agentic 基准测试启用 EAGLE 投机解码,并在 256k 数据集上以 CCU 16 和 48 运行完整测试。
Add the SGLang cache-aware DP router, set the DP4 ladder to CCU 8/12/20/32, and use 32k chunked prefill. 中文:添加 SGLang 缓存感知 DP 路由,将 DP4 并发扫描设置为 CCU 8/12/20/32,并使用 32k 分块预填充。
…ndpoint) Adds a REMOTE_BASE_URL override to the three hardcoded http://0.0.0.0:$port paths in benchmark_lib.sh (run_benchmark_serving, run_lm_eval, build_replay_cmd), a new hw-agnostic *-remote-bench.sh recipe convention that benchmarks an externally-managed endpoint instead of launching a local server, and the runner/workflow plumbing to dispatch it (cluster:remote-bench label, launch_bench-client.sh, remote-bench.yml workflow_dispatch entrypoint). All additive; no existing recipe, workflow input, or config schema changes behavior when unset. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Smoke-tested against a live sglang-vanilla endpoint and hit a reproducible silent 100%-GPU hang, always on the same trace, always after a partial chunked-prefill chunk. Root cause: without a client-side context cap, aiperf replayed trace turns up to 255,999 tokens against a model deployed with only a 131,072-token context, relying entirely on server-side --allow-auto-truncate - which triggers the hang. Capping via --max-context-length (aiperf already supports this; local recipes never needed it since they hardcode --context-length for the server they launch themselves) avoids the pathological case entirely - confirmed with a full successful end-to-end run afterward. Required, not optional, since every remote target has a real context limit an operator can self-report, same as the other required remote-bench inputs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Define configurable coding, short-chat, and RAG token mixes with pinned sources, timing semantics, cache measurement, and bilingual delivery constraints. 中文:设计 AgentX 混合工作负载,定义可配置的编程、短对话与 RAG token 比例,并固定数据源、时间语义、缓存测量方式和中英文交付约束。
Retarget the h200-greennode DPA8 router HiCache r2 configs for full 3600s agentic runs: existing mtp/EAGLE key -> CCU [1,8,32], add a no-mtp sibling key -> CCU [1,4]. Note the h200-greennode SSH target in the dispatch skill for model verification. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Allow failed-only CCU subsets in agentic retry preflight and floor max-running-requests at the DP size. 中文:修复数据并行注意力模式下请求槽位为零的问题,并允许 Agentic 重试预检仅选择失败的 CCU 子集。
Documents the full workflow for bringing up a new remote-bench recipe: required env-var contract, why config behind the endpoint still needs self-reporting even though InferenceX doesn't control it, what file to write (one per model+precision+framework, not per hardware - the body is server-agnostic), the debug loop for pre-merge validation, and what a real ingest-able run produces. Also fixes remote-bench.yml: `image` was hardcoded to a placeholder string that flows straight into the ingested artifact's `image` field; `dp-attn` was hardcoded false. Both are now real, self-reported inputs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
feat: support benchmarking remote/existing inference endpoints (BYO endpoint)
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
…cident Concrete rule: cherry-pick exact remote-bench commits onto main, never merge the whole vng-benchmark branch (PR #29 did this by accident, pulling in 21 unrelated commits, and had to be cleaned up after the fact).
Summary
Brings 25 commits from
vng-benchmarkintomain, most recently PR #27 (remote-bench / BYO-endpoint support). This is required soremote-bench.yml'sworkflow_dispatchtrigger becomes usable — GH Actions only allows dispatching a workflow that exists on the repo's default branch (main), andvng-benchmarkhas never been merged there before.Fast-forwardable: 25 ahead, 0 behind
main.Test plan
gh workflow run remote-bench.yml -R vngcloud/InferenceXno longer 404s