GLM-5.3-Flash on 8x B200: dual TP4/EP4 serving profile with long-context measurements #37153
stewtong
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
I measured an SGLang serving profile for
zai-org/GLM-5.3-Flashon one 8x NVIDIA B200 node. It uses two independent TP4/EP4 replicas, FP8 KV cache, TRT-LLM DSA backends, the FlashInfer TRT-LLM MoE runner, static NEXTN 5-1-6, and session-affine routing.The pinned software stack, engine commands, proxy configuration, client setup, methods, and results are documented here: https://github.com/stewtong/b200-glm53
Validation observations:
reasoning_effort: low. The conditions used different output limits and are reported separately in the repository.The serving table and cache study used three repetitions per condition. The custom harness, prompt corpus, and raw traces are unpublished, so exact numerical reproduction requires an equivalent harness and prompt set. The repository documents the tested boundaries and excludes unsupported aggregate throughput, decode TPOT, route-latency, and causal claims.
All reactions