[None][feat] Do not REVIEW: Stack PRs for sweep perfing and accuracy checking (v6) by Wanli-Jiang · Pull Request #14356 · NVIDIA/TensorRT-LLM

Wanli-Jiang · 2026-05-20T09:12:35Z

Features:

about replay, default is OFF, while set TRTLLM_USE_MAMBA_REPLAY=1 will enable default replay behaviour. (on B200, it should be disabled)
about legacy mamba cache manager, set TRTLLM_USE_PY_MAMBA=1 to enable it to have old mamba cache manager testing. Default is off, which will use the new unifed cache manager.

Known optimal config

For perf sweeping, we can set TRTLLM_USE_PY_MAMBA=1 now, to enable max throughput.
For accuracy, we can set TRTLLM_USE_MAMBA_REPLAY=0, which can ensure MTP also with matched accuracy.

Included PRs:

[None][fix] Fix CppMambaHybridCacheManager functional and perf issues #14003
[None][feat] Refactor to support legacy and 1.x modelopt quant config format #14088
[None][feature] Add env variables to help debugging mamba modules. #14170
- add env to control replay (enable by default)
- add env to control use _legacy_py_mamba_cache_manager (not used by default)
[None][fix] ADP router crashes on serve when scheduling_params.attent… #14267
[None][fix] Hide trailing EOS from generated text #14336
A fix for chat-template with tool call parser
Disable replay algorithm.

Wanli-Jiang · 2026-05-20T09:23:40Z

/bot run --stage-list "Build-Docker-Images" --disable-fail-fast

tensorrt-cicd · 2026-05-20T09:30:04Z

PR_Github #49404 [ run ] triggered by Bot. Commit: 0368186 Link to invocation

tensorrt-cicd · 2026-05-20T17:24:25Z

PR_Github #49404 [ run ] completed with state FAILURE. Commit: 0368186
/LLM/main/L0_MergeRequest_PR pipeline #39053 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

Please check the failed tests and fix your PR
If you cannot view the failures, ask the CI triggerer to share details
Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

commit c9933a8 Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Mon May 18 17:58:20 2026 +0800 [None][fix] Scale recurrent-state pool by pp_size for hybrid Mamba models (By Agent) `_calculate_max_num_blocks_for_linear_attention` sizes the recurrent-state slot pool as `max_batch_size + 1`, ignoring `pp_size`. With pipeline parallelism (e.g. `TestNemotronV3Super::test_nvfp4_parallelism[TP4_PP2]`), multiple microbatches are in-flight on the same rank simultaneously, each holding up to `max_batch_size` sequences' Mamba state — so the live-state slot count must be `max_batch_size * pp_size + 1`. The KVCacheManagerV2 path already does this (`max_num_sequences = max_batch_size * pp_size`); this commit aligns the linear-attention sizing path with that invariant. Three call sites are fixed: 1. `KVCacheManager._calculate_max_num_blocks_for_linear_attention` — both `intercept` (memory reservation in the affine model) and `max_snapshots` (slot count) now scale with `pp_size`. 2. The `is_linear_attention` branch in the estimation dry-run code path, which previously used a bare `max_batch_size`. 3. `CppMambaHybridCacheManager.get_cache_size_per_token` — the per-request fixed-cost intercept that drives upstream `_tokens_for_budget`. Without this fix, `PP > 1` runs trip `No free block found. This shouldn't happen!` in `evictionPolicy::getFreeBlock` once requests beyond the first microbatch arrive at the pool. Added a regression unit test that asserts `recurrent_primary >= max_batch_size * pp_size + 1` for `pp_size ∈ {2, 4}`, and verified the fix end-to-end with the TP4_PP2 model run on a TP2_PP2 configuration locally (test passed in 8m53s, no `No free block` error). Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 7a6e90f Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Mon May 18 01:45:33 2026 +0800 [None][fix] Cap CppMamba warmup tokens by recurrent-state pool capacity (By Agent) `_create_cuda_graph_warmup_request` constructs the longest dummy request using `available_tokens` returned by `kv_cache_manager.get_num_available_tokens`. For `CppMambaHybridCacheManager`, the parent's implementation only bounds this by the attention KV cache, so the long dummy's `token_num` could balloon to near `max_seq_len` (262K for Qwen3.5-397B-A17B). With block reuse on and `mamba_state_cache_interval=256`, that one dummy then needs ~1024 recurrent-state blocks, exhausting the RS pool when stacked with the (batch_size - 1) short dummies in the same warmup batch and tripping `No free block found` in `evictionPolicy::getFreeBlock`. Override `get_num_available_tokens` in CppMambaHybridCacheManager to also clamp by `(rs_free_blocks - 1) * states_snapshot_interval` whenever the snapshot interval is positive, leaving one block of headroom for the always-allocated trailing snapshot. The override is a no-op when `enable_block_reuse=False` (interval == 0). Verified with `tests/integration/defs/accuracy/test_llm_api_pytorch.py::TestQwen3_5_397B_A17B::test_nvfp4[tep4_block_reuse]`. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 717edeb Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Sat May 16 16:27:02 2026 +0800 fix schedule issue, improve runtime perf, allow change manager Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 1813212 Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Fri May 15 08:47:09 2026 +0800 [None][fix] Address recurrent-state pool sizing and layer-first transfer bugs (By Agent) Concern 1 (kvCacheTransferManager.cpp): In the layer-first cudaMemcpy2DAsync path, source and destination pitches were both derived from pool.primaryPtr. For host-offload transfers, the secondary pool has mNumSecondaryBlocks which can differ from mNumPrimaryBlocks, giving a different per-layer stride. Fix: compute srcLayerStrideBytes and dstLayerStrideBytes independently from srcPool and dstPool, and use them as separate pitches in cudaMemcpy2DAsync. Concern 2 (resource_manager.py): In _calculate_max_num_blocks_for_linear_attention, the block-reuse branch unconditionally overwrote max_snapshots with max_tokens // interval, discarding the live-state + CUDA-graph-padding floor (max_batch_size + 1 + max_draft_len) set earlier. Fix: use max(max_tokens // interval, max_snapshots) to preserve the floor. Add a regression test (enable_block_reuse=True) that would have caught the dropped floor: with max_batch_size=4 and max_tokens=512/interval=256, the naive branch returns 2 slots (< required floor of 5). Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit e00c230 Merge: 69899c5 be4f57e Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Wed May 13 18:12:00 2026 +0800 Merge branch 'main' into user/xiweny/mamba-recurrent-cuda-graph-snapshots Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 69899c5 Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Wed May 13 18:01:04 2026 +0800 fix replay check on sm120 Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 3f0e36c Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Wed May 13 17:57:48 2026 +0800 fix missing is_draft Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 516ea84 Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Tue May 12 23:16:54 2026 +0800 reduce sync overhead Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit ace03a2 Merge: 5a85831 5aad39a Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Tue May 12 21:27:53 2026 +0800 Merge branch 'user/xiweny/mamba-hybrid-zero-mamba-layers' into mamba-recurrent-cuda-graph-snapshots (By Agent) Brings in the CppMambaHybridCacheManager zero-local-mamba-layers fix and its CodeRabbit follow-ups (PR NVIDIA#13999). Conflict in tests/unittest/_torch/executor/test_mamba_cache_manager.py is resolved by keeping both test sections (recurrent-state pool sizing from this branch, zero-local-mamba early-exit from the merged branch) and unioning the import set. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 5aad39a Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Tue May 12 11:15:21 2026 +0800 [None][fix] CppMambaHybridCacheManager: address CodeRabbit review (By Agent) Two fixes from CodeRabbit's review on NVIDIA#13999: 1. Handle layer_mask=None defensively. The signature has always typed it as Optional[List[bool]], and the new constructor was unconditionally calling layer_mask.copy() / indexing before the zero-mamba guard ran. Treat None as an all-False full-attention mask and validate length against mamba_layer_mask. 2. Seed externally visible mamba fields before the early-return guard. get_mamba_ssm_cache_dtype() reads self.ssm_state_dtype, and the use_replay_state_update property reads self._use_replay_state_update; both were only assigned on the non-early-exit path, so any caller that hit those accessors on a zero-local-mamba rank would have AttributeError'd. Move both assignments above the early-return. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 5a85831 Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Tue May 12 09:41:25 2026 +0800 [None][fix] KVCacheManager: store spec_config on self in __init__ (By Agent) Address CodeRabbit review on NVIDIA#14003: the linear-attention sizing branch reads self.spec_config, but KVCacheManager.__init__ never assigned the incoming spec_config argument to self. The code worked only by accident, because CppMambaHybridCacheManager's _setup_mtp_intermediate_states runs before super().__init__ and writes self.spec_config on the partially initialized instance. Any other KVCacheManager path that exercises the linear-attention branch with a spec_config would AttributeError. Assign self.spec_config = spec_config at the top of KVCacheManager.__init__ so the attribute is always present regardless of subclass. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 3310de0 Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Mon May 11 23:46:42 2026 +0800 [None][fix] KVCacheManager: fix max_draft_tokens typo and add slot reservation tests (By Agent) Follow-up to the recurrent-state slot reservation: spec_config has no max_draft_tokens attribute (DecodingBaseConfig uses StrictBaseModel with extra="forbid"). The correct field is max_draft_len, which is also the count of additional dummy request ids issued by CUDAGraphRunner._get_padded_batch. Add two unit tests that construct a real CppMambaHybridCacheManager and inspect blocks_per_window[RECURRENT_STATES]: - without spec_config: pool reserves max_batch_size + 1 slots - with spec_config(max_draft_len=N): pool reserves max_batch_size + 1 + N Tests disable block reuse so the override (max_tokens // interval) does not mask the addition under test. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit 992ec1b Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Mon May 11 23:34:04 2026 +0800 [None][fix] KVCacheManager: reserve recurrent-state slots for CUDA graph padding and draft-len sentinels (By Agent) The recurrent-state snapshot pool was sized at exactly max_batch_size, which leaves no room for the CUDA graph padding sentinel (CUDA_GRAPH_DUMMY_REQUEST_ID) and the per-draft-len sentinels that CUDAGraphRunner::_get_padded_batch issues during speculative decoding. Reserve one extra slot for the CUDA graph padding dummy, and when a spec_config is present, reserve one additional slot per draft length so all distinct sentinel request IDs can coexist without evicting live recurrent state. The block-reuse path (max_tokens // interval) still overrides this when active. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> commit ac30ded Author: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Date: Mon May 11 20:26:14 2026 +0800 [None][fix] CppMambaHybridCacheManager: handle ranks with zero local mamba layers (By Agent) Under PP sharding, a rank can end up with only full-attention layers and no mamba layers. The previous constructor unconditionally allocated SSM / conv state and set up linear-attention metadata, leading to invalid state on these ranks. Take an early-exit after super().__init__ when local_num_mamba_layers == 0: forward num_layers (not mamba_num_layers + num_layers) and the union of mamba+attention layer masks to the parent KVCacheManager, skip all mamba-only setup. Guard prepare_resources / update_mamba_states / _setup_state_indices with the same condition so they no-op on these ranks. Add a regression test that constructs the manager with a real parent KVCacheManager and verifies: early-exit invariants on mgr state, no mamba-only attributes set, and the three guarded methods no-op without raising. Signed-off-by: Xiwen Yu <13230610+VALLIS-NERIA@users.noreply.github.com> Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

commit 1d76cf7 Author: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com> Date: Wed May 13 01:57:44 2026 -0700 [None][feat] Refactor to support legacy and 1.x modelopt quant config format Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com> Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

…ion_dp_relax is None Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

* only with issue with transformers 4.x Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

Wanli-Jiang · 2026-05-21T04:18:18Z

/bot run --stage-list "Build-Docker-Images" --disable-fail-fast

tensorrt-cicd · 2026-05-21T04:24:34Z

PR_Github #49585 [ run ] triggered by Bot. Commit: c6c611d Link to invocation

tensorrt-cicd · 2026-05-21T11:19:22Z

PR_Github #49585 [ run ] completed with state FAILURE. Commit: c6c611d
/LLM/main/L0_MergeRequest_PR pipeline #39210 (Partly Tested) completed with status: 'FAILURE'

CI Report

⚠️ Action Required:

Please check the failed tests and fix your PR
If you cannot view the failures, ask the CI triggerer to share details
Once fixed, request an NVIDIA team member to trigger CI again

CI Agent Failure Analysis

Link to invocation

Wanli-Jiang requested review from a team as code owners May 20, 2026 09:12

Wanli-Jiang requested review from bmarimuthu-nv, brb-nv, byshiue, schetlur-nv, shaharmor98 and suyoggupta May 20, 2026 09:12

github-actions Bot assigned Wanli-Jiang May 20, 2026

Wanli-Jiang changed the title ~~[None][feat] Stack PRs for sweep perfing and accuracy checking (v6)~~ [None][feat] Do not REVIEW: Stack PRs for sweep perfing and accuracy checking (v6) May 20, 2026

Wanli-Jiang removed request for bmarimuthu-nv, brb-nv, byshiue, schetlur-nv, shaharmor98 and suyoggupta May 20, 2026 09:23

This comment was marked as off-topic.

Sign in to view

Wanli-Jiang and others added 4 commits May 20, 2026 21:09

[None][feature] Add env variables to help debugging mamba modules

ddb8a79

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

[None][fix] ADP router crashes on serve when scheduling_params.attent…

521c8c6

…ion_dp_relax is None Signed-off-by: nv-guomingz <137257613+nv-guomingz@users.noreply.github.com>

Wanli-Jiang added 3 commits May 20, 2026 21:09

[None][fix] Hide trailing EOS from generated text

c5b2ac7

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

Fix chat-template missing bug

2b44d0d

* only with issue with transformers 4.x Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

Disable replay by default

c6c611d

Signed-off-by: Wanli Jiang <35160485+Wanli-Jiang@users.noreply.github.com>

Wanli-Jiang force-pushed the user/williamj/stack-pr-v6 branch from 0368186 to c6c611d Compare May 21, 2026 04:15

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[None][feat] Do not REVIEW: Stack PRs for sweep perfing and accuracy checking (v6)#14356

[None][feat] Do not REVIEW: Stack PRs for sweep perfing and accuracy checking (v6)#14356
Wanli-Jiang wants to merge 7 commits into
NVIDIA:mainfrom
Wanli-Jiang:user/williamj/stack-pr-v6

Wanli-Jiang commented May 20, 2026 •

edited

Loading

Uh oh!

Wanli-Jiang commented May 20, 2026

Uh oh!

This comment was marked as off-topic.

This comment was marked as off-topic.

Uh oh!

tensorrt-cicd commented May 20, 2026

Uh oh!

tensorrt-cicd commented May 20, 2026

Uh oh!

Wanli-Jiang commented May 21, 2026

Uh oh!

tensorrt-cicd commented May 21, 2026

Uh oh!

tensorrt-cicd commented May 21, 2026

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Conversation

Wanli-Jiang commented May 20, 2026 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Features:

Known optimal config

Included PRs:

Uh oh!

Wanli-Jiang commented May 20, 2026

Uh oh!

This comment was marked as off-topic.

This comment was marked as off-topic.

Uh oh!

tensorrt-cicd commented May 20, 2026

Uh oh!

tensorrt-cicd commented May 20, 2026

Uh oh!

Wanli-Jiang commented May 21, 2026

Uh oh!

tensorrt-cicd commented May 21, 2026

Uh oh!

tensorrt-cicd commented May 21, 2026

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

Wanli-Jiang commented May 20, 2026 •

edited

Loading