Checklist
Describe the bug
HiCache Causes Illegal Memory Access
how to reproduce
Model
export MODEL_PATH=qwen/qwen3-14b
client
python3 benchmark/hicache/bench_multiturn.py --model-path $MODEL_PATH --port 30000 --disable-random-sample --output-length 16 --request-length 2048 --num-clients 20 --num-rounds 10 --max-parallel 4 --request-rate 20 --ready-queue-policy random --disable-auto-run --seed 42
server
python3 -m sglang.launch_server --model-path $MODEL_PATH --tp-size 2 --page-size 64 --enable-hierarchical-cache --hicache-write-policy write_through --hicache-storage-backend file --hicache-ratio 2 --hicache-size 0
Problem disappears when not enabling HiCache
Run the server without HiCache, inference is successful:
python3 -m sglang.launch_server --model-path $MODEL_PATH --tp-size 2 --page-size 64
Error Message
torch.AcceleratorError: CUDA error: an illegal memory access was encountered
Search for `cudaErrorIllegalAddress' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
terminate called after throwing an instance of 'c10::AcceleratorError'
[2026-01-08 18:34:47] SIGQUIT received. signum=None, frame=None. It usually means one child failed.
what(): CUDA error: an illegal memory access was encountered
Search for `cudaErrorIllegalAddress' in https://docs.nvidia.com/cuda/cuda-runtime-api/group__CUDART__TYPES.html for more information.
CUDA kernel errors might be asynchronously reported at some other API call, so the stacktrace below might be incorrect.
For debugging consider passing CUDA_LAUNCH_BLOCKING=1
Compile with `TORCH_USE_CUDA_DSA` to enable device-side assertions.
Exception raised from c10_cuda_check_implementation at /pytorch/c10/cuda/CUDAException.cpp:44 (most recent call first):
Initial Investigation
By iterating the commits in the main branch, it seem that the commit 4935344 ([AMD] Fix aiter page-size handling, DeepSeek MLA tuple inputs, and HiCache/FA3 decode-backend override (#16531)) introduces some bugs which lead to the above error.
git diff 63cc97f4e 4935344fc
diff --git a/python/sglang/srt/layers/attention/aiter_backend.py b/python/sglang/srt/layers/attention/aiter_backend.py
index a2c4a2e19..e71c051a4 100644
--- a/python/sglang/srt/layers/attention/aiter_backend.py
+++ b/python/sglang/srt/layers/attention/aiter_backend.py
@@ -279,7 +279,7 @@ class AiterAttnBackend(AttentionBackend):
):
nhead_kv = 1
- page_size = 1
+ page_size = self.page_size
dtype = self.kv_cache_dtype
meta = get_mla_metadata_v1(
@@ -1654,7 +1654,6 @@ class AiterMultiStepDraftBackend:
# Cached variables for generate_draft_decode_kv_indices
self.pool_len = model_runner.req_to_token_pool.req_to_token.shape[1]
self.page_size = model_runner.server_args.page_size
- assert self.page_size == 1, "Page size must be 1"
def common_template(
self, forward_batch: ForwardBatch, kv_indices_buffer: torch.Tensor, call_fn: int
diff --git a/python/sglang/srt/models/deepseek_v2.py b/python/sglang/srt/models/deepseek_v2.py
index 81a2058f5..96bb3812d 100644
--- a/python/sglang/srt/models/deepseek_v2.py
+++ b/python/sglang/srt/models/deepseek_v2.py
@@ -2036,8 +2036,15 @@ class DeepseekV2AttentionMLA(nn.Module):
enable_rope_fusion = (
os.getenv("SGLANG_FUSED_MLA_ENABLE_ROPE_FUSION", "1") == "1"
)
- q_len = hidden_states.shape[0]
- q_input = hidden_states.new_empty(
+ # NOTE: hidden_states can be a tuple for some quantization paths.
+ # For shape/device/dtype, use the first tensor; still pass the original
+ # hidden_states through linear ops which may accept tuple inputs.
+ hidden_states_tensor = (
+ hidden_states[0] if isinstance(hidden_states, tuple) else hidden_states
+ )
+
+ q_len = hidden_states_tensor.shape[0]
+ q_input = hidden_states_tensor.new_empty(
q_len, self.num_local_heads, self.kv_lora_rank + self.qk_rope_head_dim
)
if self.q_lora_rank is not None:
diff --git a/python/sglang/srt/server_args.py b/python/sglang/srt/server_args.py
index 35ed102ed..a76486a3b 100644
--- a/python/sglang/srt/server_args.py
+++ b/python/sglang/srt/server_args.py
@@ -1993,21 +1993,31 @@ class ServerArgs:
or self.disaggregation_decode_enable_offload_kvcache
) and self.hicache_io_backend == "kernel":
# fix for the compatibility issue with FlashAttention3 decoding and HiCache kernel backend
- if self.decode_attention_backend is None:
- if not self.use_mla_backend():
- self.decode_attention_backend = (
- "flashinfer" if is_flashinfer_available() else "triton"
- )
+ # Only override when the *effective* decode backend would be FA3.
+ # Otherwise, respect the user's chosen attention backend (e.g., aiter on ROCm).
+ effective_decode_backend = (
+ self.decode_attention_backend
+ if self.decode_attention_backend is not None
+ else self.attention_backend
+ )
+ if effective_decode_backend == "fa3":
+ if self.decode_attention_backend is None:
+ # If decode backend wasn't explicitly set, pick a safe default that works with HiCache kernel IO.
+ if not self.use_mla_backend():
+ self.decode_attention_backend = (
+ "flashinfer" if is_flashinfer_available() else "triton"
+ )
+ else:
+ self.decode_attention_backend = (
+ "flashinfer" if is_sm100_supported() else "triton"
+ )
else:
- self.decode_attention_backend = (
- "flashinfer" if is_sm100_supported() else "triton"
+ # If user explicitly requested FA3 decode, fall back to direct IO.
+ self.hicache_io_backend = "direct"
+ logger.warning(
+ "FlashAttention3 decode backend is not compatible with hierarchical cache. "
+ "Setting hicache_io_backend to vanilla I/O, which may lead to suboptimal performance with small page sizes."
)
- elif self.decode_attention_backend == "fa3":
- self.hicache_io_backend = "direct"
- logger.warning(
- "FlashAttention3 decode backend is not compatible with hierarchical cache. "
- "Setting hicache_io_backend to vanilla I/O, which may lead to suboptimal performance with small page sizes."
- )
def _handle_speculative_decoding(self):
if (
Reproduction
see the description in the above summary
Environment
$ python3 -m sglang.check_env
Python: 3.12.3 (main, Aug 14 2025, 17:47:21) [GCC 13.3.0]
CUDA available: True
GPU 0,1,2,3,4,5,6,7: NVIDIA H100 80GB HBM3
GPU 0,1,2,3,4,5,6,7 Compute Capability: 9.0
CUDA_HOME: /usr/local/cuda-13
NVCC: Cuda compilation tools, release 13.0, V13.0.88
CUDA Driver Version: 580.105.08
PyTorch: 2.9.1+cu128
sglang: 0.5.6.post3.dev975+g4935344fc
sgl_kernel: 0.3.20
flashinfer_python: 0.5.3
flashinfer_cubin: 0.5.3
flashinfer_jit_cache: Module Not Found
triton: 3.5.1
transformers: 4.57.1
torchao: 0.9.0
numpy: 2.4.0
aiohttp: 3.13.3
fastapi: 0.128.0
hf_transfer: 0.1.9
huggingface_hub: 0.36.0
interegular: 0.3.3
modelscope: 1.33.0
orjson: 3.11.5
outlines: 0.1.11
packaging: 25.0
psutil: 7.2.1
pydantic: 2.12.5
python-multipart: 0.0.21
pyzmq: 27.1.0
uvicorn: 0.40.0
uvloop: 0.22.1
vllm: Module Not Found
xgrammar: 0.1.27
openai: 2.6.1
tiktoken: 0.12.0
anthropic: 0.75.0
litellm: Module Not Found
decord2: 3.0.0
NVIDIA Topology:
GPU0 GPU1 GPU2 GPU3 GPU4 GPU5 GPU6 GPU7 NIC0 NIC1 NIC2 NIC3 NIC4 NIC5 NIC6 NIC7 NIC8 NIC9 NIC10 NIC11 CPU Affinity NUMA Affinity GPU NUMA ID
GPU0 X NV18 NV18 NV18 NV18 NV18 NV18 NV18 PXB NODE NODE NODE NODE NODE SYS SYS SYS SYS SYS SYS 0-55,112-167 0 N/A
GPU1 NV18 X NV18 NV18 NV18 NV18 NV18 NV18 NODE NODE NODE PXB NODE NODE SYS SYS SYS SYS SYS SYS 0-55,112-167 0 N/A
GPU2 NV18 NV18 X NV18 NV18 NV18 NV18 NV18 NODE NODE NODE NODE PXB NODE SYS SYS SYS SYS SYS SYS 0-55,112-167 0 N/A
GPU3 NV18 NV18 NV18 X NV18 NV18 NV18 NV18 NODE NODE NODE NODE NODE PXB SYS SYS SYS SYS SYS SYS 0-55,112-167 0 N/A
GPU4 NV18 NV18 NV18 NV18 X NV18 NV18 NV18 SYS SYS SYS SYS SYS SYS PXB NODE NODE NODE NODE NODE 56-111,168-223 1 N/A
GPU5 NV18 NV18 NV18 NV18 NV18 X NV18 NV18 SYS SYS SYS SYS SYS SYS NODE NODE NODE PXB NODE NODE 56-111,168-223 1 N/A
GPU6 NV18 NV18 NV18 NV18 NV18 NV18 X NV18 SYS SYS SYS SYS SYS SYS NODE NODE NODE NODE PXB NODE 56-111,168-223 1 N/A
GPU7 NV18 NV18 NV18 NV18 NV18 NV18 NV18 X SYS SYS SYS SYS SYS SYS NODE NODE NODE NODE NODE PXB 56-111,168-223 1 N/A
NIC0 PXB NODE NODE NODE SYS SYS SYS SYS X NODE NODE NODE NODE NODE SYS SYS SYS SYS SYS SYS
NIC1 NODE NODE NODE NODE SYS SYS SYS SYS NODE X PIX NODE NODE NODE SYS SYS SYS SYS SYS SYS
NIC2 NODE NODE NODE NODE SYS SYS SYS SYS NODE PIX X NODE NODE NODE SYS SYS SYS SYS SYS SYS
NIC3 NODE PXB NODE NODE SYS SYS SYS SYS NODE NODE NODE X NODE NODE SYS SYS SYS SYS SYS SYS
NIC4 NODE NODE PXB NODE SYS SYS SYS SYS NODE NODE NODE NODE X NODE SYS SYS SYS SYS SYS SYS
NIC5 NODE NODE NODE PXB SYS SYS SYS SYS NODE NODE NODE NODE NODE X SYS SYS SYS SYS SYS SYS
NIC6 SYS SYS SYS SYS PXB NODE NODE NODE SYS SYS SYS SYS SYS SYS X NODE NODE NODE NODE NODE
NIC7 SYS SYS SYS SYS NODE NODE NODE NODE SYS SYS SYS SYS SYS SYS NODE X PIX NODE NODE NODE
NIC8 SYS SYS SYS SYS NODE NODE NODE NODE SYS SYS SYS SYS SYS SYS NODE PIX X NODE NODE NODE
NIC9 SYS SYS SYS SYS NODE PXB NODE NODE SYS SYS SYS SYS SYS SYS NODE NODE NODE X NODE NODE
NIC10 SYS SYS SYS SYS NODE NODE PXB NODE SYS SYS SYS SYS SYS SYS NODE NODE NODE NODE X NODE
NIC11 SYS SYS SYS SYS NODE NODE NODE PXB SYS SYS SYS SYS SYS SYS NODE NODE NODE NODE NODE X
Legend:
X = Self
SYS = Connection traversing PCIe as well as the SMP interconnect between NUMA nodes (e.g., QPI/UPI)
NODE = Connection traversing PCIe as well as the interconnect between PCIe Host Bridges within a NUMA node
PHB = Connection traversing PCIe as well as a PCIe Host Bridge (typically the CPU)
PXB = Connection traversing multiple PCIe bridges (without traversing the PCIe Host Bridge)
PIX = Connection traversing at most a single PCIe bridge
NV# = Connection traversing a bonded set of # NVLinks
NIC Legend:
NIC0: mlx5_0
NIC1: mlx5_1
NIC2: mlx5_2
NIC3: mlx5_3
NIC4: mlx5_4
NIC5: mlx5_5
NIC6: mlx5_6
NIC7: mlx5_7
NIC8: mlx5_8
NIC9: mlx5_9
NIC10: mlx5_10
NIC11: mlx5_11
ulimit soft: 500000
Checklist
Describe the bug
HiCache Causes Illegal Memory Access
how to reproduce
Model
export MODEL_PATH=qwen/qwen3-14bclient
python3 benchmark/hicache/bench_multiturn.py --model-path $MODEL_PATH --port 30000 --disable-random-sample --output-length 16 --request-length 2048 --num-clients 20 --num-rounds 10 --max-parallel 4 --request-rate 20 --ready-queue-policy random --disable-auto-run --seed 42server
python3 -m sglang.launch_server --model-path $MODEL_PATH --tp-size 2 --page-size 64 --enable-hierarchical-cache --hicache-write-policy write_through --hicache-storage-backend file --hicache-ratio 2 --hicache-size 0Problem disappears when not enabling HiCache
Run the server without HiCache, inference is successful:
python3 -m sglang.launch_server --model-path $MODEL_PATH --tp-size 2 --page-size 64Error Message
Initial Investigation
By iterating the commits in the main branch, it seem that the commit 4935344 ([AMD] Fix aiter page-size handling, DeepSeek MLA tuple inputs, and HiCache/FA3 decode-backend override (#16531)) introduces some bugs which lead to the above error.
Reproduction
see the description in the above summary
Environment