Skip to content

nvcc fatal : Unsupported gpu architecture 'compute_90a' #1658

Description

@XingyuLu206

I am trying to user FlashInfer on H800
I encounterd this error:
/usr/local/cuda/bin/nvcc --generate-dependencies-with-compile --dependency-output single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_12│
8_posenc_0_use_swa_False_use_logits_cap_False/single_decode_jit_pybind.cuda.o.d -DTORCH_EXTENSION_NAME=single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_│
qk_128_head_dim_vo_128_posenc_0_use_swa_False_use_logits_cap_False -DTORCH_API_INCLUDE_EXTENSION_H -DPy_LIMITED_API=0x03090000 -DPYBIND11_COMPILER_TYPE="_gcc" -DPYBIND11_STDLI│
B="_libstdcpp" -DPYBIND11_BUILD_ABI="_cxxabi1016" -D_GLIBCXX_USE_CXX11_ABI=1 -isystem include/python3.10 -isystem
/lib/python3.10/site-packages/torch/include -isystem python3.10/site-packages/torch/include/torch/csrc/api/include -isystem /usr│
/local/cuda/include -isystem /usr/local/cuda/include/cccl -isystem /lib/python3.10/site-packages/flashinfer/data/include -isystem /lib/python3.10/site-packages/flashinfer/data/csrc -isystem /lib/python3.10/site-packages/flashinfer/data/cutlass/include │
-isystem /lib/python3.10/site-packages/flashinfer/data/cutlass/tools/util/include -isystem /lib/python3.10/si│
te-packages/flashinfer/data/spdlog/include --compiler-options=-fPIC --expt-relaxed-constexpr -gencode=arch=compute_90a,code=sm_90a -DFLASHINFER_ENABLE_FP8_E8M0 -DFLASHINFER_ENAB│
LE_FP4_E2M1 -O3 -std=c++17 --threads=1 -use_fast_math -DFLASHINFER_ENABLE_F16 -DFLASHINFER_ENABLE_BF16 -DFLASHINFER_ENABLE_FP8_E4M3 -DFLASHINFER_ENABLE_FP8_E5M2 -DNDEBUG -c /roo│
t/.cache/flashinfer/52_60_61_70_75_80_86_90/generated/single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_128_posenc_0_use_swa_False_use│
_logits_cap_False/single_decode_jit_pybind.cu -o single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_128_posenc_0_use_swa_False_use_logi│
ts_cap_False/single_decode_jit_pybind.cuda.o │
nvcc fatal : Unsupported gpu architecture 'compute_90a' │
ninja: build stopped: subcommand failed.

Here is my environment:
bitsandbytes 0.47.0
cuda-bindings 13.0.1
cuda-pathfinder 1.2.1
cuda-python 13.0.1
cupy-cuda12x 13.6.0
cycler 0.12.1
deepspeed 0.17.5
depyf 0.19.0
dill 0.3.8
diskcache 5.6.3
fastapi 0.116.1
fastapi-cli 0.0.10
fastapi-cloud-cli 0.1.5
fastrlock 0.8.3
fastuuid 0.12.0
ffmpy 0.6.1
filelock 3.19.1
fire 0.7.1
flash_attn 2.8.3
flashinfer-python 0.3.1
gradio 5.32.1
gradio_client 1.10.2
hf_transfer 0.1.9
hf-xet 1.1.9
jieba 0.42.1
Jinja2 3.1.6
jiter 0.10.0
jmespath 0.10.0
joblib 1.5.2
jsonschema 4.25.1
jsonschema-specifications 2025.4.1
kiwisolver 1.4.9
lark 1.2.2
latex2sympy2_extended 1.10.2
lazy_loader 0.4
librosa 0.11.0
liger_kernel 0.6.2
litellm 1.76.2
llguidance 0.7.30
llvmlite 0.44.0
lm-format-enforcer 0.10.12
lmdeploy 0.9.2.post1
lxml 6.0.1
math-verify 0.8.0
matplotlib 3.10.6
matplotlib-inline 0.1.7
mdurl 0.1.2
mistral_common 1.8.4
mmengine-lite 0.10.7
modelscope 1.29.2
mpmath 1.3.0
ms_swift 3.8.0.dev0
msgpack 1.1.1
msgspec 0.19.0
multidict 6.6.4
multiprocess 0.70.16
nest-asyncio 1.6.0
networkx 3.4.2
ninja 1.13.0
nltk 3.9.1
nodeenv 1.9.1
numba 0.61.2
numpy 2.2.6
nvidia-cublas-cu12 12.6.4.1
nvidia-cuda-cupti-cu12 12.6.80
nvidia-cuda-nvrtc-cu12 12.6.77
nvidia-cuda-runtime-cu12 12.6.77
nvidia-cudnn-cu12 9.5.1.17
nvidia-cudnn-frontend 1.14.0
nvidia-cufft-cu12 11.3.0.4
nvidia-cufile-cu12 1.11.1.6
nvidia-curand-cu12 10.3.7.77
nvidia-cusolver-cu12 11.7.1.2
nvidia-cusparse-cu12 12.5.4.2
nvidia-cusparselt-cu12 0.6.3
nvidia-ml-py 12.575.51
nvidia-nccl-cu12 2.26.2
nvidia-nvjitlink-cu12 12.6.85
nvidia-nvshmem-cu12 3.3.24
nvidia-nvtx-cu12 12.6.77
nvitop 1.5.3
openai 1.90.0
openai-harmony 0.0.4
opencv-python-headless 4.12.0.88
optimum 1.27.0
orjson 3.11.3
oss2 2.19.1
outlines 0.1.11
outlines_core 0.2.10
packaging 25.0
pandas 2.3.2
parso 0.8.5
partial-json-parser 0.2.1.1.post6
peft 0.14.0
pyparsing 3.2.3
pyproject_hooks 1.2.0
pytz 2025.2
PyYAML 6.0.2
pyzmq 27.0.2
qwen-omni-utils 0.0.8
qwen-vl-utils 0.0.11
ray 2.49.1
referencing 0.36.2
regex 2025.9.1
requests 2.32.5
rich 13.9.4
rich-toolkit 0.15.0
rignore 0.6.4
rouge 1.0.1
rpds-py 0.27.1
ruff 0.12.11
s3transfer 0.13.1
safehttpx 0.1.6
safetensors 0.6.2
scikit-learn 1.7.1
scipy 1.15.3
semantic-version 2.10.0
sentencepiece 0.2.1
sentry-sdk 2.36.0
setproctitle 1.3.6
setuptools 78.1.1
sgl-kernel 0.2.8
sglang 0.4.10.post2
shellingham 1.5.4
shortuuid 1.0.13
simplejson 3.20.1
six 1.17.0
smmap 5.0.2
sniffio 1.3.1
sortedcontainers 2.4.0
soundfile 0.13.1
soxr 0.5.0.post1
stack-data 0.6.3
starlette 0.47.3
tiktoken 0.11.0
timeout-decorator 0.5.0
torch 2.7.1
torch_memory_saver 0.0.8
torchao 0.9.0
torchaudio 2.7.1
torchvision 0.22.1
traitlets 5.14.3
transformers 4.55.4
transformers-stream-generator 0.0.5
triton 3.3.1
trl 0.20.0
vllm 0.10.0

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions