I am trying to user FlashInfer on H800
I encounterd this error:
/usr/local/cuda/bin/nvcc --generate-dependencies-with-compile --dependency-output single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_12│
8_posenc_0_use_swa_False_use_logits_cap_False/single_decode_jit_pybind.cuda.o.d -DTORCH_EXTENSION_NAME=single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_│
qk_128_head_dim_vo_128_posenc_0_use_swa_False_use_logits_cap_False -DTORCH_API_INCLUDE_EXTENSION_H -DPy_LIMITED_API=0x03090000 -DPYBIND11_COMPILER_TYPE="_gcc" -DPYBIND11_STDLI│
B="_libstdcpp" -DPYBIND11_BUILD_ABI="_cxxabi1016" -D_GLIBCXX_USE_CXX11_ABI=1 -isystem include/python3.10 -isystem
/lib/python3.10/site-packages/torch/include -isystem python3.10/site-packages/torch/include/torch/csrc/api/include -isystem /usr│
/local/cuda/include -isystem /usr/local/cuda/include/cccl -isystem /lib/python3.10/site-packages/flashinfer/data/include -isystem /lib/python3.10/site-packages/flashinfer/data/csrc -isystem /lib/python3.10/site-packages/flashinfer/data/cutlass/include │
-isystem /lib/python3.10/site-packages/flashinfer/data/cutlass/tools/util/include -isystem /lib/python3.10/si│
te-packages/flashinfer/data/spdlog/include --compiler-options=-fPIC --expt-relaxed-constexpr -gencode=arch=compute_90a,code=sm_90a -DFLASHINFER_ENABLE_FP8_E8M0 -DFLASHINFER_ENAB│
LE_FP4_E2M1 -O3 -std=c++17 --threads=1 -use_fast_math -DFLASHINFER_ENABLE_F16 -DFLASHINFER_ENABLE_BF16 -DFLASHINFER_ENABLE_FP8_E4M3 -DFLASHINFER_ENABLE_FP8_E5M2 -DNDEBUG -c /roo│
t/.cache/flashinfer/52_60_61_70_75_80_86_90/generated/single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_128_posenc_0_use_swa_False_use│
_logits_cap_False/single_decode_jit_pybind.cu -o single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_128_posenc_0_use_swa_False_use_logi│
ts_cap_False/single_decode_jit_pybind.cuda.o │
nvcc fatal : Unsupported gpu architecture 'compute_90a' │
ninja: build stopped: subcommand failed.
Here is my environment:
bitsandbytes 0.47.0
cuda-bindings 13.0.1
cuda-pathfinder 1.2.1
cuda-python 13.0.1
cupy-cuda12x 13.6.0
cycler 0.12.1
deepspeed 0.17.5
depyf 0.19.0
dill 0.3.8
diskcache 5.6.3
fastapi 0.116.1
fastapi-cli 0.0.10
fastapi-cloud-cli 0.1.5
fastrlock 0.8.3
fastuuid 0.12.0
ffmpy 0.6.1
filelock 3.19.1
fire 0.7.1
flash_attn 2.8.3
flashinfer-python 0.3.1
gradio 5.32.1
gradio_client 1.10.2
hf_transfer 0.1.9
hf-xet 1.1.9
jieba 0.42.1
Jinja2 3.1.6
jiter 0.10.0
jmespath 0.10.0
joblib 1.5.2
jsonschema 4.25.1
jsonschema-specifications 2025.4.1
kiwisolver 1.4.9
lark 1.2.2
latex2sympy2_extended 1.10.2
lazy_loader 0.4
librosa 0.11.0
liger_kernel 0.6.2
litellm 1.76.2
llguidance 0.7.30
llvmlite 0.44.0
lm-format-enforcer 0.10.12
lmdeploy 0.9.2.post1
lxml 6.0.1
math-verify 0.8.0
matplotlib 3.10.6
matplotlib-inline 0.1.7
mdurl 0.1.2
mistral_common 1.8.4
mmengine-lite 0.10.7
modelscope 1.29.2
mpmath 1.3.0
ms_swift 3.8.0.dev0
msgpack 1.1.1
msgspec 0.19.0
multidict 6.6.4
multiprocess 0.70.16
nest-asyncio 1.6.0
networkx 3.4.2
ninja 1.13.0
nltk 3.9.1
nodeenv 1.9.1
numba 0.61.2
numpy 2.2.6
nvidia-cublas-cu12 12.6.4.1
nvidia-cuda-cupti-cu12 12.6.80
nvidia-cuda-nvrtc-cu12 12.6.77
nvidia-cuda-runtime-cu12 12.6.77
nvidia-cudnn-cu12 9.5.1.17
nvidia-cudnn-frontend 1.14.0
nvidia-cufft-cu12 11.3.0.4
nvidia-cufile-cu12 1.11.1.6
nvidia-curand-cu12 10.3.7.77
nvidia-cusolver-cu12 11.7.1.2
nvidia-cusparse-cu12 12.5.4.2
nvidia-cusparselt-cu12 0.6.3
nvidia-ml-py 12.575.51
nvidia-nccl-cu12 2.26.2
nvidia-nvjitlink-cu12 12.6.85
nvidia-nvshmem-cu12 3.3.24
nvidia-nvtx-cu12 12.6.77
nvitop 1.5.3
openai 1.90.0
openai-harmony 0.0.4
opencv-python-headless 4.12.0.88
optimum 1.27.0
orjson 3.11.3
oss2 2.19.1
outlines 0.1.11
outlines_core 0.2.10
packaging 25.0
pandas 2.3.2
parso 0.8.5
partial-json-parser 0.2.1.1.post6
peft 0.14.0
pyparsing 3.2.3
pyproject_hooks 1.2.0
pytz 2025.2
PyYAML 6.0.2
pyzmq 27.0.2
qwen-omni-utils 0.0.8
qwen-vl-utils 0.0.11
ray 2.49.1
referencing 0.36.2
regex 2025.9.1
requests 2.32.5
rich 13.9.4
rich-toolkit 0.15.0
rignore 0.6.4
rouge 1.0.1
rpds-py 0.27.1
ruff 0.12.11
s3transfer 0.13.1
safehttpx 0.1.6
safetensors 0.6.2
scikit-learn 1.7.1
scipy 1.15.3
semantic-version 2.10.0
sentencepiece 0.2.1
sentry-sdk 2.36.0
setproctitle 1.3.6
setuptools 78.1.1
sgl-kernel 0.2.8
sglang 0.4.10.post2
shellingham 1.5.4
shortuuid 1.0.13
simplejson 3.20.1
six 1.17.0
smmap 5.0.2
sniffio 1.3.1
sortedcontainers 2.4.0
soundfile 0.13.1
soxr 0.5.0.post1
stack-data 0.6.3
starlette 0.47.3
tiktoken 0.11.0
timeout-decorator 0.5.0
torch 2.7.1
torch_memory_saver 0.0.8
torchao 0.9.0
torchaudio 2.7.1
torchvision 0.22.1
traitlets 5.14.3
transformers 4.55.4
transformers-stream-generator 0.0.5
triton 3.3.1
trl 0.20.0
vllm 0.10.0
I am trying to user FlashInfer on H800
I encounterd this error:
/usr/local/cuda/bin/nvcc --generate-dependencies-with-compile --dependency-output single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_12│
8_posenc_0_use_swa_False_use_logits_cap_False/single_decode_jit_pybind.cuda.o.d -DTORCH_EXTENSION_NAME=single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_│
qk_128_head_dim_vo_128_posenc_0_use_swa_False_use_logits_cap_False -DTORCH_API_INCLUDE_EXTENSION_H -DPy_LIMITED_API=0x03090000 -DPYBIND11_COMPILER_TYPE="_gcc" -DPYBIND11_STDLI│
B="_libstdcpp" -DPYBIND11_BUILD_ABI="_cxxabi1016" -D_GLIBCXX_USE_CXX11_ABI=1 -isystem include/python3.10 -isystem
/lib/python3.10/site-packages/torch/include -isystem python3.10/site-packages/torch/include/torch/csrc/api/include -isystem /usr│
/local/cuda/include -isystem /usr/local/cuda/include/cccl -isystem /lib/python3.10/site-packages/flashinfer/data/include -isystem /lib/python3.10/site-packages/flashinfer/data/csrc -isystem /lib/python3.10/site-packages/flashinfer/data/cutlass/include │
-isystem /lib/python3.10/site-packages/flashinfer/data/cutlass/tools/util/include -isystem /lib/python3.10/si│
te-packages/flashinfer/data/spdlog/include --compiler-options=-fPIC --expt-relaxed-constexpr -gencode=arch=compute_90a,code=sm_90a -DFLASHINFER_ENABLE_FP8_E8M0 -DFLASHINFER_ENAB│
LE_FP4_E2M1 -O3 -std=c++17 --threads=1 -use_fast_math -DFLASHINFER_ENABLE_F16 -DFLASHINFER_ENABLE_BF16 -DFLASHINFER_ENABLE_FP8_E4M3 -DFLASHINFER_ENABLE_FP8_E5M2 -DNDEBUG -c /roo│
t/.cache/flashinfer/52_60_61_70_75_80_86_90/generated/single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_128_posenc_0_use_swa_False_use│
_logits_cap_False/single_decode_jit_pybind.cu -o single_decode_with_kv_cache_dtype_q_f16_dtype_kv_f16_dtype_o_f16_head_dim_qk_128_head_dim_vo_128_posenc_0_use_swa_False_use_logi│
ts_cap_False/single_decode_jit_pybind.cuda.o │
nvcc fatal : Unsupported gpu architecture 'compute_90a' │
ninja: build stopped: subcommand failed.
Here is my environment:
bitsandbytes 0.47.0
cuda-bindings 13.0.1
cuda-pathfinder 1.2.1
cuda-python 13.0.1
cupy-cuda12x 13.6.0
cycler 0.12.1
deepspeed 0.17.5
depyf 0.19.0
dill 0.3.8
diskcache 5.6.3
fastapi 0.116.1
fastapi-cli 0.0.10
fastapi-cloud-cli 0.1.5
fastrlock 0.8.3
fastuuid 0.12.0
ffmpy 0.6.1
filelock 3.19.1
fire 0.7.1
flash_attn 2.8.3
flashinfer-python 0.3.1
gradio 5.32.1
gradio_client 1.10.2
hf_transfer 0.1.9
hf-xet 1.1.9
jieba 0.42.1
Jinja2 3.1.6
jiter 0.10.0
jmespath 0.10.0
joblib 1.5.2
jsonschema 4.25.1
jsonschema-specifications 2025.4.1
kiwisolver 1.4.9
lark 1.2.2
latex2sympy2_extended 1.10.2
lazy_loader 0.4
librosa 0.11.0
liger_kernel 0.6.2
litellm 1.76.2
llguidance 0.7.30
llvmlite 0.44.0
lm-format-enforcer 0.10.12
lmdeploy 0.9.2.post1
lxml 6.0.1
math-verify 0.8.0
matplotlib 3.10.6
matplotlib-inline 0.1.7
mdurl 0.1.2
mistral_common 1.8.4
mmengine-lite 0.10.7
modelscope 1.29.2
mpmath 1.3.0
ms_swift 3.8.0.dev0
msgpack 1.1.1
msgspec 0.19.0
multidict 6.6.4
multiprocess 0.70.16
nest-asyncio 1.6.0
networkx 3.4.2
ninja 1.13.0
nltk 3.9.1
nodeenv 1.9.1
numba 0.61.2
numpy 2.2.6
nvidia-cublas-cu12 12.6.4.1
nvidia-cuda-cupti-cu12 12.6.80
nvidia-cuda-nvrtc-cu12 12.6.77
nvidia-cuda-runtime-cu12 12.6.77
nvidia-cudnn-cu12 9.5.1.17
nvidia-cudnn-frontend 1.14.0
nvidia-cufft-cu12 11.3.0.4
nvidia-cufile-cu12 1.11.1.6
nvidia-curand-cu12 10.3.7.77
nvidia-cusolver-cu12 11.7.1.2
nvidia-cusparse-cu12 12.5.4.2
nvidia-cusparselt-cu12 0.6.3
nvidia-ml-py 12.575.51
nvidia-nccl-cu12 2.26.2
nvidia-nvjitlink-cu12 12.6.85
nvidia-nvshmem-cu12 3.3.24
nvidia-nvtx-cu12 12.6.77
nvitop 1.5.3
openai 1.90.0
openai-harmony 0.0.4
opencv-python-headless 4.12.0.88
optimum 1.27.0
orjson 3.11.3
oss2 2.19.1
outlines 0.1.11
outlines_core 0.2.10
packaging 25.0
pandas 2.3.2
parso 0.8.5
partial-json-parser 0.2.1.1.post6
peft 0.14.0
pyparsing 3.2.3
pyproject_hooks 1.2.0
pytz 2025.2
PyYAML 6.0.2
pyzmq 27.0.2
qwen-omni-utils 0.0.8
qwen-vl-utils 0.0.11
ray 2.49.1
referencing 0.36.2
regex 2025.9.1
requests 2.32.5
rich 13.9.4
rich-toolkit 0.15.0
rignore 0.6.4
rouge 1.0.1
rpds-py 0.27.1
ruff 0.12.11
s3transfer 0.13.1
safehttpx 0.1.6
safetensors 0.6.2
scikit-learn 1.7.1
scipy 1.15.3
semantic-version 2.10.0
sentencepiece 0.2.1
sentry-sdk 2.36.0
setproctitle 1.3.6
setuptools 78.1.1
sgl-kernel 0.2.8
sglang 0.4.10.post2
shellingham 1.5.4
shortuuid 1.0.13
simplejson 3.20.1
six 1.17.0
smmap 5.0.2
sniffio 1.3.1
sortedcontainers 2.4.0
soundfile 0.13.1
soxr 0.5.0.post1
stack-data 0.6.3
starlette 0.47.3
tiktoken 0.11.0
timeout-decorator 0.5.0
torch 2.7.1
torch_memory_saver 0.0.8
torchao 0.9.0
torchaudio 2.7.1
torchvision 0.22.1
traitlets 5.14.3
transformers 4.55.4
transformers-stream-generator 0.0.5
triton 3.3.1
trl 0.20.0
vllm 0.10.0