INFO 06-03 15:25:43 [importing.py:16] Triton not installed or not compatible; certain GPU-related functions will not be available.
WARNING 06-03 15:25:43 [importing.py:28] Triton is not installed. Using dummy decorators. Install it via `pip install triton` to enable kernel compilation.
INFO 06-03 15:25:46 [__init__.py:30] Available plugins for group vllm.platform_plugins:
INFO 06-03 15:25:46 [__init__.py:32] name=ascend, value=vllm_ascend:register
INFO 06-03 15:25:46 [__init__.py:44] plugin ascend loaded.
INFO 06-03 15:25:46 [__init__.py:239] Platform plugin ascend is activated
WARNING 06-03 15:25:49 [_custom_ops.py:21] Failed to import from vllm._C with ModuleNotFoundError("No module named 'vllm._C'")
INFO 06-03 15:25:51 [__init__.py:30] Available plugins for group vllm.general_plugins:
INFO 06-03 15:25:51 [__init__.py:32] name=lora_filesystem_resolver, value=vllm.plugins.lora_resolvers.filesystem_resolver:register_filesystem_resolver
INFO 06-03 15:25:51 [__init__.py:32] name=ascend_enhanced_model, value=vllm_ascend:register_model
INFO 06-03 15:25:53 [config.py:1903] Disabled the custom all-reduce kernel because it is not supported on current platform.
INFO 06-03 15:25:55 [api_server.py:1289] vLLM API server version 0.1.dev6389+g246e3e0.empty
INFO 06-03 15:25:57 [config.py:1903] Disabled the custom all-reduce kernel because it is not supported on current platform.
INFO 06-03 15:25:57 [cli_args.py:300] non-default args: {'host': '127.0.0.1', 'port': 18080, 'model': '/home/model/Qwen2.5-7B-Instruct', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_model_len': 4096, 'disable_custom_all_reduce': True, 'gpu_memory_utilization': 0.8, 'max_num_seqs': 4}
INFO 06-03 15:26:14 [config.py:787] This model supports multiple tasks: {'classify', 'embed', 'generate', 'reward', 'score'}. Defaulting to 'generate'.
WARNING 06-03 15:26:14 [arg_utils.py:1595] Detected VLLM_USE_V1=1 with npu. Usage should be considered experimental. Please report any issues on Github.
INFO 06-03 15:26:14 [config.py:1903] Disabled the custom all-reduce kernel because it is not supported on current platform.
INFO 06-03 15:26:14 [config.py:2112] Chunked prefill is enabled with max_num_batched_tokens=2048.
WARNING 06-03 15:26:14 [platform.py:142] NPU compilation support pending. Will be available in future CANN and torch_npu releases. NPU graph mode is currently experimental and disabled by default. You can just adopt additional_config={'enable_graph_mode': True} to serve deepseek models with NPU graph mode on vllm-ascend with V0 engine.
INFO 06-03 15:26:14 [platform.py:150] Compilation disabled, using eager mode by default
INFO 06-03 15:26:21 [importing.py:16] Triton not installed or not compatible; certain GPU-related functions will not be available.
WARNING 06-03 15:26:21 [importing.py:28] Triton is not installed. Using dummy decorators. Install it via `pip install triton` to enable kernel compilation.
INFO 06-03 15:26:24 [__init__.py:30] Available plugins for group vllm.platform_plugins:
INFO 06-03 15:26:24 [__init__.py:32] name=ascend, value=vllm_ascend:register
INFO 06-03 15:26:24 [__init__.py:44] plugin ascend loaded.
INFO 06-03 15:26:24 [__init__.py:239] Platform plugin ascend is activated
WARNING 06-03 15:26:27 [_custom_ops.py:21] Failed to import from vllm._C with ModuleNotFoundError("No module named 'vllm._C'")
INFO 06-03 15:26:30 [core.py:427] Waiting for init message from front-end.
WARNING 06-03 15:26:30 [platform.py:142] NPU compilation support pending. Will be available in future CANN and torch_npu releases. NPU graph mode is currently experimental and disabled by default. You can just adopt additional_config={'enable_graph_mode': True} to serve deepseek models with NPU graph mode on vllm-ascend with V0 engine.
INFO 06-03 15:26:30 [platform.py:150] Compilation disabled, using eager mode by default
INFO 06-03 15:26:30 [core.py:61] Initializing a V1 LLM engine (v0.1.dev6389+g246e3e0.empty) with config: model='/home/model/Qwen2.5-7B-Instruct', speculative_config=None, tokenizer='/home/model/Qwen2.5-7B-Instruct', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, override_neuron_config={}, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=4096, download_dir=None, load_format=LoadFormat.AUTO, tensor_parallel_size=1, pipeline_parallel_size=1, disable_custom_all_reduce=True, quantization=None, enforce_eager=False, kv_cache_dtype=auto, device_config=npu, decoding_config=DecodingConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_backend=''), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None), seed=0, served_model_name=/home/model/Qwen2.5-7B-Instruct, num_scheduler_steps=1, multi_step_stream_outputs=True, enable_prefix_caching=True, chunked_prefill_enabled=True, use_async_output_proc=True, pooler_config=None, compilation_config={"custom_ops": ["all"], "splitting_ops": ["vllm.unified_attention", "vllm.unified_attention_with_output"], "compile_sizes": [], "use_cudagraph": true, "cudagraph_num_of_warmups": 1, "cudagraph_capture_sizes": [512, 504, 496, 488, 480, 472, 464, 456, 448, 440, 432, 424, 416, 408, 400, 392, 384, 376, 368, 360, 352, 344, 336, 328, 320, 312, 304, 296, 288, 280, 272, 264, 256, 248, 240, 232, 224, 216, 208, 200, 192, 184, 176, 168, 160, 152, 144, 136, 128, 120, 112, 104, 96, 88, 80, 72, 64, 56, 48, 40, 32, 24, 16, 8, 4, 2, 1], "max_capture_size": 512}
INFO 06-03 15:26:30 [__init__.py:30] Available plugins for group vllm.general_plugins:
INFO 06-03 15:26:30 [__init__.py:32] name=lora_filesystem_resolver, value=vllm.plugins.lora_resolvers.filesystem_resolver:register_filesystem_resolver
INFO 06-03 15:26:30 [__init__.py:32] name=ascend_enhanced_model, value=vllm_ascend:register_model
ERROR 06-03 15:26:30 [core.py:489] EngineCore failed to start.
ERROR 06-03 15:26:30 [core.py:489] Traceback (most recent call last):
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 480, in run_engine_core
ERROR 06-03 15:26:30 [core.py:489] engine_core = EngineCoreProc(*args, **kwargs)
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 379, in __init__
ERROR 06-03 15:26:30 [core.py:489] super().__init__(vllm_config, executor_class, log_stats,
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/v1/engine/core.py", line 67, in __init__
ERROR 06-03 15:26:30 [core.py:489] self.model_executor = executor_class(vllm_config)
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/executor/executor_base.py", line 52, in __init__
ERROR 06-03 15:26:30 [core.py:489] self._init_executor()
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/executor/uniproc_executor.py", line 45, in _init_executor
ERROR 06-03 15:26:30 [core.py:489] self.collective_rpc("init_worker", args=([kwargs], ))
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/executor/uniproc_executor.py", line 56, in collective_rpc
ERROR 06-03 15:26:30 [core.py:489] answer = run_method(self.driver_worker, method, args, kwargs)
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/utils.py", line 2598, in run_method
ERROR 06-03 15:26:30 [core.py:489] return func(*args, **kwargs)
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/worker/worker_base.py", line 558, in init_worker
ERROR 06-03 15:26:30 [core.py:489] worker_class = resolve_obj_by_qualname(
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/utils.py", line 2184, in resolve_obj_by_qualname
ERROR 06-03 15:26:30 [core.py:489] module = importlib.import_module(module_name)
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/importlib/__init__.py", line 126, in import_module
ERROR 06-03 15:26:30 [core.py:489] return _bootstrap._gcd_import(name[level:], package, level)
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "<frozen importlib._bootstrap>", line 1204, in _gcd_import
ERROR 06-03 15:26:30 [core.py:489] File "<frozen importlib._bootstrap>", line 1176, in _find_and_load
ERROR 06-03 15:26:30 [core.py:489] File "<frozen importlib._bootstrap>", line 1147, in _find_and_load_unlocked
ERROR 06-03 15:26:30 [core.py:489] File "<frozen importlib._bootstrap>", line 690, in _load_unlocked
ERROR 06-03 15:26:30 [core.py:489] File "<frozen importlib._bootstrap_external>", line 940, in exec_module
ERROR 06-03 15:26:30 [core.py:489] File "<frozen importlib._bootstrap>", line 241, in _call_with_frames_removed
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm_ascend/worker/worker_v1.py", line 46, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm_ascend.worker.model_runner_v1 import NPUModelRunner
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm_ascend/worker/model_runner_v1.py", line 39, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.fused_moe import FusedMoE
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/fused_moe/__init__.py", line 6, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.fused_moe.layer import (
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/fused_moe/layer.py", line 51, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.fused_moe.fused_moe import grouped_topk
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/fused_moe/fused_moe.py", line 14, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.fused_moe.deep_gemm_moe import (
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/fused_moe/deep_gemm_moe.py", line 10, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.fused_moe.moe_permute_unpermute import (
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/fused_moe/moe_permute_unpermute.py", line 9, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.fused_moe.utils import _fp8_perm
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/fused_moe/utils.py", line 8, in <module>
ERROR 06-03 15:26:30 [core.py:489] from vllm.model_executor.layers.quantization.utils.fp8_utils import (
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/model_executor/layers/quantization/utils/fp8_utils.py", line 183, in <module>
ERROR 06-03 15:26:30 [core.py:489] direct_register_custom_op(
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/vllm/utils.py", line 2166, in direct_register_custom_op
ERROR 06-03 15:26:30 [core.py:489] schema_str = torch.library.infer_schema(op_func,
ERROR 06-03 15:26:30 [core.py:489] ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/torch/_library/infer_schema.py", line 106, in infer_schema
ERROR 06-03 15:26:30 [core.py:489] error_fn(
ERROR 06-03 15:26:30 [core.py:489] File "/opt/huawei/miniconda3/lib/python3.11/site-packages/torch/_library/infer_schema.py", line 58, in error_fn
ERROR 06-03 15:26:30 [core.py:489] raise ValueError(
ERROR 06-03 15:26:30 [core.py:489] ValueError: infer_schema(func): Parameter block_size has unsupported type list[int]. The valid types are: dict_keys([<class 'torch.Tensor'>, typing.Optional[torch.Tensor], typing.Sequence[torch.Tensor], typing.List[torch.Tensor], typing.Sequence[typing.Optional[torch.Tensor]], typing.List[typing.Optional[torch.Tensor]], <class 'int'>, typing.Optional[int], typing.Sequence[int], typing.List[int], typing.Optional[typing.Sequence[int]], typing.Optional[typing.List[int]], <class 'float'>, typing.Optional[float], typing.Sequence[float], typing.List[float], typing.Optional[typing.Sequence[float]], typing.Optional[typing.List[float]], <class 'bool'>, typing.Optional[bool], typing.Sequence[bool], typing.List[bool], typing.Optional[typing.Sequence[bool]], typing.Optional[typing.List[bool]], <class 'str'>, typing.Optional[str], typing.Union[int, float, bool], typing.Union[int, float, bool, NoneType], typing.Sequence[typing.Union[int, float, bool]], typing.List[typing.Union[int, float, bool]], <class 'torch.dtype'>, typing.Optional[torch.dtype], <class 'torch.device'>, typing.Optional[torch.device]]). Got func with signature (input: torch.Tensor, weight: torch.Tensor, block_size: list[int], weight_scale: torch.Tensor, input_scale: Optional[torch.Tensor] = None, bias: Optional[torch.Tensor] = None, cutlass_block_fp8_supported: bool = False, use_aiter_and_is_supported: bool = False) -> torch.Tensor)
[ERROR] 2025-06-03-15:26:33 (PID:1399, Device:-1, RankID:-1) ERR99999 UNKNOWN applicaiton exception
Your current environment
The output of `python collect_env.py`
🐛 Describe the bug
使用 vllm-ascend main 分支代码启动 openai api_server:
报错如下:
同样的环境, 能正常进行离线推理。