Skip to content

Releases: cosdt/text-generation-inference

v3.3.7-npu: fix(scripts): harden start-tgi.sh model path handling

Choose a tag to compare

@VenusTZZ VenusTZZ released this 21 Sep 07:49

基于上游 v3.3.7 的 Ascend NPU 适配版本(910B / CANN 9.1.0)。

验证栈:torch 2.9.0 / torch_npu 2.9.0.post2 / transformers 4.57.6 / kernels 0.5.0 / Python 3.12

包含:

  • 单卡 + 多卡 HCCL 张量并行(v3 后端,bf16)
  • 可移植服务脚本 start-tgi.sh / stop-tgi.sh / run-npu.sh:默认模型 Qwen/Qwen3-0.6B(ModelScope 自动下载)、默认端口
    8080,不再硬编码本机路径
  • Dockerfile_ascend 修复:aarch64 protobuf-compiler、Python 3.12 构建与运行一致、torch_npu 从华为昇腾源安装、kernels==0.5.0
  • 脚本加固:MODEL_ID 尾随空白清理、模型下载失败快速失败

使用:见 docs/npu/quick-start.md