用qwen很慢?
#3745
Replies: 4 comments
|
想帮你定位,但还缺几个关键信息:
如果 qwen 是本地跑的,7.7 tok/s 在 CPU 推理或小显存下是正常水平;如果是 API 却这么慢,则要看网络和并发配置。贴一下部署方式和 |
0 replies
|
1、qwen3.8-max 云端API(百炼) |
0 replies
|
qwen3.8-max 是百炼的 max 档,本来就是"慢但聪明"那类,跟 v4 flash 这种快速档比 tok/s 没啥意义,两个不是一个速度等级。7.7 对 max 档来说算正常水平 想确认是不是 harness 转发拖慢的,最省事的办法是绕开它直连百炼测一次: import time
from openai import OpenAI
client = OpenAI(api_key="<key>", base_url="https://dashscope.aliyuncs.com/compatible-mode/v1")
t0 = time.time()
text = ""
for chunk in client.chat.completions.create(
model="qwen3-8-max",
messages=[{"role": "user", "content": "写300字短文"}],
stream=True,
):
if chunk.choices[0].delta.content:
text += chunk.choices[0].delta.content
print(len(text) / (time.time() - t0), "字/s")直连也这速度就是模型/百炼的锅;直连快很多再回来查 harness。 还有qwen3 系列默认可能带思考模式,思考不算 tok 但特别费时间,看看有没有传 |
0 replies
|
不能直接用 7.7 vs 81 tok/s 判断 Harness 有性能问题:qwen3.8-max 与 DeepSeek 快速档不是同一速度级别,thinking 也会增加首 token 延迟。正确控制组是同模型、同 prompt、同 reasoning 分别直连百炼和经 Harness,记录 TTFT、tok/s、reasoning/output tokens;只有直连明显更快才查 gateway。 |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
用qwen,7.7tok/s 切回ds,81tok/s 又遇到相似问题的吗?
All reactions