Replies: 2 comments 1 reply
|
Your root-cause analysis is accurate, and this is a clean two-part bug. I read the alpha.2 source ( Part 1 — static catalog short-circuit (confirmed).
if (request.provider !== undefined) {
const installed = catalogModels(request.provider)
if (installed.size > 0) {
return [...installed.values()].map(model => ({ id: model.id, ... })) // ← never hits network
}
}The Part 2 — 410 not classified (confirmed, and this is the higher-risk half).
The risk is the retry classification. Proposed fix directions (highest-to-lowest blast radius):
For the test suite, a good pairing is: (a) a discovery test that a catalog provider marked "dynamic" calls network rather than returning the static list; (b) a This is the same error-classification class as the missing-402 handling in |



Uh oh!
There was an error while loading. Please reload this page.
问题描述
当用户使用 DSH 内置的 nvidia provider(不显式配置 models 列表)时,模型列表来自 pi-ai 的静态 catalog,其中的模型条目已过时(pre-DeepSeek / pre-Gemma-4 时代)。调用这些模型会触发 NVIDIA API 返回 HTTP 410 Gone,错误码被模糊地归类为 PI_AI_ERROR。
相比之下,用户自定义的 nvidia-nim provider(不在 pi-ai catalog 中)会走动态 GET /models 路径,能获取最新列表,正常工作。
复现步骤
在 settings.yaml 中配置:
yaml
复制
llm-pi-ai:
providers:
nvidia:
apiKeyEnv: NVIDIA_API_KEY
不显式声明 models 列表
尝试使用内置 nvidia provider 的任意模型(如 meta/llama-3.1-70b-instruct)
观察错误:PI_AI_ERROR / 410 status code (no body)
期望行为
内置 nvidia provider 应像自定义 provider 一样,动态从 NVIDIA API 拉取最新可用模型列表,或在 catalog 层面自动替换已过时的模型。
实际行为
返回 HTTP 410 Gone,错误码为不透明的 PI_AI_ERROR。
根因分析
packages/llm/llm-pi-ai/src/discovery.ts 第 199-211 行:
typescript
复制
// catalog 路由直接短路返回静态列表,从不调 API
if (request.provider !== undefined) {
const installed = catalogModels(request.provider)
if (installed.size > 0) {
return [...installed.values()] // ← 永远不走网络
}
}
而 nvidia-nim 等自定义 provider 因为不在 pi-ai catalog 中,会走到第 232 行的 fetch(url) 动态获取逻辑。
建议修复方向
discovery.ts:对 nvidia provider 跳过静态 catalog 短路,改为动态探测
catalog.ts:新增 NVIDIA_CATALOG_PATCH 常量,在 resolveRouteModels 中自动替换过时模型
stream.ts:classifyPiAiError 增加 410 分类(返回 GONE),错误信息更明确
All reactions