Replies: 5 comments
EN TL;DRThe per-route credential is resolved once per request, so rotating keys between attempts is possible today; and the 先把现状核实清楚,再回答你最后那两个边界问题 —— 因为那两个问题的答案决定了插件兜底能走多远。 1. 单凭据是「配置面」的事实,不是解析面的事实
⇒ 「换一把 key 再重试」在机制上今天就能成立:不需要改适配器,只要让下一次请求解析出另一个值。这正是你想要的形态,只是没有一条「声明多把 + 自动轮换」的配置路径。 2. 你问的两个边界,源码已有的答案(a) 哪个错误码可以换 key 重试失败分类是稳定机器码而非文案( 已有的可重试集合是 关键区分: (b) 换 key 重试是否只在响应体开始前你担心的那半已经是对的:重试决策发生在 但缺口恰好也在这里: ⇒ 这条是我给上游的具体诉求:把「本次尝试是否已产出可见内容」作为 3. 插件形态:不需要再改
|
|
先把版本钉住,不然后面全是空转:我这边跑的是 给 core 的两条:一个判别位,和一条别再靠巧合我要的东西很小:
回 argszero先认两条:凭据是每次请求解析的,所以「换一把再重试」机制上今天就成立; 三处对不上,都在版本上:
还有一处我这边打不通的:你插件 pod 的 我这边的处置:本机那跳轮换代理先留着 —— 不是因为你形状不好,是我这条路由现在换不动。上面那个嵌套寻址一旦能通, |
|
补充一条版本更新后的复测(本机 先回上次三条对不上的现状:
然后是这次升级里我真正想报的(正好接住你「限流和额度尽得分开」那句):
我本地的处置(本机改编译产物自用,不是提 PR):判定顺序改成 401 → 终态额度白名单( |
|
Pinning the version first, as asked. My line numbers were from master ( Your point 1: the timing sentence was mine to loseOn master today: The waterfall fires after the attempt's stream is already durable. So "fires before any partial stream is consumed" is wrong for this seam even at master, not only on 0.1.2-rc.1. It is true of exactly one thing — the llm layer's own retry, where Your point 1b:
|
|
版本先钉上:我核的是装着的这版 —— 你那颗旋钮在我这版上有,但 pi-ai 是每个 provider 一行: 但我方不把 QUOTA 放进去,而且拿你这条报文量过了:从 npm 拉未打补丁的 0.1.5-rc.2 三个产物跑同一套用例,23 条 10 红,其中两条正好相反 —— 周额度尽那条 429 判 AUTH 抹消息那半我这版也在: pod 那条我读了 0.1.1 的产物,跟你讲的一致: 脚本在这儿。数目先更正:上一条写「19 例」是 v3,09-17 加了端点侧内容审核四条,现在 23 例。它从装好的产物里按函数名抠 test_classification.js(23 例,纯字符串判定,不联网)// 从已打补丁的 dsh 产物里抽出两个分类器,用历史真实报错验证 v2/v3 语义。
// 覆盖 dsh-llm-pi-ai 的 classifyPiAiError + dsh-llm-deepseek 的 httpErrorCode。
// 本文件刻意不含反斜杠:换行统一用 String.fromCharCode(10) 构造。
// 两条真报文里各有一处账户侧的值(服务端 request id、周额度的重置时刻)已换成占位;分类只看措辞。
// v3 新增的用例全部是 2026-09-13 本轮 429 的真实报文与它的两个邻居(周额度尽 / FreeTierOnly)。
// v4 新增的用例是 2026-09-17 输出侧绿网(tokenplan/qwen3.8-flash 真报文)与它的 400/403 包裹形状。
const fs = require('fs'), path = require('path');
const NL = String.fromCharCode(10);
function findRoot() {
if (process.argv[2]) return process.argv[2];
if (process.env.DSH_INSTALL_ROOT) return process.env.DSH_INSTALL_ROOT;
// PATH 上的 dsh -> 同目录 node_modules/@deepseek-ai/dsh(2026-09-12 迁移后别再硬编码盘符)
const pathexes = (process.env.PATH || '').split(path.delimiter)
.map((d) => path.join(d, 'dsh'))
.filter((p) => { try { return fs.statSync(p).isFile(); } catch (e) { return false; } });
const cands = [];
for (const p of pathexes) {
cands.push(path.join(path.dirname(p), 'node_modules', '@deepseek-ai', 'dsh'));
try { cands.push(path.join(path.dirname(fs.realpathSync(p)), 'node_modules', '@deepseek-ai', 'dsh')); } catch (e) {}
}
for (const c of cands) {
if (fs.existsSync(path.join(c, 'lib'))) return c;
}
throw new Error('找不到 dsh 安装目录,位置参数传入或用 DSH_INSTALL_ROOT。试过: ' + cands.join(', '));
}
const root = findRoot();
console.log('dsh root:', root);
const pi = path.join(root, 'node_modules', '@deepseek-ai', 'dsh-llm-pi-ai', 'lib', 'index.js');
const ds = path.join(root, 'node_modules', '@deepseek-ai', 'dsh-llm-deepseek', 'lib', 'index.js');
const core = path.join(root, 'node_modules', '@deepseek-ai', 'dsh-llm', 'lib', 'index.js');
const srcPi = fs.readFileSync(pi, 'utf8');
const srcDs = fs.readFileSync(ds, 'utf8');
const srcCore = fs.readFileSync(core, 'utf8');
function grab(text, sig) {
const i = text.indexOf(sig);
if (i < 0) throw new Error('missing ' + sig);
const j = text.indexOf(NL + '}', i);
if (j < 0) throw new Error('unterminated ' + sig);
return text.slice(i, j + 2);
}
const code = [
grab(srcCore, 'function isQuotaExceededError('),
'const QUOTA_EXCEEDED_CODE = "QUOTA";',
'const CONTEXT_WINDOW_EXCEEDED_CODE = "CONTEXT_WINDOW_EXCEEDED";',
grab(srcPi, 'function classifyPiAiError('),
grab(srcDs, 'function httpErrorCode('),
'return { classifyPiAiError, httpErrorCode };',
].join(NL);
const { classifyPiAiError, httpErrorCode } = new Function(code)();
const piCases = [
// --- v2 就有的:不能退化 ---
['401+配额措辞(应仍AUTH)', '401: {"message":"Free quota exhausted","code":"insufficient_quota"}', 'AUTH'],
['真·密钥错(应仍AUTH)', '401: {"error":{"message":"Invalid Authentication api key"}}', 'AUTH'],
['百炼免费额度耗尽(403)', '403: {"message":"Free quota exhausted. To continue accessing the model"}', 'QUOTA'],
['OpenRouter 未购额度(402)', '402: {"message":"Insufficient credits. This account never purchased credits."}', 'QUOTA'],
['DeepSeek 余额不足', 'Insufficient Balance', 'QUOTA'],
['区域不可用(403)', '403: {"message":"This model is not available in your region.","metadata":{"failed_routing_step":"Gate Endpoints with Geo Restrictions"}}', 'GEO_OR_ENTITLEMENT'],
['纯 403 拒绝(应AUTH)', 'HTTP 403: Forbidden', 'AUTH'],
['429 限流', '429 Too Many Requests: rate limit exceeded', 'RATE_LIMIT'],
['上下文超限', '400: This model model max is 128000 tokens', 'INVALID_REQUEST'],
// --- v3 新增:2026-09-13 实测三类 429 必须分得开 ---
['TPM 打满(本轮真报文,应RATE_LIMIT)', '429: {"message":"Allocated quota exceeded, please increase your quota limit. For details, see: https://www.alibabacloud.com/help/en/model-studio/error-code#token-limit","id":"<request-id>","type":"insufficient_quota","code":"insufficient_quota"}', 'RATE_LIMIT'],
['tokenplan 周额度尽(应仍QUOTA)', '429: {"message":"Your token-plan 1-week quota has been exhausted. The quota will reset at <reset>."}', 'QUOTA'],
['安心模式用完即停(403,应QUOTA)', '403: {"code":"AllocationQuota.FreeTierOnly","message":"Free quota only. The free quota has been exhausted."}', 'QUOTA'],
['429 通用超额措辞(应RATE_LIMIT)', '429: {"type":"insufficient_quota","message":"You exceeded your current quota, please check your plan and billing details."}', 'RATE_LIMIT'],
['429 RPM 请求数限流(应RATE_LIMIT)', '429: {"code":"Throttling.RateQuota","message":"Requests call limit exceeded."}', 'RATE_LIMIT'],
// --- v4 新增:2026-09-17 端点侧内容审核(绿网)必须与 INVALID_REQUEST 分开 ---
['绿网输出拦截(实测裸报文,应MODERATION_BLOCKED)', 'Output data may contain inappropriate content.', 'MODERATION_BLOCKED'],
['绿网输入拦截(400 包裹,应MODERATION_BLOCKED)', '400: {"code":"data_inspection_failed","message":"Input data may contain inappropriate content.","request_id":"abc"}', 'MODERATION_BLOCKED'],
['绿网输出拦截(403 包裹,应MODERATION_BLOCKED)', '403: {"code":"data_inspection_failed","message":"Output data may contain inappropriate content."}', 'MODERATION_BLOCKED'],
['普通 400 不被误判成审核(应INVALID_REQUEST)', '400: {"code":"invalid_parameter","message":"content length exceeds the limit"}', 'INVALID_REQUEST'],
];
const dsCases = [
['http 403 + insufficient_quota -> QUOTA', 403, { code: 'insufficient_quota', message: 'Free quota exhausted. Please recharge.' }, 'QUOTA'],
['http 403 plain -> AUTH', 403, { message: 'forbidden' }, 'AUTH'],
['http 401 + 配额措辞 -> AUTH', 401, { code: 'insufficient_quota', message: 'Free quota exhausted' }, 'AUTH'],
['http 429 + 额度耗尽措辞 -> QUOTA', 429, { code: 'insufficient_quota', message: 'account credits exhausted' }, 'QUOTA'],
['http 429 plain -> RATE_LIMIT', 429, { message: 'request rate limit exceeded' }, 'RATE_LIMIT'],
];
let bad = 0;
function check(ok, name, got, want) {
if (!ok) bad++;
console.log((ok ? 'PASS ' : 'FAIL ') + name + ' -> ' + got + ' (want ' + want + ')');
}
for (const [name, msg, want] of piCases) {
const got = classifyPiAiError(msg);
check(got === want, '[pi-ai] ' + name, got, want);
}
for (const [name, status, body, want] of dsCases) {
const got = httpErrorCode(status, body);
check(got === want, '[llm-deepseek] ' + name, got, want);
}
console.log(bad ? bad + ' 项不符' : '全部通过:401 保持 AUTH;403 配额/区域显示真实原因;429 的 TPM/RPM 限流按 RATE_LIMIT 可重试,周额度尽仍判 QUOTA;绿网拦下判 MODERATION_BLOCKED、普通 400 不受影响');
process.exit(bad ? 1 : 0); |
Uh oh!
There was an error while loading. Please reload this page.
场景
同一个 provider 路由下想挂多把 API key,在同一份模型列表上轮换使用:
429 ... error-code#token-limit);现状
llm-pi-ai的 provider profile 只有单个apiKeyEnv(一个凭据引用),路由与密钥是一对一。多个凭据目前只能靠"复制一个 provider"绕,代价是模型列表、显示名、选择器条目全部翻倍,且不会自动轮换。
请求
能否支持一个路由多凭据 + 自动轮换?形态上我不确定哪种更合你们的设计,列几个想到的:
apiKeyEnvs: [REF_A, REF_B](与apiKeyEnv二选一),失败重试时按顺序换下一把;.credentials.yaml里BAILIAN_KEYS存多行/多值),由适配器内部轮换;另外希望明确轮换的判定与重试边界:哪些错误码可以换 key 重试(
429、403 free quota exhausted、401?)、换 key 重试是否只在响应体开始前进行(流式开始后不重试)——这点影响我们这边怎么设计兜底。
我们现在的兜底(说明这不是空想)
在插件里做了一个本机轮换代理:provider 的
baseURL指向本机的转发端点,代理持有密钥池(读
.credentials.yaml的同名前缀多条 ref),遇到可轮换的错误码就换下一把重试,正常请求流式透传。它解决了"今天就要用"的问题,但代价是:要改 provider 的 baseURL、多一跳本地转发、
还得自己实现 token 鉴权防止本机其它进程蹭额度(插件路由不走浏览器鉴权)。如果上游原生支持,这些都能省掉。
环境
@deepseek-ai/dsh0.1.2-rc.1,Windows 本机部署,Web profile。All reactions