GitHub Copilot's /models carries the capacities the catalog short-circuit assumes are missing #3200
Replies: 4 comments 1 reply
|
Follow-up: as a third-party plugin this does not build today, for two packaging reasons — neither of them in the plugin. I tried repackaging the above as a standalone plugin (the shape 1. Two copies of pi-ai. The published That is the compile-time face of a runtime hazard: a 2. Minor, but it costs an install to discover: the I am not asking for anything specific — recording it because the first two are the difference between "a plugin reuses the adapter" and "a plugin reimplements streaming", which seems worth knowing when deciding how much of this belongs upstream. The logic itself is framework-independent and its tests (protocol derivation against the installed catalog, the Responses |
|
对照当前 main HEAD 1. 目录短路(packages/llm/llm-pi-ai/src/discovery.ts) 模块头注释明写:"A route the installed pi-ai catalog ships is answered from that catalog, with no network call at all: pi-ai's registry is the authoritative list for its own providers, and it carries the capacities a listing endpoint would not disclose."——目录条目零网络调用,与你的描述一致。同时 2. profiles 记忆化接缝(packages/llm/llm-pi-ai/src/adapter.ts)
3. 对两条建议的倾向
4. 一个可落地的加固(把 400 变成配置期错误)
5. 关于 follow-up 的打包问题 两副本 pi-ai 与 #2660/#3033 的 dsh-tools 双实例同族(Symbol/单例分裂)。建议:插件声明 |
|
The code is public now in case it is easier to read than prose. Two stacked branches on my fork — not a PR, since
The test worth pointing at is |
|
Consumer-side datapoint for this proposal, from the one place that reads these capacities downstream of the seam you're changing: We maintain pi2dsh (runs Pi plugins on DSH; plugins read the model directory through our projection), and our examples regression caught exactly the split your proposal touches: DSH separates "directory membership" ( Why that matters for this thread: the short-circuit premise ("no listing endpoint reports capacities") being false for Copilot means the resolve half is exactly where live No opinion on the stacked branches' internals — the above is evidence that the capacity-resolution seam has real downstream consumers whose failure mode is silent-wrong, so whichever fix lands, a regression asserting "declared/live capacity actually reaches a route consumer" is worth including. |
Uh oh!
There was an error while loading. Please reload this page.
Summary
discovery.tsanswers a catalog-backed route from the installed pi-ai catalog and deliberately makes no network call, because "the installed entries carry context windows and output caps no listing endpoint reports". For GitHub Copilot that premise does not hold: itsGET /modelsreports everything a catalog entry needs. So a Copilot route could discover models live instead of waiting for a pi-ai release.This is a report rather than a patch, since
CONTRIBUTING.mdsays external PRs cannot be accepted.What Copilot's listing actually carries
Per model,
GET {endpoint}/modelsreturns:supported_endpointscapabilities.limits.max_context_window_tokenscapabilities.limits.max_output_tokenscapabilities.supports.reasoning_effort[]capabilities.supports.visionmodel_picker_enabled,policy.stateA real entry (2026-08-19):
{ "id": "gemini-3.7-flash", "name": "Gemini 3.7 Flash", "vendor": "Google", "model_picker_enabled": true, "policy": { "state": "enabled" }, "supported_endpoints": ["/chat/completions"], "capabilities": { "type": "chat", "limits": { "max_context_window_tokens": 1000000, "max_output_tokens": 64000 }, "supports": { "reasoning_effort": ["low","medium","high"], "vision": true, "tool_calls": true } } }Why it matters today
reuseCatalogProviderdrops catalog-owned dynamic refresh, andfilterModelsonly intersects, so a model GitHub ships after pi-ai's last release cannot be selected at all — not for want of entitlement, but because the id is absent from a compiled-in JSON file. Adopting it means raising a dependency range, rebuilding, and restarting the host.Measured on my install:
@earendil-works/pi-ai0.82.1's Copilot catalog had neithergemini-3.6-flashnorgemini-3.7-flash; 0.84.0 (2026-08-06) added 3.6, and 0.84.2 (2026-08-14) added 3.7. Upstream tracks Copilot within about a week — the lag is the dependency bump, not the data.Also worth noting: the caret range on a
0.xdependency (^0.82.1) pins the minor, sopnpm upcannot cross it; and 0.84 requires two drift-gate updates inllm-pi-ai(basetenadded toOpenAICompletionsCompat['thinkingFormat'], andStopReasongainingpending/deferred, which breaks the exhaustive switch instream.ts).Two rules a live path has to respect
I built a local plugin to check whether this is actually feasible, and both of these came out of running it rather than reading code.
1. Protocol choice is derivable, and the ordering matters. Many Copilot models accept several endpoints — every Anthropic model also answers
/chat/completions, and some GPT models answer both/responsesand/chat/completions. This order reproduces pi-ai's own curatedapifor every model that both the listing and the catalog describe (23 models on my account, asserted in a unit test againstgetBuiltinModels('github-copilot')):Picking the first match in listing order instead would silently drop thinking blocks and cache control on Anthropic models.
2.
offmust not be offered onopenai-responses. Copilot rejects such a turn outright:observed on
grok-4.6. This agrees with pi-ai's curated catalog, where all 14openai-responsesCopilot entries pinthinkingLevelMap.offtonull, while the Anthropic and Completions entries leave it open. The listing itself never mentionsoffamongreasoning_effort, so this cannot be derived from the payload — it has to be encoded as a rule.What a working version looks like
The local plugin reuses
PiAiAdapterand pi-ai's owngithub-copilotprovider (OAuth exchange, all three protocols, andfilterModelsentitlement all unchanged), and replaces exactly one member:getModels. Precedence is curated-first — the pi-ai entry wins for every id both describe, because it carries compat quirks and thinking-level spellings the listing cannot express (claude-opus-4.6carriesforceAdaptiveThinkingand amaxeffort the listing does not report). Discovery only contributes ids the catalog has never described, so the route is only ever widened; an unreachable endpoint or a missing credential costs the new ids and nothing else.Result on my account: 22 models from the catalog alone, 23 with live discovery (
grok-4.6is the addition), and a real streaming turn through that discovered model completes normally.Two implementation notes that may be useful regardless of what upstream decides:
credentialsininject. Reading it throughctx.get('credentials')at mount time races plugin registration; the read falls back to the launch environment, discovery fails, and — because discovery is designed to degrade rather than throw — the route silently keeps serving the static catalog. The symptom is "one model missing", with no error anywhere.PiAiAdapterOptions.profilesbeing a callback memoized by map identity is what makes hot refresh possible at all: swapping in a new map is enough for the next call to see new models, with no restart. That seam is doing more work than its docs suggest.Suggestion
Either relax the catalog short-circuit for providers whose listing is rich enough (Copilot is the clear case, and it could be a provider-level capability flag rather than a special case), or document the merge rules — curated-first precedence, protocol derivation, the
off-on-Responses restriction — so that plugins can do this safely without rediscovering the 400 the hard way.Environment: Windows 11, source checkout at
0.1.0-rc.5,@earendil-works/pi-ai0.84.2.All reactions