Reasoning on its own channel, switchable tiers, and upstream-parity gates
Reasoning stops leaking into the answer
- DevEco's GLM models write their thinking inline in
content, closed by a stray</think>with no opening tag, so clients rendered the model talking to itself as part of the reply — and the pollution accumulated over turns. The proxy now asks the upstream to report reasoning on its own channel:reasoning_contenton the OpenAI endpoint, a standardthinkingblock (andthinking_delta) on the Anthropic one. Assistant messages that still carry a</think>scratchpad are cleaned before forwarding. - Thinking is not switched off, and the level is untouched: GLM models cannot stop thinking upstream (
enable_thinking:falseandthinking:{type:"disabled"}were both measured to have no effect), andreasoning_effort— including the Anthropic thinking-budget mapping — still passes through and takes effect.
Thinking tiers follow the cloud declaration
- The tiers the cloud publishes per model (GLM-5.3:
low/high/max, defaulthigh) are parsed and exposed as opencode model variants (Ctrl+T), with the declared default applied to every request — the same projection the official client builds. - A tier from outside that list is snapped onto the nearest declared one, ties going to the stronger tier. Measured on GLM-5.3:
mediumused to land on the heaviest tier (352 chars of reasoning against 103 forhigh); it now lands onhigh. A model that declares no tiers (GLM-5.1) is left alone, exactly as the official client leaves it.
The model list reaches opencode by itself
- A successful fetch is persisted to
<config>/opencode-deveco/models.json(0600, atomic write) and injected by the config hook, so the model picker lists every cloud model with its context limit, tiers and modalities — no hand-writtenopencode.jsonentry needed. - Models already written into
opencode.jsonare merged: your fields win, cloud metadata fills the rest. Sizes the cloud omits fall back to 32768/8192 (opencode would otherwise treat the model as 0-context), and modalities opencode does not recognise are dropped. small_modelnow follows the cloud'stask_default_model_map(currently GLM-5.3) unless you set your own.
Region, account and real-name gates
- A European timezone — including the Russian/Central-Asian zones the upstream groups with it — answers
451with the official wording, before any login or token refresh can start. Metadata endpoints are unaffected, as upstream. - The account region is no longer hardcoded: a non-China
nationalCodegets403. - An account without HUAWEI real-name verification is told to complete it and retry, instead of failing upstream. The check runs only while the account is unverified (30s verdict cache, concurrent turns share one round-trip), so a verified account pays nothing extra.
Race guards and request hygiene
- Signing out (or cancelling a login) supersedes a login exchange still in flight, so tokens can no longer be written back after a logout; a refresh re-checks that the stored jwtToken has not changed before installing the session.
- Inference requests carry
X-DevEco-Improvement-Enabled(DEVECO_TOOL_IMPROVEMENT=0turns it off), and login/token requests send a browser User-Agent and detect gateway HTML error pages, failing with a readable reason instead ofUnexpected token '<'.
Packaging and CI
- The project is no longer published to npm — install from source:
git clone,npm ci,npm run build. The READMEs and the project page were updated accordingly, andpackage-lock.jsonnow resolves from the public registry instead of a China mirror, so CI installs are reproducible anywhere. - CI added (
.github/workflows/ci.yml): every push tomain, every pull request and manual dispatch runs type check, lint, build and the full 155-test suite on Node 20 and Node 22. - Node 20+ is now required (
engines.node) — Node 18 is EOL and the test toolchain cannot even start there.
Verification: 155 tests passing (npm test), typecheck / lint / build clean, plus live checks: reasoning separated on all four protocol/stream combinations, medium snapping to the high tier, model catalog persisted with the cloud's real limits and tiers, 451 from a TZ=Europe/Berlin instance on both protocols, and normal traffic unaffected.