Skip to content

v1.3.0

Latest

Choose a tag to compare

@wenyinos wenyinos released this 27 Sep 14:21

Reasoning on its own channel, switchable tiers, and upstream-parity gates

Reasoning stops leaking into the answer

  • DevEco's GLM models write their thinking inline in content, closed by a stray </think> with no opening tag, so clients rendered the model talking to itself as part of the reply — and the pollution accumulated over turns. The proxy now asks the upstream to report reasoning on its own channel: reasoning_content on the OpenAI endpoint, a standard thinking block (and thinking_delta) on the Anthropic one. Assistant messages that still carry a </think> scratchpad are cleaned before forwarding.
  • Thinking is not switched off, and the level is untouched: GLM models cannot stop thinking upstream (enable_thinking:false and thinking:{type:"disabled"} were both measured to have no effect), and reasoning_effort — including the Anthropic thinking-budget mapping — still passes through and takes effect.

Thinking tiers follow the cloud declaration

  • The tiers the cloud publishes per model (GLM-5.3: low/high/max, default high) are parsed and exposed as opencode model variants (Ctrl+T), with the declared default applied to every request — the same projection the official client builds.
  • A tier from outside that list is snapped onto the nearest declared one, ties going to the stronger tier. Measured on GLM-5.3: medium used to land on the heaviest tier (352 chars of reasoning against 103 for high); it now lands on high. A model that declares no tiers (GLM-5.1) is left alone, exactly as the official client leaves it.

The model list reaches opencode by itself

  • A successful fetch is persisted to <config>/opencode-deveco/models.json (0600, atomic write) and injected by the config hook, so the model picker lists every cloud model with its context limit, tiers and modalities — no hand-written opencode.json entry needed.
  • Models already written into opencode.json are merged: your fields win, cloud metadata fills the rest. Sizes the cloud omits fall back to 32768/8192 (opencode would otherwise treat the model as 0-context), and modalities opencode does not recognise are dropped.
  • small_model now follows the cloud's task_default_model_map (currently GLM-5.3) unless you set your own.

Region, account and real-name gates

  • A European timezone — including the Russian/Central-Asian zones the upstream groups with it — answers 451 with the official wording, before any login or token refresh can start. Metadata endpoints are unaffected, as upstream.
  • The account region is no longer hardcoded: a non-China nationalCode gets 403.
  • An account without HUAWEI real-name verification is told to complete it and retry, instead of failing upstream. The check runs only while the account is unverified (30s verdict cache, concurrent turns share one round-trip), so a verified account pays nothing extra.

Race guards and request hygiene

  • Signing out (or cancelling a login) supersedes a login exchange still in flight, so tokens can no longer be written back after a logout; a refresh re-checks that the stored jwtToken has not changed before installing the session.
  • Inference requests carry X-DevEco-Improvement-Enabled (DEVECO_TOOL_IMPROVEMENT=0 turns it off), and login/token requests send a browser User-Agent and detect gateway HTML error pages, failing with a readable reason instead of Unexpected token '<'.

Packaging and CI

  • The project is no longer published to npm — install from source: git clone, npm ci, npm run build. The READMEs and the project page were updated accordingly, and package-lock.json now resolves from the public registry instead of a China mirror, so CI installs are reproducible anywhere.
  • CI added (.github/workflows/ci.yml): every push to main, every pull request and manual dispatch runs type check, lint, build and the full 155-test suite on Node 20 and Node 22.
  • Node 20+ is now required (engines.node) — Node 18 is EOL and the test toolchain cannot even start there.

Verification: 155 tests passing (npm test), typecheck / lint / build clean, plus live checks: reasoning separated on all four protocol/stream combinations, medium snapping to the high tier, model catalog persisted with the cloud's real limits and tiers, 451 from a TZ=Europe/Berlin instance on both protocols, and normal traffic unaffected.