Releases: wenyinos/opencode-deveco
Release list
v1.3.0
Reasoning on its own channel, switchable tiers, and upstream-parity gates
Reasoning stops leaking into the answer
- DevEco's GLM models write their thinking inline in
content, closed by a stray</think>with no opening tag, so clients rendered the model talking to itself as part of the reply — and the pollution accumulated over turns. The proxy now asks the upstream to report reasoning on its own channel:reasoning_contenton the OpenAI endpoint, a standardthinkingblock (andthinking_delta) on the Anthropic one. Assistant messages that still carry a</think>scratchpad are cleaned before forwarding. - Thinking is not switched off, and the level is untouched: GLM models cannot stop thinking upstream (
enable_thinking:falseandthinking:{type:"disabled"}were both measured to have no effect), andreasoning_effort— including the Anthropic thinking-budget mapping — still passes through and takes effect.
Thinking tiers follow the cloud declaration
- The tiers the cloud publishes per model (GLM-5.3:
low/high/max, defaulthigh) are parsed and exposed as opencode model variants (Ctrl+T), with the declared default applied to every request — the same projection the official client builds. - A tier from outside that list is snapped onto the nearest declared one, ties going to the stronger tier. Measured on GLM-5.3:
mediumused to land on the heaviest tier (352 chars of reasoning against 103 forhigh); it now lands onhigh. A model that declares no tiers (GLM-5.1) is left alone, exactly as the official client leaves it.
The model list reaches opencode by itself
- A successful fetch is persisted to
<config>/opencode-deveco/models.json(0600, atomic write) and injected by the config hook, so the model picker lists every cloud model with its context limit, tiers and modalities — no hand-writtenopencode.jsonentry needed. - Models already written into
opencode.jsonare merged: your fields win, cloud metadata fills the rest. Sizes the cloud omits fall back to 32768/8192 (opencode would otherwise treat the model as 0-context), and modalities opencode does not recognise are dropped. small_modelnow follows the cloud'stask_default_model_map(currently GLM-5.3) unless you set your own.
Region, account and real-name gates
- A European timezone — including the Russian/Central-Asian zones the upstream groups with it — answers
451with the official wording, before any login or token refresh can start. Metadata endpoints are unaffected, as upstream. - The account region is no longer hardcoded: a non-China
nationalCodegets403. - An account without HUAWEI real-name verification is told to complete it and retry, instead of failing upstream. The check runs only while the account is unverified (30s verdict cache, concurrent turns share one round-trip), so a verified account pays nothing extra.
Race guards and request hygiene
- Signing out (or cancelling a login) supersedes a login exchange still in flight, so tokens can no longer be written back after a logout; a refresh re-checks that the stored jwtToken has not changed before installing the session.
- Inference requests carry
X-DevEco-Improvement-Enabled(DEVECO_TOOL_IMPROVEMENT=0turns it off), and login/token requests send a browser User-Agent and detect gateway HTML error pages, failing with a readable reason instead ofUnexpected token '<'.
Packaging and CI
- The project is no longer published to npm — install from source:
git clone,npm ci,npm run build. The READMEs and the project page were updated accordingly, andpackage-lock.jsonnow resolves from the public registry instead of a China mirror, so CI installs are reproducible anywhere. - CI added (
.github/workflows/ci.yml): every push tomain, every pull request and manual dispatch runs type check, lint, build and the full 155-test suite on Node 20 and Node 22. - Node 20+ is now required (
engines.node) — Node 18 is EOL and the test toolchain cannot even start there.
Verification: 155 tests passing (npm test), typecheck / lint / build clean, plus live checks: reasoning separated on all four protocol/stream combinations, medium snapping to the high tier, model catalog persisted with the cloud's real limits and tiers, 451 from a TZ=Europe/Berlin instance on both protocols, and normal traffic unaffected.
v1.2.0
Serial turn queueing, GLM-5.3, and self-updating vision routing
Proxy hardening (#2, thanks @ming-14)
- DevEco API compatibility —
role: "developer"is downgraded tosystem,max_completion_tokensis translated tomax_tokens(DevEco ignored the former, so the cap never applied and the non-streaming gateway dropped long generations mid-flight), andreasoning_effort: none/offbecomesthinking: {type: "disabled"}. - Crash safety —
uncaughtException/unhandledRejectionguards on the standalone proxy, socketerrorlisteners on every request and on the login callback server, a.catch()on the login chain, write guards for clients that already went away, and client disconnects now cancel the upstream turn instead of draining a dead pipe. - Serial turn queueing — DevEco throttles bursts per account, so upstream generations run one at a time by default and latecomers queue in arrival order.
DEVECO_MAX_CONCURRENCY(default1),DEVECO_QUEUE_COOLDOWN_SEC(default1s pause before a queued turn starts,0disables),DEVECO_MAX_QUEUE(default3; beyond that the caller gets a429 rate_limit_errorinstead of an endless backlog). Metadata endpoints never queue. - Process supervision —
npm startruns the proxy underdist/daemon.js, which relaunches it with capped exponential backoff (1s → 30s, reset after 60s of uptime); the Windows start/stop scripts drive the supervisor. - Defects fixed on top of the PR before merging: a concurrency-slot leak (a synchronous throw while building the abort signal could leave the only slot unreleased and wedge the queue for good), an idle timeout that was misreported as a client disconnect, and two Windows script problems (unquoted entry path,
stopleaving an orphan proxy behind on a non-default port).
GLM-5.3
- DevEco now advertises
GLM-5.3(input_modalities: ["text"], 170k context, structured tool calls). The model list is dynamic, so it appears in clients with no change; it is now covered by the vision fallback as well.
Vision routing updates itself
- Which models are text-only is now derived from the upstream model config's
input_modalities(cached together with the model list for an hour) rather than a hardcoded list, so the upstream adding or re-classifying a model needs no code change here. Precedence: explicitDEVECO_TEXT_ONLY_MODELS→ upstream metadata → built-in fallback while no config is cached (cold start, offline). GLM-5.1/GLM-5.3image turns are still transparently rerouted toQwen3_VL_235B_A22B_Instruct; the reroute dropstools/tool_choicebecause that model emits no structured tool calls.
Verification: 125 tests passing (npm test), npm run typecheck / lint / build clean, plus end-to-end checks against the live backend (image reroute on cold start and after the model config is cached, tool calls, streaming, Anthropic endpoint, history-image stripping).
v1.1.1
npm package now ships the autostart scripts + synced README
fileswhitelist now includesscripts/andREADME_zh.md—npm install opencode-devecogives you the systemd / launchd / Windows autostart files atnode_modules/opencode-deveco/scripts/and both language docs.- README updated: npm install path (
$(npm root -g)/opencode-deveco), autostart script source for npm vs source installs, session-key anchoring on the first user message, honest/v2/statussemantics, disconnect handling, 128 MB request-body limit, verified on opencode 1.18.x.
See v1.1.0 for the underlying proxy defect fixes.
v1.1.0
Proxy chain defect fixes
All defects found in a full audit of the proxy chain are fixed:
- Stable Chat-Id per conversation — the session key now anchors on the conversation's first user message (not
messages[0], which is the system prompt on the OpenAI wire format). opencode's volatile system prompt (current time, cwd…) no longer mints a new DevEco session every turn, which previously tripped the upstream session limit. system-firstmode fixed — the system prompt is picked from OpenAImessages[0]as well as the Anthropic top-levelsystemfield.- Graceful shutdown no longer hangs on long-lived SSE streams (5s grace, then force-close); removed unused request tracking.
- Client disconnects cancel the upstream read loop in both OpenAI and Anthropic streaming paths (was draining the backend connection into a dead pipe).
- Honest
/v2/status—logged_in:truewhile credentials exist and can be silently refreshed, not only while the access token is unexpired. /chat/completionsrequires POST, matching the Anthropic route.- Non-streaming usage logging keeps the whole JSON body (
TAIL_KEEPnow only caps SSE tails); Anthropicmessage_startkeeps the requested model id. - token-store keeps the fresher legacy token when migration fails; plugin session map bounded at 500; request bodies capped at 128 MB.
- Added proxy integration regression tests (75 total).
Verified end-to-end against real DevEco models (GLM-5.1, vision routing), opencode 1.18.x, and Claude Code.
v1.0.0
opencode-deveco v1.0.0
首个正式版本。本地代理把 opencode / Claude Code 桥接到华为 DevEco Code 模型服务。
功能
- OpenAI 兼容端点
POST /v2/chat/completions(流式 / 非流式) - Anthropic Messages 端点
POST /anthropic/v1/messages(双向协议转换、工具调用、thinking、图片) - 华为 DevEco OAuth 登录、jwtToken 持久化与自动刷新
- 动态模型列表 + 静态兜底,systemd / launchd / Windows 启动脚本
- 透明识图路由:GLM-5.1 遇到图片自动转 Qwen3 VL,支持
tool_result图片 - 会话稳定性:默认按首条用户消息复用 Chat-Id,避免上游“新会话限流”
主要修复
- 401 刷新后 exitSessionQueue 使用新 token
- fetch 失败路径补调 exitSessionQueue
- JWT 存储路径统一与旧路径自动迁移
- realName 布尔/字符串兼容
- OpenAI
tool_choice对象归一化 - Anthropic 流式
input_tokens/cache_read_input_tokens - 流式中断输出显式错误事件
- 独立 CLI 支持
DEVECO_PROXY_PORT - 未登录访问
/v2/models不再自动弹浏览器
资源
- 源码压缩包:
opencode-deveco-1.0.0.zip