Skip to content

Releases: wenyinos/opencode-deveco

v1.3.0

Choose a tag to compare

@wenyinos wenyinos released this 27 Sep 14:21

Reasoning on its own channel, switchable tiers, and upstream-parity gates

Reasoning stops leaking into the answer

  • DevEco's GLM models write their thinking inline in content, closed by a stray </think> with no opening tag, so clients rendered the model talking to itself as part of the reply — and the pollution accumulated over turns. The proxy now asks the upstream to report reasoning on its own channel: reasoning_content on the OpenAI endpoint, a standard thinking block (and thinking_delta) on the Anthropic one. Assistant messages that still carry a </think> scratchpad are cleaned before forwarding.
  • Thinking is not switched off, and the level is untouched: GLM models cannot stop thinking upstream (enable_thinking:false and thinking:{type:"disabled"} were both measured to have no effect), and reasoning_effort — including the Anthropic thinking-budget mapping — still passes through and takes effect.

Thinking tiers follow the cloud declaration

  • The tiers the cloud publishes per model (GLM-5.3: low/high/max, default high) are parsed and exposed as opencode model variants (Ctrl+T), with the declared default applied to every request — the same projection the official client builds.
  • A tier from outside that list is snapped onto the nearest declared one, ties going to the stronger tier. Measured on GLM-5.3: medium used to land on the heaviest tier (352 chars of reasoning against 103 for high); it now lands on high. A model that declares no tiers (GLM-5.1) is left alone, exactly as the official client leaves it.

The model list reaches opencode by itself

  • A successful fetch is persisted to <config>/opencode-deveco/models.json (0600, atomic write) and injected by the config hook, so the model picker lists every cloud model with its context limit, tiers and modalities — no hand-written opencode.json entry needed.
  • Models already written into opencode.json are merged: your fields win, cloud metadata fills the rest. Sizes the cloud omits fall back to 32768/8192 (opencode would otherwise treat the model as 0-context), and modalities opencode does not recognise are dropped.
  • small_model now follows the cloud's task_default_model_map (currently GLM-5.3) unless you set your own.

Region, account and real-name gates

  • A European timezone — including the Russian/Central-Asian zones the upstream groups with it — answers 451 with the official wording, before any login or token refresh can start. Metadata endpoints are unaffected, as upstream.
  • The account region is no longer hardcoded: a non-China nationalCode gets 403.
  • An account without HUAWEI real-name verification is told to complete it and retry, instead of failing upstream. The check runs only while the account is unverified (30s verdict cache, concurrent turns share one round-trip), so a verified account pays nothing extra.

Race guards and request hygiene

  • Signing out (or cancelling a login) supersedes a login exchange still in flight, so tokens can no longer be written back after a logout; a refresh re-checks that the stored jwtToken has not changed before installing the session.
  • Inference requests carry X-DevEco-Improvement-Enabled (DEVECO_TOOL_IMPROVEMENT=0 turns it off), and login/token requests send a browser User-Agent and detect gateway HTML error pages, failing with a readable reason instead of Unexpected token '<'.

Packaging and CI

  • The project is no longer published to npm — install from source: git clone, npm ci, npm run build. The READMEs and the project page were updated accordingly, and package-lock.json now resolves from the public registry instead of a China mirror, so CI installs are reproducible anywhere.
  • CI added (.github/workflows/ci.yml): every push to main, every pull request and manual dispatch runs type check, lint, build and the full 155-test suite on Node 20 and Node 22.
  • Node 20+ is now required (engines.node) — Node 18 is EOL and the test toolchain cannot even start there.

Verification: 155 tests passing (npm test), typecheck / lint / build clean, plus live checks: reasoning separated on all four protocol/stream combinations, medium snapping to the high tier, model catalog persisted with the cloud's real limits and tiers, 451 from a TZ=Europe/Berlin instance on both protocols, and normal traffic unaffected.

v1.2.0

Choose a tag to compare

@wenyinos wenyinos released this 27 Sep 08:37

Serial turn queueing, GLM-5.3, and self-updating vision routing

Proxy hardening (#2, thanks @ming-14)

  • DevEco API compatibility — role: "developer" is downgraded to system, max_completion_tokens is translated to max_tokens (DevEco ignored the former, so the cap never applied and the non-streaming gateway dropped long generations mid-flight), and reasoning_effort: none/off becomes thinking: {type: "disabled"}.
  • Crash safety — uncaughtException / unhandledRejection guards on the standalone proxy, socket error listeners on every request and on the login callback server, a .catch() on the login chain, write guards for clients that already went away, and client disconnects now cancel the upstream turn instead of draining a dead pipe.
  • Serial turn queueing — DevEco throttles bursts per account, so upstream generations run one at a time by default and latecomers queue in arrival order. DEVECO_MAX_CONCURRENCY (default 1), DEVECO_QUEUE_COOLDOWN_SEC (default 1s pause before a queued turn starts, 0 disables), DEVECO_MAX_QUEUE (default 3; beyond that the caller gets a 429 rate_limit_error instead of an endless backlog). Metadata endpoints never queue.
  • Process supervision — npm start runs the proxy under dist/daemon.js, which relaunches it with capped exponential backoff (1s → 30s, reset after 60s of uptime); the Windows start/stop scripts drive the supervisor.
  • Defects fixed on top of the PR before merging: a concurrency-slot leak (a synchronous throw while building the abort signal could leave the only slot unreleased and wedge the queue for good), an idle timeout that was misreported as a client disconnect, and two Windows script problems (unquoted entry path, stop leaving an orphan proxy behind on a non-default port).

GLM-5.3

  • DevEco now advertises GLM-5.3 (input_modalities: ["text"], 170k context, structured tool calls). The model list is dynamic, so it appears in clients with no change; it is now covered by the vision fallback as well.

Vision routing updates itself

  • Which models are text-only is now derived from the upstream model config's input_modalities (cached together with the model list for an hour) rather than a hardcoded list, so the upstream adding or re-classifying a model needs no code change here. Precedence: explicit DEVECO_TEXT_ONLY_MODELS → upstream metadata → built-in fallback while no config is cached (cold start, offline).
  • GLM-5.1 / GLM-5.3 image turns are still transparently rerouted to Qwen3_VL_235B_A22B_Instruct; the reroute drops tools / tool_choice because that model emits no structured tool calls.

Verification: 125 tests passing (npm test), npm run typecheck / lint / build clean, plus end-to-end checks against the live backend (image reroute on cold start and after the model config is cached, tool calls, streaming, Anthropic endpoint, history-image stripping).

v1.1.1

Choose a tag to compare

@wenyinos wenyinos released this 19 Aug 16:29

npm package now ships the autostart scripts + synced README

  • files whitelist now includes scripts/ and README_zh.md — npm install opencode-deveco gives you the systemd / launchd / Windows autostart files at node_modules/opencode-deveco/scripts/ and both language docs.
  • README updated: npm install path ($(npm root -g)/opencode-deveco), autostart script source for npm vs source installs, session-key anchoring on the first user message, honest /v2/status semantics, disconnect handling, 128 MB request-body limit, verified on opencode 1.18.x.

See v1.1.0 for the underlying proxy defect fixes.

v1.1.0

Choose a tag to compare

@wenyinos wenyinos released this 19 Aug 16:29

Proxy chain defect fixes

All defects found in a full audit of the proxy chain are fixed:

  • Stable Chat-Id per conversation — the session key now anchors on the conversation's first user message (not messages[0], which is the system prompt on the OpenAI wire format). opencode's volatile system prompt (current time, cwd…) no longer mints a new DevEco session every turn, which previously tripped the upstream session limit.
  • system-first mode fixed — the system prompt is picked from OpenAI messages[0] as well as the Anthropic top-level system field.
  • Graceful shutdown no longer hangs on long-lived SSE streams (5s grace, then force-close); removed unused request tracking.
  • Client disconnects cancel the upstream read loop in both OpenAI and Anthropic streaming paths (was draining the backend connection into a dead pipe).
  • Honest /v2/status — logged_in:true while credentials exist and can be silently refreshed, not only while the access token is unexpired.
  • /chat/completions requires POST, matching the Anthropic route.
  • Non-streaming usage logging keeps the whole JSON body (TAIL_KEEP now only caps SSE tails); Anthropic message_start keeps the requested model id.
  • token-store keeps the fresher legacy token when migration fails; plugin session map bounded at 500; request bodies capped at 128 MB.
  • Added proxy integration regression tests (75 total).

Verified end-to-end against real DevEco models (GLM-5.1, vision routing), opencode 1.18.x, and Claude Code.

v1.0.0

Choose a tag to compare

@wenyinos wenyinos released this 15 Aug 06:40

opencode-deveco v1.0.0

首个正式版本。本地代理把 opencode / Claude Code 桥接到华为 DevEco Code 模型服务。

功能

  • OpenAI 兼容端点 POST /v2/chat/completions(流式 / 非流式)
  • Anthropic Messages 端点 POST /anthropic/v1/messages(双向协议转换、工具调用、thinking、图片)
  • 华为 DevEco OAuth 登录、jwtToken 持久化与自动刷新
  • 动态模型列表 + 静态兜底,systemd / launchd / Windows 启动脚本
  • 透明识图路由:GLM-5.1 遇到图片自动转 Qwen3 VL,支持 tool_result 图片
  • 会话稳定性:默认按首条用户消息复用 Chat-Id,避免上游“新会话限流”

主要修复

  • 401 刷新后 exitSessionQueue 使用新 token
  • fetch 失败路径补调 exitSessionQueue
  • JWT 存储路径统一与旧路径自动迁移
  • realName 布尔/字符串兼容
  • OpenAI tool_choice 对象归一化
  • Anthropic 流式 input_tokens / cache_read_input_tokens
  • 流式中断输出显式错误事件
  • 独立 CLI 支持 DEVECO_PROXY_PORT
  • 未登录访问 /v2/models 不再自动弹浏览器

资源

  • 源码压缩包:opencode-deveco-1.0.0.zip