v1.2.0
Serial turn queueing, GLM-5.3, and self-updating vision routing
Proxy hardening (#2, thanks @ming-14)
- DevEco API compatibility —
role: "developer"is downgraded tosystem,max_completion_tokensis translated tomax_tokens(DevEco ignored the former, so the cap never applied and the non-streaming gateway dropped long generations mid-flight), andreasoning_effort: none/offbecomesthinking: {type: "disabled"}. - Crash safety —
uncaughtException/unhandledRejectionguards on the standalone proxy, socketerrorlisteners on every request and on the login callback server, a.catch()on the login chain, write guards for clients that already went away, and client disconnects now cancel the upstream turn instead of draining a dead pipe. - Serial turn queueing — DevEco throttles bursts per account, so upstream generations run one at a time by default and latecomers queue in arrival order.
DEVECO_MAX_CONCURRENCY(default1),DEVECO_QUEUE_COOLDOWN_SEC(default1s pause before a queued turn starts,0disables),DEVECO_MAX_QUEUE(default3; beyond that the caller gets a429 rate_limit_errorinstead of an endless backlog). Metadata endpoints never queue. - Process supervision —
npm startruns the proxy underdist/daemon.js, which relaunches it with capped exponential backoff (1s → 30s, reset after 60s of uptime); the Windows start/stop scripts drive the supervisor. - Defects fixed on top of the PR before merging: a concurrency-slot leak (a synchronous throw while building the abort signal could leave the only slot unreleased and wedge the queue for good), an idle timeout that was misreported as a client disconnect, and two Windows script problems (unquoted entry path,
stopleaving an orphan proxy behind on a non-default port).
GLM-5.3
- DevEco now advertises
GLM-5.3(input_modalities: ["text"], 170k context, structured tool calls). The model list is dynamic, so it appears in clients with no change; it is now covered by the vision fallback as well.
Vision routing updates itself
- Which models are text-only is now derived from the upstream model config's
input_modalities(cached together with the model list for an hour) rather than a hardcoded list, so the upstream adding or re-classifying a model needs no code change here. Precedence: explicitDEVECO_TEXT_ONLY_MODELS→ upstream metadata → built-in fallback while no config is cached (cold start, offline). GLM-5.1/GLM-5.3image turns are still transparently rerouted toQwen3_VL_235B_A22B_Instruct; the reroute dropstools/tool_choicebecause that model emits no structured tool calls.
Verification: 125 tests passing (npm test), npm run typecheck / lint / build clean, plus end-to-end checks against the live backend (image reroute on cold start and after the model config is cached, tool calls, streaming, Anthropic endpoint, history-image stripping).