Skip to content

v1.2.0

Choose a tag to compare

@wenyinos wenyinos released this 27 Sep 08:37
· 16 commits to main since this release

Serial turn queueing, GLM-5.3, and self-updating vision routing

Proxy hardening (#2, thanks @ming-14)

  • DevEco API compatibility — role: "developer" is downgraded to system, max_completion_tokens is translated to max_tokens (DevEco ignored the former, so the cap never applied and the non-streaming gateway dropped long generations mid-flight), and reasoning_effort: none/off becomes thinking: {type: "disabled"}.
  • Crash safety — uncaughtException / unhandledRejection guards on the standalone proxy, socket error listeners on every request and on the login callback server, a .catch() on the login chain, write guards for clients that already went away, and client disconnects now cancel the upstream turn instead of draining a dead pipe.
  • Serial turn queueing — DevEco throttles bursts per account, so upstream generations run one at a time by default and latecomers queue in arrival order. DEVECO_MAX_CONCURRENCY (default 1), DEVECO_QUEUE_COOLDOWN_SEC (default 1s pause before a queued turn starts, 0 disables), DEVECO_MAX_QUEUE (default 3; beyond that the caller gets a 429 rate_limit_error instead of an endless backlog). Metadata endpoints never queue.
  • Process supervision — npm start runs the proxy under dist/daemon.js, which relaunches it with capped exponential backoff (1s → 30s, reset after 60s of uptime); the Windows start/stop scripts drive the supervisor.
  • Defects fixed on top of the PR before merging: a concurrency-slot leak (a synchronous throw while building the abort signal could leave the only slot unreleased and wedge the queue for good), an idle timeout that was misreported as a client disconnect, and two Windows script problems (unquoted entry path, stop leaving an orphan proxy behind on a non-default port).

GLM-5.3

  • DevEco now advertises GLM-5.3 (input_modalities: ["text"], 170k context, structured tool calls). The model list is dynamic, so it appears in clients with no change; it is now covered by the vision fallback as well.

Vision routing updates itself

  • Which models are text-only is now derived from the upstream model config's input_modalities (cached together with the model list for an hour) rather than a hardcoded list, so the upstream adding or re-classifying a model needs no code change here. Precedence: explicit DEVECO_TEXT_ONLY_MODELS → upstream metadata → built-in fallback while no config is cached (cold start, offline).
  • GLM-5.1 / GLM-5.3 image turns are still transparently rerouted to Qwen3_VL_235B_A22B_Instruct; the reroute drops tools / tool_choice because that model emits no structured tool calls.

Verification: 125 tests passing (npm test), npm run typecheck / lint / build clean, plus end-to-end checks against the live backend (image reroute on cold start and after the model config is cached, tool calls, streaming, Anthropic endpoint, history-image stripping).