Skip to content

ThinkWatch Core 0.56.0

Latest

Choose a tag to compare

@github-actions github-actions released this 29 Sep 16:31
· 2 commits to main since this release
f09d002

This release keeps a conversation on the upstream that holds its prompt cache and decides a turn's route once, moves a request to the next upstream when a stream fails before its first content, sets an upstream aside for as long as the reason it gives calls for, answers token counts that no upstream can serve with a local estimate, and resends a request whose reasoning another account sealed.

Upgrade notes

  • The control-plane protocol version (CONTROL_API_VERSION) is now 30. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.24 includes 0.55.2 (protocol 28) and does not connect to 0.56.0: a server used with it stays on 0.55.2 until the app is updated to a release that includes 0.56.0. sudo twcore upgrade --version 0.55.2 --restart switches a server back to 0.55.2.
  • The request store's schema is unchanged (22): the request history is kept.
  • groups[].session_affinity is removed. A configuration that still sets it (only session_affinity: false was ever written) is refused; delete the line.
  • Changed in the protocol:
    • Removed: GroupView.session_affinity, GroupView.hurts_cache, GroupInput.session_affinity, DryRunResult.hurts_cache.
    • RequestRouted.affinity and RoutingView.affinity (AffinityView, Stay): whether the request kept its turn's route and why it stayed with the upstream that answered before.
    • Overview.failover (FailoverView).
    • AttemptOutcome has estimated.
  • New in the configuration: the failover section (see Setting a failing upstream aside).
  • Message codes: new config.failover_range and gw.upstream.stream_opening_error; removed gw.count_tokens.bedrock_model.

A conversation stays where its cache is. Routing rules used to be evaluated for every request, and only load-balance groups kept a session on one upstream. A rule on input_tokens or on images could therefore switch upstreams in the middle of an agent's turn, and a failover moved the conversation away from the upstream holding its prompt cache for good.

  • A conversation is recognised by the client's session header (x-claude-code-session-id, session_id, conversation_id and a few others) together with the content fingerprint, or by the fingerprint alone. A turn lasts from a user message until the next one; tool results belong to the turn.
  • The routing decision made at the start of a turn (rule, group, rewrites) holds for the rest of the turn. It is made again when the input approaches the smallest context window among the candidate models, where the price data gives one.
  • The next request goes first to the upstream that last answered the conversation, including one that took over after a failover: always within a turn, and across turns while its last answer read or wrote at least 1024 cached tokens less than five minutes ago. An upstream that is set aside, not among the candidates or in another group is not held to.
  • load-balance takes turns between new conversations only. session_affinity and the warnings that a group lowers the cache hit rate are gone, since no strategy drops a warm cache in the middle of a conversation any more.

Errors at the start of a stream fail over. An upstream can answer 200, open a stream and report an error as its first event: Anthropic's overloaded_error, the Codex backend's response.failed when the usage limit is reached, Bedrock's throttlingException, an error chunk from a Chat or Gemini upstream. The client had received nothing yet, and the error still reached it. A streamed answer is now held until its first content arrives, and an error before that point moves the request to the next candidate, as an error status does. What was read while waiting is passed on unchanged. The wait ends after failover.stream_start_wait_secs (15) or 1 MiB; the last candidate is not held.

Setting a failing upstream aside. Failures are sorted by the reason the upstream gives.

  • 401, 403, 402 and 404 now move the request to the next candidate. A 400 or 422 does so only when its body names an insufficient balance, a used-up quota or an unavailable model; other 4xx still go to the client unchanged. The last candidate's own 4xx reaches the client as the upstream wrote it.
  • An insufficient balance sets the upstream aside for no_balance_pause_secs (1800). A used-up quota sets it aside until the reset time the upstream names in the body, in its quota headers or in a GLM 429, and for quota_pause_secs (3600) when it names none. A rate limit with Retry-After (or Gemini's retryDelay) sets it aside for that long, at most rate_limit_max_pause_secs (3600). A missing model moves on without counting against the upstream.
  • Failures without a stated reason pause the upstream after failures_to_pause (3) in a row, for pause_secs (60), doubling with each further pause up to max_pause_secs (600) and starting over after a success.
  • A request with a single candidate is never affected, and when every candidate is set aside they are all tried.

Counting tokens. /v1/messages/count_tokens and Gemini's :countTokens routed to an upstream of another format, or to a same-format upstream that answers 404 or 405, are answered by the gateway with an estimate of the system prompt, messages, tool calls and results, and tool definitions; nothing is sent to that upstream. The answer carries x-thinkwatch-local: 1, and the request is recorded as a local answer with no cost. A count does not fail over to another format or another model. Counting through an AWS Bedrock upstream still answers 501 not_supported, after which Claude Code counts precisely itself.

Reasoning sealed by another account. Reasoning items carry content encrypted or signed for the account that produced them. After a failover from one account to another, the upstream refuses them (Responses' invalid_encrypted_content, Anthropic's invalid signature in a thinking block). The request is now sent once more without them, and the refused items are left out up front on the conversation's later turns to that upstream.

Downloads

Platform Binary Archive for server installation
Linux, x86_64 twcore-x86_64-unknown-linux-gnu twcore-x86_64-unknown-linux-gnu.tar.gz
Linux, aarch64 twcore-aarch64-unknown-linux-gnu twcore-aarch64-unknown-linux-gnu.tar.gz
macOS, Apple silicon twcore-aarch64-apple-darwin —
Windows, x64 twcore-x86_64-pc-windows-msvc.exe —
Windows, ARM64 twcore-aarch64-pc-windows-msvc.exe —

Each file is published with a .sha256 file beside it. A Linux archive contains twcore, the systemd unit twcore.service and LICENSE. ThinkWatch Lite includes its own copy of twcore; the files here are for running core separately, such as on a server.

Server installation

On Linux (x86_64 or aarch64), the install script sets up twcore as a systemd service. This installs 0.56.0:

curl -fsSL https://raw.githubusercontent.com/ThinkWatchProject/ThinkWatch-Core/main/scripts/install.sh | sudo sh -s -- --version 0.56.0

An installation made with the script switches to 0.56.0 with:

sudo twcore upgrade --version 0.56.0 --restart

Configuration, the remote control port and connecting ThinkWatch Lite are described in docs/server.md.

Verifying a download

A .sha256 file holds the SHA-256 of the file followed by its name. With both files in the current directory, on Linux:

sha256sum -c twcore-x86_64-unknown-linux-gnu.tar.gz.sha256

On macOS:

shasum -a 256 -c twcore-aarch64-apple-darwin.sha256

On Windows, in PowerShell, the following prints True when the binary matches:

(Get-FileHash .\twcore-x86_64-pc-windows-msvc.exe).Hash -eq (Get-Content .\twcore-x86_64-pc-windows-msvc.exe.sha256).Split()[0]

The install script and twcore upgrade check the SHA-256 themselves.

What's Changed

  • Keep a conversation on the upstream that answered it, and hold a turn's route by @fylorn in #237
  • Estimate count_tokens locally; resend without reasoning sealed by another account by @fylorn in #238
  • Fail over on errors at the start of a stream; pause upstreams by the reason they give by @fylorn in #239
  • chore: v0.56.0 by @fylorn in #240

Full Changelog: v0.55.2...v0.56.0