This release keeps a conversation on the upstream that holds its prompt cache and decides a turn's route once, moves a request to the next upstream when a stream fails before its first content, sets an upstream aside for as long as the reason it gives calls for, answers token counts that no upstream can serve with a local estimate, and resends a request whose reasoning another account sealed.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 30. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.24 includes 0.55.2 (protocol 28) and does not connect to 0.56.0: a server used with it stays on 0.55.2 until the app is updated to a release that includes 0.56.0.sudo twcore upgrade --version 0.55.2 --restartswitches a server back to 0.55.2. - The request store's schema is unchanged (22): the request history is kept.
groups[].session_affinityis removed. A configuration that still sets it (onlysession_affinity: falsewas ever written) is refused; delete the line.- Changed in the protocol:
- Removed:
GroupView.session_affinity,GroupView.hurts_cache,GroupInput.session_affinity,DryRunResult.hurts_cache. RequestRouted.affinityandRoutingView.affinity(AffinityView,Stay): whether the request kept its turn's route and why it stayed with the upstream that answered before.Overview.failover(FailoverView).AttemptOutcomehasestimated.
- Removed:
- New in the configuration: the
failoversection (see Setting a failing upstream aside). - Message codes: new
config.failover_rangeandgw.upstream.stream_opening_error; removedgw.count_tokens.bedrock_model.
A conversation stays where its cache is. Routing rules used to be evaluated for every request, and only load-balance groups kept a session on one upstream. A rule on input_tokens or on images could therefore switch upstreams in the middle of an agent's turn, and a failover moved the conversation away from the upstream holding its prompt cache for good.
- A conversation is recognised by the client's session header (
x-claude-code-session-id,session_id,conversation_idand a few others) together with the content fingerprint, or by the fingerprint alone. A turn lasts from a user message until the next one; tool results belong to the turn. - The routing decision made at the start of a turn (rule, group, rewrites) holds for the rest of the turn. It is made again when the input approaches the smallest context window among the candidate models, where the price data gives one.
- The next request goes first to the upstream that last answered the conversation, including one that took over after a failover: always within a turn, and across turns while its last answer read or wrote at least 1024 cached tokens less than five minutes ago. An upstream that is set aside, not among the candidates or in another group is not held to.
load-balancetakes turns between new conversations only.session_affinityand the warnings that a group lowers the cache hit rate are gone, since no strategy drops a warm cache in the middle of a conversation any more.
Errors at the start of a stream fail over. An upstream can answer 200, open a stream and report an error as its first event: Anthropic's overloaded_error, the Codex backend's response.failed when the usage limit is reached, Bedrock's throttlingException, an error chunk from a Chat or Gemini upstream. The client had received nothing yet, and the error still reached it. A streamed answer is now held until its first content arrives, and an error before that point moves the request to the next candidate, as an error status does. What was read while waiting is passed on unchanged. The wait ends after failover.stream_start_wait_secs (15) or 1 MiB; the last candidate is not held.
Setting a failing upstream aside. Failures are sorted by the reason the upstream gives.
- 401, 403, 402 and 404 now move the request to the next candidate. A 400 or 422 does so only when its body names an insufficient balance, a used-up quota or an unavailable model; other 4xx still go to the client unchanged. The last candidate's own 4xx reaches the client as the upstream wrote it.
- An insufficient balance sets the upstream aside for
no_balance_pause_secs(1800). A used-up quota sets it aside until the reset time the upstream names in the body, in its quota headers or in a GLM 429, and forquota_pause_secs(3600) when it names none. A rate limit withRetry-After(or Gemini'sretryDelay) sets it aside for that long, at mostrate_limit_max_pause_secs(3600). A missing model moves on without counting against the upstream. - Failures without a stated reason pause the upstream after
failures_to_pause(3) in a row, forpause_secs(60), doubling with each further pause up tomax_pause_secs(600) and starting over after a success. - A request with a single candidate is never affected, and when every candidate is set aside they are all tried.
Counting tokens. /v1/messages/count_tokens and Gemini's :countTokens routed to an upstream of another format, or to a same-format upstream that answers 404 or 405, are answered by the gateway with an estimate of the system prompt, messages, tool calls and results, and tool definitions; nothing is sent to that upstream. The answer carries x-thinkwatch-local: 1, and the request is recorded as a local answer with no cost. A count does not fail over to another format or another model. Counting through an AWS Bedrock upstream still answers 501 not_supported, after which Claude Code counts precisely itself.
Reasoning sealed by another account. Reasoning items carry content encrypted or signed for the account that produced them. After a failover from one account to another, the upstream refuses them (Responses' invalid_encrypted_content, Anthropic's invalid signature in a thinking block). The request is now sent once more without them, and the refused items are left out up front on the conversation's later turns to that upstream.
Downloads
| Platform | Binary | Archive for server installation |
|---|---|---|
| Linux, x86_64 | twcore-x86_64-unknown-linux-gnu |
twcore-x86_64-unknown-linux-gnu.tar.gz |
| Linux, aarch64 | twcore-aarch64-unknown-linux-gnu |
twcore-aarch64-unknown-linux-gnu.tar.gz |
| macOS, Apple silicon | twcore-aarch64-apple-darwin |
— |
| Windows, x64 | twcore-x86_64-pc-windows-msvc.exe |
— |
| Windows, ARM64 | twcore-aarch64-pc-windows-msvc.exe |
— |
Each file is published with a .sha256 file beside it. A Linux archive contains twcore, the systemd unit twcore.service and LICENSE. ThinkWatch Lite includes its own copy of twcore; the files here are for running core separately, such as on a server.
Server installation
On Linux (x86_64 or aarch64), the install script sets up twcore as a systemd service. This installs 0.56.0:
curl -fsSL https://raw.githubusercontent.com/ThinkWatchProject/ThinkWatch-Core/main/scripts/install.sh | sudo sh -s -- --version 0.56.0An installation made with the script switches to 0.56.0 with:
sudo twcore upgrade --version 0.56.0 --restartConfiguration, the remote control port and connecting ThinkWatch Lite are described in docs/server.md.
Verifying a download
A .sha256 file holds the SHA-256 of the file followed by its name. With both files in the current directory, on Linux:
sha256sum -c twcore-x86_64-unknown-linux-gnu.tar.gz.sha256On macOS:
shasum -a 256 -c twcore-aarch64-apple-darwin.sha256On Windows, in PowerShell, the following prints True when the binary matches:
(Get-FileHash .\twcore-x86_64-pc-windows-msvc.exe).Hash -eq (Get-Content .\twcore-x86_64-pc-windows-msvc.exe.sha256).Split()[0]The install script and twcore upgrade check the SHA-256 themselves.
What's Changed
- Keep a conversation on the upstream that answered it, and hold a turn's route by @fylorn in #237
- Estimate count_tokens locally; resend without reasoning sealed by another account by @fylorn in #238
- Fail over on errors at the start of a stream; pause upstreams by the reason they give by @fylorn in #239
- chore: v0.56.0 by @fylorn in #240
Full Changelog: v0.55.2...v0.56.0