ThinkWatch Core 0.54.0
This release records when each response's first token arrives and how fast the response is generated, and corrects the generation rate that GET /live reports.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 27. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. ThinkWatch Lite 2026.9.21 includes 0.53.0 (protocol 26) and does not connect to 0.54.0: a server used with it stays on 0.53.0 until the app is updated to a release that includes 0.54.0.sudo twcore upgrade --version 0.53.0 --restartswitches a server back to 0.53.0. - The format of the request history changed (request store schema 22, from 21). On its first start, 0.54.0 replaces the request history and the stored request and response bodies of an earlier version with an empty store; the configuration is kept. Switching back to an earlier version empties it again.
- Changed in the protocol: the new event
RequestFirstToken { id, ttft_ms },RequestFinished.tokens_per_sec,HistoryRow.ttft_msandHistoryRow.tokens_per_sec, and the new endpointsGET /token-rateandGET /token-rate/provider, which returnTokenRateView { model, p50, samples }.GET /latencyandGET /latency/providernow report the time to the first token. No message code changed.
First token. The headers of a streamed response usually come back as soon as the upstream receives the request. The upstream then queues the request and reads the prompt before it produces anything, and the time a user waits is the time to the first token, not to the headers. The gateway now reads each successful streamed response in the upstream's format and reports the moment the first text, reasoning or tool call arrives; opening frames such as message_start and response.created do not count. A response that is not streamed arrives whole and has no first token. GET /latency and GET /latency/provider now give percentiles of this time, so responses that are not streamed are left out of them. url-test groups still pick the upstream with the fastest response headers.
Generation speed. Each finished streamed request has a generation speed: the output produced after the first token, divided by the time from the first token to the end.
- When the model reasoned before the first token without streaming its reasoning, those reasoning tokens were generated outside that time. They are subtracted when the upstream reports how many there were, as OpenAI and Gemini do. When it does not, as with Anthropic, the request has no speed rather than one several times too high.
- Reasoning that is streamed as it happens is counted.
- Responses that are not streamed, cancelled and failed requests, and generations shorter than half a second have no speed.
GET /token-rate and GET /token-rate/provider give the median speed by model and by upstream, with the number of samples.
The live rate. GET /live used to divide each request's output by the time after its response headers. A response that is not streamed gets its headers only when generation is done, so that time was a few milliseconds, and one such request could push the rate to tens of thousands of tokens per second. The rate now combines the speeds of the requests that have one, weighted by their generation time.
Downloads
| Platform | Binary | Archive for server installation |
|---|---|---|
| Linux, x86_64 | twcore-x86_64-unknown-linux-gnu |
twcore-x86_64-unknown-linux-gnu.tar.gz |
| Linux, aarch64 | twcore-aarch64-unknown-linux-gnu |
twcore-aarch64-unknown-linux-gnu.tar.gz |
| macOS, Apple silicon | twcore-aarch64-apple-darwin |
— |
| Windows, x64 | twcore-x86_64-pc-windows-msvc.exe |
— |
| Windows, ARM64 | twcore-aarch64-pc-windows-msvc.exe |
— |
Each file is published with a .sha256 file beside it. A Linux archive contains twcore, the systemd unit twcore.service and LICENSE. ThinkWatch Lite includes its own copy of twcore; the files here are for running core separately, such as on a server.
Server installation
On Linux (x86_64 or aarch64), the install script sets up twcore as a systemd service. This installs 0.54.0:
curl -fsSL https://raw.githubusercontent.com/ThinkWatchProject/ThinkWatch-Core/main/scripts/install.sh | sudo sh -s -- --version 0.54.0An installation made with the script switches to 0.54.0 with:
sudo twcore upgrade --version 0.54.0 --restartConfiguration, the remote control port and connecting ThinkWatch Lite are described in docs/server.md.
Verifying a download
A .sha256 file holds the SHA-256 of the file followed by its name. With both files in the current directory, on Linux:
sha256sum -c twcore-x86_64-unknown-linux-gnu.tar.gz.sha256On macOS:
shasum -a 256 -c twcore-aarch64-apple-darwin.sha256On Windows, in PowerShell, the following prints True when the binary matches:
(Get-FileHash .\twcore-x86_64-pc-windows-msvc.exe).Hash -eq (Get-Content .\twcore-x86_64-pc-windows-msvc.exe.sha256).Split()[0]The install script and twcore upgrade check the SHA-256 themselves.
What's Changed
- Record when the first token arrives and how fast a response is generated by @fylorn in #220
- chore: v0.54.0 by @fylorn in #221
Full Changelog: v0.53.0...v0.54.0