ThinkWatch Core 0.57.0
This release redacts Chinese resident ID numbers and bank card numbers out of the box, searches the whole request history including the text of requests and answers, and adds an upstream check-up: whether the model named in an upstream's answers matches the one sent, how its reported input compares with other upstreams serving the same model, how much of its input comes from the prompt cache, and its time to first token and speed.
Upgrade notes
- The control-plane protocol version (
CONTROL_API_VERSION) is now 31. ThinkWatch Lite connects only to a core with the same protocol version, and has to regenerate its protocol types from this release to use it. A server used with ThinkWatch Lite 2026.9.26 or earlier (protocol 30) stays on 0.56.0 until the app is updated to a release that includes 0.57.0;sudo twcore upgrade --version 0.56.0 --restartswitches a server back. - The request store's schema goes from 22 to 23 (three new columns): the request history is cleared when the store is rebuilt on first start.
- Two redaction rules are on by default:
cn-resident-idandbank-card(see below). With outbound redaction inenforce, such numbers in a request are replaced before it leaves; inobservethey are recorded.security.redact.disableturns either off. - Routing estimates input size the way token counting does. Rules on
input_tokensnow see one token per CJK character, 1,600 per image or file, and no thinking, instead of four bytes per token with images and files left out. A request with many images can now cross a size threshold it used to stay under. - Changed in the protocol:
POST /history/search(HistorySearchQuery→HistorySearchPage, withHistoryCursor,ContentHit,ContentSide,SearchStop).GET /upstreams/health(UpstreamHealth,UpstreamCheckup,ModelConsistency,ModelPair,InputVsEstimate,InputForModel,RatioView,CacheReads,CacheForModel,CacheTally,MedianView).Event::RequestStartedhasinput_estimate;RequestFinished,RequestFailedandRequestCancelledhaveanswered_model(both left out when absent).Matcherhascn-resident-id(born_since) andbank-card(networks:CardNetwork,CardPrefix);SecretKindhaspersonal.
- No message codes change.
Chinese resident ID numbers and bank card numbers. Outbound redaction looked only for credentials. Two built-in rules now look for personal numbers, held to the same bar as the credential rules: a value must check out structurally, not merely look like one.
cn-resident-id: an 18-character resident ID number whose first two digits are a province-level code, whose birth date is a real date between 1900-01-01 and today, and whose check character (ISO 7064 MOD 11-2) is right. The old 15-digit numbers are not matched.bank-card: a number whose prefix and length belong to UnionPay, Visa, Mastercard, American Express, JCB, Discover or Diners Club and which passes the Luhn check, written in one run or in groups of four separated by single spaces or hyphens (and American Express's 4-6-5, Diners Club's 4-6-4). The public test card numbers of Stripe, Braintree and Adyen are left alone.- A number inside a longer run of letters or digits is not matched, and neither is a number written as a JSON number (a tool call's arguments, for example): replacing it would leave the body invalid JSON.
- Placeholders say what was there:
<<TW_ID_NUMBER_1>>,<<TW_CARD_NUMBER_1>>. They are restored in the answer as before, streams included. The security log shows only the last four characters. - Scanning a request body also got faster: the token scan no longer allocates per token.
Searching the whole history. The Traffic page searched only the requests it had loaded. POST /history/search applies the same filter to every stored request, in pages that go back in time.
- It matches the free text against the path, client, client hint, peer, upstream, model and error message, and takes the message codes whose translation matched, so a query in the app's own language finds errors stored in English.
- With
content, it also searches, for requests whose bodies are still kept, what is new in each request: the last user turn (text and tool results) and the answer (text, and each tool call's name and input). Clients that resend the whole conversation every turn therefore do not match every later request of a session. Each match comes with a short excerpt, with credentials masked. - Content search reads at most 500 requests' bodies or 1.5 seconds per call, newest first, and says how far back it got, so a client can offer to search further.
Upstream check-up. GET /upstreams/health reports, for every upstream over a window (the last 7 days by default), facts and sample sizes rather than verdicts:
- requests, failures and cancellations;
- whether the model named in its answers matches the model sent (after rule rewrites), with up to three example pairs. Names are compared after removing differences that are not a different model (vendor prefixes, dates and snapshot markers, Bedrock's inference-profile and version forms,
4.5against4-5), never the family, size, generation or variant; - the median ratio of the input it reported to the gateway's local estimate, overall and, for every model that other upstreams also served, next to the ratio of those other upstreams;
- for later turns of a conversation sent to it within five minutes of the previous one, the share of input read from the prompt cache and how many such turns read none, also next to other upstreams serving the same model;
- median time to first token and generation speed.
- The answered model is read from the upstream's own response in the same pass that reads usage (Bedrock's responses do not name one). The estimate is the one routing already computes.
Downloads
| Platform | Binary | Archive for server installation |
|---|---|---|
| Linux, x86_64 | twcore-x86_64-unknown-linux-gnu |
twcore-x86_64-unknown-linux-gnu.tar.gz |
| Linux, aarch64 | twcore-aarch64-unknown-linux-gnu |
twcore-aarch64-unknown-linux-gnu.tar.gz |
| macOS, Apple silicon | twcore-aarch64-apple-darwin |
— |
| Windows, x64 | twcore-x86_64-pc-windows-msvc.exe |
— |
| Windows, ARM64 | twcore-aarch64-pc-windows-msvc.exe |
— |
Each file is published with a .sha256 file beside it. A Linux archive contains twcore, the systemd unit twcore.service and LICENSE. ThinkWatch Lite includes its own copy of twcore; the files here are for running core separately, such as on a server.
Server installation
On Linux (x86_64 or aarch64), the install script sets up twcore as a systemd service. This installs 0.57.0:
curl -fsSL https://raw.githubusercontent.com/ThinkWatchProject/ThinkWatch-Core/main/scripts/install.sh | sudo sh -s -- --version 0.57.0An installation made with the script switches to 0.57.0 with:
sudo twcore upgrade --version 0.57.0 --restartConfiguration, the remote control port and connecting ThinkWatch Lite are described in docs/server.md.
Verifying a download
A .sha256 file holds the SHA-256 of the file followed by its name. With both files in the current directory, on Linux:
sha256sum -c twcore-x86_64-unknown-linux-gnu.tar.gz.sha256On macOS:
shasum -a 256 -c twcore-aarch64-apple-darwin.sha256On Windows, in PowerShell, the following prints True when the binary matches:
(Get-FileHash .\twcore-x86_64-pc-windows-msvc.exe).Hash -eq (Get-Content .\twcore-x86_64-pc-windows-msvc.exe.sha256).Split()[0]The install script and twcore upgrade check the SHA-256 themselves.
What's Changed
- Rewrite the README to be concise, with security near the top by @fylorn in #241
- README: name the relay threat that tool-call inspection guards against by @fylorn in #242
- README: the ThinkWatch Core logo at the top by @fylorn in #243
- Redact Chinese resident ID and bank card numbers out of the box by @fylorn in #244
- Search the whole request history, including request and answer text by @fylorn in #245
- Upstream check-up: model consistency, input usage vs the local estimate, cache reads per upstream by @fylorn in #246
- chore: v0.57.0 by @fylorn in #247
Full Changelog: v0.56.0...v0.57.0