Releases: BEKO2210/statim
Release list
Statim 0.6.2
Statim 0.6.2 · watch it decide
- Films on the site: the 60-second film and the 45-second story in a new Watch it decide section, the 30-second showreel in One pass. Every option scored. Landscape screens get 16:9, phones get 9:16; nothing loads before you press play.
- README: a ten-second animation that links to the film (GitHub does not play videos in READMEs).
- Model cards embed the film.
Site: beko2210.github.io/statim/#film · Changelog: CHANGELOG.md · All changes: v0.6.1...v0.6.2
Statim 0.6.1
Statim 0.6.1 · checked by a stranger's machine
- Independent clean-room reproduction on a fresh CPU-only VM (#16): ticket triage 0.896 and 0.915 at the 0.9 threshold exactly, typed-decisions 0.7585 and Banking77 0.9035 exactly, AG News and Emotion within 1 and 4 rows of 2,000 (PyTorch reference on CPU instead of CUDA). Report: docs/reproductions/clean-room.md.
REPRODUCE.mdfixes from that run: clone first, v0.5.2 archive and model checksums, ggml submodule,hf downloadinstead ofhuggingface-cli(which exits 1 inhuggingface_hub1.x), 10 CPU tests.- The
securitytests pass when run as root (Docker, cloud sandboxes); the unreadable-key-file check is skipped only for root. tools/finetune/data_licenses.pylists mixture v6 sources and synthetic gap data for the next model'sDATA_LICENSES.md.
Demo: huggingface.co/spaces/Beko2210/statim · Changelog: CHANGELOG.md · All changes: v0.6.0...v0.6.1
Statim 0.6.0
Statim 0.6.0 · concurrent requests can share one forward pass
- Server: opt-in micro-batching. With
--batch-window-ms 2 --max-batch 16, concurrentPOST /v1/systemonecalls that arrive within the window run as one packed batch; every client still gets its own response, deadline and request ID, and the answers match single runs within the 1e-4 parity gate. English model on an RTX 3070 (Vulkan, exact f32, long requests): +13 % throughput with 8 concurrent clients, +27 % with 16. It does not help on a CPU, so the default stays off. - Training tooling for the next models: mixture v6 builder with 121 licence-checked sources across every decision category (evaluation texts removed, flat memory), synthetic gap data v2 from a local model only (blind verification, near-duplicate removal, quality verdict), and trainer changes (per-category temperature mixing, streamed mixture, no held-out texts in distillation).
- The published models are unchanged:
statim-decide-en-large0.5.0 andstatim-decide-multilingual-base0.4.0.
Demo: huggingface.co/spaces/Beko2210/statim · Changelog: CHANGELOG.md · All changes: v0.5.2...v0.6.0
Statim 0.5.2
Statim 0.5.2 · the demo now says when a person should look
- Playground: the review threshold defaults to 0.6. Answers the model is unsure about show Needs review out of the box (the demo's German ticket splits urgency between "soon" and "critical" at confidence 0.11); the setting can still be switched off.
docs/API.mdstates that the checkpoint'saction.act_probabilitysaturates at 1.0 and thatmin_confidence/escalateis the abstain mechanism.
Demo: huggingface.co/spaces/Beko2210/statim · Changelog: CHANGELOG.md · All changes: v0.5.1...v0.5.2
Statim 0.5.1
Statim 0.5.1 · check every published number yourself
- REPRODUCE.md: commands and expected values for the ten-minute ticket-triage run, the parity gates, the full gate evaluation of a published model against its published
eval.json, the comparison with the base checkpoint, and speed. gate.pyruns on CPU-only machines (STATIM_BIN,STATIM_GATE_DEVICE,STATIM_GATE_PORT,STATIM_CONVERT_PY).- Website hero shows the 0.5.0 English model's real answers; the demo Space follows the current release.
Models: statim-decide-en-large · statim-decide-multilingual-base · Demo: huggingface.co/spaces/Beko2210/statim · Changelog: CHANGELOG.md · All changes: v0.5.0...v0.5.1
Statim 0.5.0
Statim 0.5.0 · the first English Statim Decide model, published weights, a public demo, CUDA
Highlights
- statim-decide-en-large (ModernBERT-large, 395M): the no-harm gate reports PROMOTE against the English base checkpoint on 54 held-out suites, 11 significant gains, 0 regressions. typed-decisions 0.361 → 0.768, on par with the best published result (meraGPT 0.768); Banking77 0.550 → 0.928; MASSIVE English 0.533 → 0.867; HWU64 0.607 → 0.833. Suites never trained on stay within noise.
- Published weights on Hugging Face, f32 and q8_0 GGUF plus the checkpoint, with model cards generated from the gate evaluation: statim-decide-en-large · statim-decide-multilingual-base. The API reports these names.
- Live demo, no install and no key: huggingface.co/spaces/Beko2210/statim. The playground is rebuilt as forms: text on one side, questions as cards, each answer inside its question, a review threshold, a developer view with cURL and Python, four worked examples including German. Mobile-safe.
- CUDA backend (
-DSTATIM_CUDA=ON) with exact f32 parity gates (4/4). Measured on an RTX 3070: exact f32 equally fast on Vulkan and CUDA;--gpu-fast16.4 req/s on Vulkan, q8_0 on CUDA 15.7 req/s for the English model. - Selective prediction:
min_confidencemarks answers below a threshold withescalate: true, in the server, the docs and both SDKs. On Banking77, keeping the 90 % most confident answers raises accuracy from 91.5 % to 95.6 %. - Licensing: weights may now also be used under PolyForm Small Business 1.0.0 (free commercial use for companies below 100 people and 1 M USD revenue) and PolyForm Free Trial 1.0.0 (32-day evaluation), next to PolyForm Noncommercial and the commercial licence.
- Ticket triage example with real Banking77 tickets (examples/ticket-triage): 0.896 on a seeded 500-ticket sample, 0.915 on the 96 % answered at a 0.9 review threshold.
- Synthetic-data pipeline that uses only a local Apache-2.0 generator, with a quality report and a scale-up verdict; 8-bit AdamW for large encoders on 8 GB GPUs.
- Correction: the 0.4.0 model scores 0.9315 on AG News zero-shot (the 0.4.0 notes said 0.9385).
Links
- Changelog: https://github.com/BEKO2210/statim/blob/v0.5.0/CHANGELOG.md
- Models: statim-decide-en-large · statim-decide-multilingual-base · Licences: LICENSE-MODEL.md · COMMERCIAL.md
- Demo: https://huggingface.co/spaces/Beko2210/statim · Website: https://beko2210.github.io/statim/
- API: docs/API.md · OpenAPI · Data licences: DATA_LICENSES.md
- Roadmap: docs/ROADMAP.md · Pull request: #7 · All changes: v0.4.0...v0.5.0
Verification: CI green (build, parity gates, security suites, site QA); locally the CPU suite plus the Vulkan and CUDA parity gates; gate PROMOTE with 0 regressions on 54 held-out suites; Hugging Face checksums verified against the local SHA256SUMS.
Statim 0.4.0
Statim 0.4.0 · typed decisions above Jev, official SDKs, binaries on every release
Highlights
- typed-decisions 0.7585: above Jev (0.727) and the dataset's teacher agreement (0.735). MASSIVE over 12 languages 0.772, Banking77 0.903. Trained only on licence-audited data.
- Stricter no-harm gate: suite families are pooled, so slow drifts across many small suites are caught.
- Official SDKs: Python (
clients/python) and TypeScript (clients/js), no runtime dependencies. - Downloads: Linux x86-64 binaries for CPU and Vulkan GPU are attached to this release by CI, with SHA-256 checksums. GPU container:
Dockerfile.vulkan. Production guide: docs/DEPLOY.md.
Results
- 0.4.0 model (multilingual checkpoint,
train_multitask.py --cleanon mixture v5: 163 audited
tasksource sources plus Nemotron-Safety, IndicGuard, MINDS-14 and SNIPS, MASSIVE 2,000 rows per
language, 20 epochs with early stopping; best epoch 19). Against 0.3.0 the gate reports
PROMOTE: typed-decisions 0.6905 → 0.7585 (above Jev's 0.727 and the dataset's teacher agreement
0.735), MASSIVE over 12 languages 0.733 → 0.772, Banking77 0.891 → 0.903, HWU64 0.760 → 0.820,
Belebele 0.273 → 0.310. Zero-shot suites pooled: −0.3 points (rows) / −1.6 (suites), both within
two standard errors; the next run strengthens distillation to reverse that trend. Against the
base checkpoint: trained tasks +41 points, sentiment +3.4 (significant), zero-shot within noise.
Chart:assets/diagrams/results-0.4.0-*.svg. Weights are not yet published. - Context length was measured, not assumed: 13.4 % of training items exceed 512 tokens, 2.4 % 1,024
and 0.2 % 2,048, and no evaluation item exceeds 1,024. typed-decisions: 0.756 at max_len 512,
0.7585 at 1,024 and at 2,048 (identical).
Added
gate.pyalso pools each suite family (trained, zero-shot, sentiment), row- and suite-weighted,
so a drift spread over many small suites counts as a regression even when no single suite is
significant.train_multitask.py --massive-langsand--max-len(per-language MASSIVE selection; context
length override saved with the model).- Official client SDKs for the HTTP API: Python package
statiminclients/python
(standard library only) and TypeScript package@statim/clientinclients/js
(fetch, no runtime dependencies). Both exposedecide,decide_batch,models,
health, andready, typed choice, score, and yes/no answers, request IDs, and
retries with backoff for HTTP 503. - GitHub release packaging (runs when a release is published) for portable Linux x86-64 CPU and Vulkan binaries, including
licence and deployment documents plus published SHA-256 checksums; model weights remain separate. - A non-root Vulkan container image with Mesa and NVIDIA Container Toolkit deployment options.
- A production deployment guide covering hardened systemd and Docker operation, TLS reverse proxying,
authenticated Prometheus scraping, health/readiness probes, and resource ceilings.
Changelog · API · Deploy · Security · Data licences · Commercial licensing · #4 · v0.3.0...v0.4.0
Statim 0.3.0
Statim 0.3.0 · better decisions in 12 languages, trained only on data you may use commercially
Highlights
- Multilingual intents 0.340 → 0.733 over 12 languages (Arabic 0.20 → 0.61, Hindi 0.27 → 0.67); Banking77 0.517 → 0.891; typed decisions 0.351 → 0.691. Zero-shot suites the model never trained on stay within noise or improve.
- No-harm gate: a new model only replaces the old one if validation improves and none of 54 held-out suites drops by more than two standard errors. This release: 16 significant gains, 0 regressions.
- Licence-clean training: every training source was audited at its origin (279 checked, 116 excluded); only permissive licences, listed in DATA_LICENSES.md.
- Licensing model: code Apache-2.0; published model weights free for noncommercial use under PolyForm Noncommercial 1.0.0, commercial use by licence (COMMERCIAL.md).
- Enterprise security hardening: 12 findings of an independent review fixed with regression tests (fail-closed auth, JSON and work limits, bounded HTTP queues, log-injection fix, cpp-httplib 0.58.0). Details in docs/SECURITY.md.
- Documentation: OpenAPI 3.1 spec and API guide; zero-shot benchmark over eight unseen suites.
Results
- The 0.3.0 model (multilingual checkpoint fine-tuned with
train_multitask.py --cleanon
licence-audited data only) against the base checkpoint, evaluated bytools/finetune/gate.pyon
54 held-out suites: MASSIVE intents over 12 languages 0.340 → 0.733 (Arabic 0.200 → 0.613,
Hindi 0.267 → 0.673), Banking77 0.5175 → 0.891, typed-decisions 0.351 → 0.6905, HWU64 (sibling
of MASSIVE, overlapping rows removed) 0.500 → 0.760; zero-shot suites never trained on stay
within noise or improve (GoEmotions +6.0, SIB-200 +2.5, SemRel +2.9, Belebele +2.7, FarsTail
−2.0, DAIR Emotion −1.5, AG News +0.1 points). 16 significant gains, 0 significant
regressions; calibration error roughly halves. Chart:assets/diagrams/results-0.3.0-*.svg.
Weights are reproducible with the scripts and not yet published.
Added
- Release results chart generated from gate evaluations (
tools/diagrams/gate_chart.py). - Updated
docs/API.mdanddocs/openapi.yamlfor the hardened server (Codex, checked against a
running server). - Licensing model: source code stays Apache-2.0; model weights published by Statim are licensed
under PolyForm Noncommercial 1.0.0 (LICENSE-MODEL.md, verbatim official text), commercial use
needs a paid licence (COMMERCIAL.md). DATA_LICENSES.md, generated bytools/finetune/data_licenses.py: base models, training data of
released weights with licences, evaluation-only data, and data excluded from releases.train_multitask.py --clean: commercial-clean training (no tyqiangz sentiment, distillation texts
from the licence-filtered mixture).
Changed
build_mixture.pykeeps only rows whose every listed licence is permissive (Apache-2.0, MIT, BSD,
CC0, CC-BY, ODC-By, AFL-3.0); ShareAlike, copyleft, custom and unknown terms are excluded.
Security
- Fail closed when any explicitly configured key file/environment value cannot supply valid
keys. Log unauthenticated local mode explicitly; require the production environment file/key. - Preflight JSON with depth, node, object-width, key-length, and duplicate-key checks before
constructing the ordered DOM, preventing deeply nested crashes and quadratic wide-object parsing. - Bound HTTP workers and pending sockets; authenticate and admit inference requests before
body reception. Reject oversized declared bodies without draining, and enforce absolute
header/body deadlines even when a peer keeps sending bytes. - Bound field sizes and aggregate batch/ensemble/calibration/consensus work, token budgets,
attention-memory estimates, and response bytes. Add cooperative inference deadlines and
systemd memory/CPU/task/file-descriptor ceilings without changing graph packing or math. - Replace the entry-count-only calibration cache with a byte-bounded LRU keyed on the validated
question only, so unknown fields (still ignored, as by laya.serve) are never retained; a test
sends 64 MiB of ignored metadata and checks that memory stays flat. - Allowlist request IDs and serialize log records as JSON to prevent log injection.
- Validate signed/unsigned token and ensemble budgets before integer narrowing, including
CLI defaults; reject effective sequence lengths beyond model capacity with 422. - Upgrade vendored cpp-httplib from 0.26.0 to 0.58.0. Explicitly reject simultaneous
Content-Length/Transfer-Encoding (including zero lengths) and duplicate Content-Length headers. - Require bearer auth for
/metricsand/v1/modelswhen configured. Keep/healthand
/readyopen; reduce/healthto status and version./v1/modelsnow also reports each
model's compute device, and the playground reads models and device from there. - Use local timestamp storage with
gmtime_r(gmtime_son Windows) for concurrent logs. - Install a global HTTP exception handler with fixed client errors and server-only details.
- Add C++, CPU model-backed, and live HTTP security regressions alongside existing parity gates.
Links
- Changelog: https://github.com/BEKO2210/statim/blob/v0.3.0/CHANGELOG.md
- Security: https://github.com/BEKO2210/statim/blob/v0.3.0/docs/SECURITY.md
- API: https://github.com/BEKO2210/statim/blob/v0.3.0/docs/API.md · OpenAPI
- Data licences: https://github.com/BEKO2210/statim/blob/v0.3.0/DATA_LICENSES.md
- Commercial licensing: https://github.com/BEKO2210/statim/blob/v0.3.0/COMMERCIAL.md
- Roadmap: https://github.com/BEKO2210/statim/blob/v0.3.0/docs/ROADMAP.md
- Pull request: #3 · All changes: v0.2.1...v0.3.0
Verification: CI green (build, parity gates, security suites); locally ctest 12/12 including GPU parity; gate PROMOTE with 0 regressions on 54 held-out suites. Model weights are reproducible with the documented commands and not yet published.
Statim 0.2.1
Statim 0.2.1 · the README now explains itself in pictures
Three diagrams in the Statim brand style, each with a light and a dark variant that follows your GitHub theme. Every number comes from the documented results in the README and CHANGELOG.
How a request becomes a decision
Prompt builder, tokenizer, one [MASK] per option, encoder, decision head, scorer and calibrated softmax. One forward pass answers every question and option.
GPU backend
RTX 3070 vs. Ryzen 7 5800X: 7.7× (multilingual) and 8.9× (English) HTTP throughput, same answers.
Fine-tuning results
Banking77 0.4885 → 0.8655 while held-out AG News and Emotion stay flat. Calibration error falls from 0.372 to 0.043.
Changed
- The ASCII sketch in "How it works" is replaced by the architecture diagram.
- Version 0.2.0 → 0.2.1.
Links
- Changelog: https://github.com/BEKO2210/statim/blob/v0.2.1/CHANGELOG.md
- Diagram sources: https://github.com/BEKO2210/statim/tree/v0.2.1/assets/diagrams
- Pull request: #2
- All changes: v0.2.0...v0.2.1
- Previous release: v0.2.0
Verification: CI green (build + parity gates). Each diagram was reviewed from renders at README width in both themes (alignment, overlap, WCAG contrast, numbers).
Statim 0.2.0
Statim 0.2.0 · typed decisions over any text, answered in one forward pass
GPU inference, a production-grade playground, a reproducible fine-tuning and evaluation toolkit, and the Statim brand.
Highlights
- Vulkan GPU backend: 7.7–8.9× HTTP throughput on an RTX 3070 vs. a Ryzen 7 5800X, with the same answers (parity 240/240, max |Δlogit| ≤ 1.6e-4).
- Many-option questions:
head_max_lenper request lifts Banking77 (77 intents) from 0.4885 to 0.5175 without training. - Fine-tuning: Banking77 0.4885 → 0.8655 (ECE 0.37 → 0.04) with held-out Emotion and AG News unchanged, thanks to distillation replay. Weights are reproducible with the scripts, not shipped.
- Multilingual benchmark: MASSIVE intents and sentiment in 12 languages, seeded and stratified.
- Brand identity: logo, wordmark, icons, social preview.
Added
- Vulkan GPU backend.
--device cpu|gpu|vulkan|<name>(orSTATIM_DEVICE) selects the ggml
backend; weights are copied to VRAM once. Exact f32 by default: ggml-vulkan's f16 matmul paths
moved logits by up to 0.12, so they are disabled unless--gpu-fast/STATIM_GPU_FAST=1.
Build with-DSTATIM_VULKAN=ON, which also registers the parity gates on the GPU. RTX 3070 vs.
Ryzen 7 5800X: 7.7–8.9× HTTP throughput, parity 240/240 at max |Δlogit| 8.8e-5 / 1.6e-4. - Option budget per request.
max_len/head_max_lenrequest fields (as in Laya's
predict_batch) and server defaults--max-len/--head-max-len. Many-option questions
(e.g. 77 intents) otherwise see one subword per option. - Playground redesign. Two-pane layout, live server status, model details from
/v1/models,
API key dialog, light/dark theme, examples (including a German support ticket with a separate
revenue-impact question), question templates, JSON linting, per-type result views with
confidence meters, JSON and cURL tabs, session history, keyboard shortcut. - Fine-tuning toolkit (
tools/finetune/):train_banking77.py: Laya's RLCD recipe on one task, frozen token embeddings, typed-decisions
replay and--distilllearning-without-forgetting replay.train_multitask.py: Banking77 + MASSIVE (51 languages) + multilingual sentiment (12
languages); train rows that also occur in a test split are dropped (14,121 MASSIVE and 452
sentiment rows, per language inbench/results/multitask_train_test_overlap.json); per-epoch
task budgets, LR warmup and EMA weights.merge.py: model soup, task arithmetic and TIES merging of checkpoints from one base.build_mixture.py: licence-clean training mixture fromtasksource/tasksource-jev-typed-decisions
(commercial rows only, per-source cap, evaluation and emotion sources excluded, exact-match
dedup against every reported test split).eval_laya.py(test suites on GPU) andeval_dev.py(model selection on validation data only).
- Multilingual benchmark
bench/eval_multilingual.py: MASSIVE intents (59 labels) and
multilingual sentiment, 12 languages each, seeded stratified samples with identical MASSIVE rows
across languages; results for the base checkpoints inbench/results/. CHANGELOG.mdanddocs/ROADMAP.md.- Brand identity (
assets/brand/): logo mark (one pass meeting a column of options, one of
them chosen), custom monoline wordmark whose i-dot repeats the decision point, light and dark
lockups, favicon, PNG icons (16–512 px) and a 1280×640 social preview. Used in the README header
and the playground.
Changed
/healthand the startup log report the actual compute device instead of a fixed"cpu".- The engine raises
max_lenwith a raisedhead_max_lenso the state keeps at least 128 tokens.
This never binds at the checkpoint defaults; CPU and GPU parity are unchanged. - The model parity test prints per-state deviations (
STATIM_VERBOSE) and can run one graph per
item (STATIM_BATCH1).
Fixed
- Mixture training: items sharing a document with the held-out mixture slice are dropped, so the
slice no longer inflates checkpoint selection (found by a pre-training audit). eval_dev.pyrefuses models trained with fewer than 500 held-out Banking77 rows.--gpu-fasthelp text understated the logit drift (up to ~0.12, not ~1e-2).- Playground: the mode selector silently overwrote the chosen model.
consensusis now a model
option, offered only when both checkpoints are loaded.
Results (weights are reproducible with the scripts; not shipped)
- Banking77 fine-tune v3 (
--distill 6000 --epochs 5), first 2,000 test rows: 0.4885 → 0.8655,
ECE 0.37 → 0.04; held-out Emotion 0.532 → 0.528 and AG News 0.938 → 0.9385 (within noise). - Multi-task checkpoint (experimental): MASSIVE macro over 12 languages 0.340 → 0.689, multilingual
sentiment 0.559 → 0.654, but Banking77 0.8275 (−3.8 vs. v3). Not yet recommended.
Links
- Changelog: https://github.com/BEKO2210/statim/blob/v0.2.0/CHANGELOG.md
- Roadmap to 1.0: https://github.com/BEKO2210/statim/blob/v0.2.0/docs/ROADMAP.md
- Pull request: #1
- All changes: v0.1.0...v0.2.0
- Brand assets: https://github.com/BEKO2210/statim/tree/v0.2.0/assets/brand
Verification: CI green (build + tokenizer/model/engine parity gates for both checkpoints); locally CPU ctest 5/5 and Vulkan ctest 9/9.