Statim 0.2.0
Statim 0.2.0 · typed decisions over any text, answered in one forward pass
GPU inference, a production-grade playground, a reproducible fine-tuning and evaluation toolkit, and the Statim brand.
Highlights
- Vulkan GPU backend: 7.7–8.9× HTTP throughput on an RTX 3070 vs. a Ryzen 7 5800X, with the same answers (parity 240/240, max |Δlogit| ≤ 1.6e-4).
- Many-option questions:
head_max_lenper request lifts Banking77 (77 intents) from 0.4885 to 0.5175 without training. - Fine-tuning: Banking77 0.4885 → 0.8655 (ECE 0.37 → 0.04) with held-out Emotion and AG News unchanged, thanks to distillation replay. Weights are reproducible with the scripts, not shipped.
- Multilingual benchmark: MASSIVE intents and sentiment in 12 languages, seeded and stratified.
- Brand identity: logo, wordmark, icons, social preview.
Added
- Vulkan GPU backend.
--device cpu|gpu|vulkan|<name>(orSTATIM_DEVICE) selects the ggml
backend; weights are copied to VRAM once. Exact f32 by default: ggml-vulkan's f16 matmul paths
moved logits by up to 0.12, so they are disabled unless--gpu-fast/STATIM_GPU_FAST=1.
Build with-DSTATIM_VULKAN=ON, which also registers the parity gates on the GPU. RTX 3070 vs.
Ryzen 7 5800X: 7.7–8.9× HTTP throughput, parity 240/240 at max |Δlogit| 8.8e-5 / 1.6e-4. - Option budget per request.
max_len/head_max_lenrequest fields (as in Laya's
predict_batch) and server defaults--max-len/--head-max-len. Many-option questions
(e.g. 77 intents) otherwise see one subword per option. - Playground redesign. Two-pane layout, live server status, model details from
/v1/models,
API key dialog, light/dark theme, examples (including a German support ticket with a separate
revenue-impact question), question templates, JSON linting, per-type result views with
confidence meters, JSON and cURL tabs, session history, keyboard shortcut. - Fine-tuning toolkit (
tools/finetune/):train_banking77.py: Laya's RLCD recipe on one task, frozen token embeddings, typed-decisions
replay and--distilllearning-without-forgetting replay.train_multitask.py: Banking77 + MASSIVE (51 languages) + multilingual sentiment (12
languages); train rows that also occur in a test split are dropped (14,121 MASSIVE and 452
sentiment rows, per language inbench/results/multitask_train_test_overlap.json); per-epoch
task budgets, LR warmup and EMA weights.merge.py: model soup, task arithmetic and TIES merging of checkpoints from one base.build_mixture.py: licence-clean training mixture fromtasksource/tasksource-jev-typed-decisions
(commercial rows only, per-source cap, evaluation and emotion sources excluded, exact-match
dedup against every reported test split).eval_laya.py(test suites on GPU) andeval_dev.py(model selection on validation data only).
- Multilingual benchmark
bench/eval_multilingual.py: MASSIVE intents (59 labels) and
multilingual sentiment, 12 languages each, seeded stratified samples with identical MASSIVE rows
across languages; results for the base checkpoints inbench/results/. CHANGELOG.mdanddocs/ROADMAP.md.- Brand identity (
assets/brand/): logo mark (one pass meeting a column of options, one of
them chosen), custom monoline wordmark whose i-dot repeats the decision point, light and dark
lockups, favicon, PNG icons (16–512 px) and a 1280×640 social preview. Used in the README header
and the playground.
Changed
/healthand the startup log report the actual compute device instead of a fixed"cpu".- The engine raises
max_lenwith a raisedhead_max_lenso the state keeps at least 128 tokens.
This never binds at the checkpoint defaults; CPU and GPU parity are unchanged. - The model parity test prints per-state deviations (
STATIM_VERBOSE) and can run one graph per
item (STATIM_BATCH1).
Fixed
- Mixture training: items sharing a document with the held-out mixture slice are dropped, so the
slice no longer inflates checkpoint selection (found by a pre-training audit). eval_dev.pyrefuses models trained with fewer than 500 held-out Banking77 rows.--gpu-fasthelp text understated the logit drift (up to ~0.12, not ~1e-2).- Playground: the mode selector silently overwrote the chosen model.
consensusis now a model
option, offered only when both checkpoints are loaded.
Results (weights are reproducible with the scripts; not shipped)
- Banking77 fine-tune v3 (
--distill 6000 --epochs 5), first 2,000 test rows: 0.4885 → 0.8655,
ECE 0.37 → 0.04; held-out Emotion 0.532 → 0.528 and AG News 0.938 → 0.9385 (within noise). - Multi-task checkpoint (experimental): MASSIVE macro over 12 languages 0.340 → 0.689, multilingual
sentiment 0.559 → 0.654, but Banking77 0.8275 (−3.8 vs. v3). Not yet recommended.
Links
- Changelog: https://github.com/BEKO2210/statim/blob/v0.2.0/CHANGELOG.md
- Roadmap to 1.0: https://github.com/BEKO2210/statim/blob/v0.2.0/docs/ROADMAP.md
- Pull request: #1
- All changes: v0.1.0...v0.2.0
- Brand assets: https://github.com/BEKO2210/statim/tree/v0.2.0/assets/brand
Verification: CI green (build + tokenizer/model/engine parity gates for both checkpoints); locally CPU ctest 5/5 and Vulkan ctest 9/9.