Skip to content

Statim 0.2.0

Choose a tag to compare

@BEKO2210 BEKO2210 released this 26 Sep 23:39
· 77 commits to main since this release
81967b6

Statim

Statim 0.2.0 · typed decisions over any text, answered in one forward pass

GPU inference, a production-grade playground, a reproducible fine-tuning and evaluation toolkit, and the Statim brand.

Highlights

  • Vulkan GPU backend: 7.7–8.9× HTTP throughput on an RTX 3070 vs. a Ryzen 7 5800X, with the same answers (parity 240/240, max |Δlogit| ≤ 1.6e-4).
  • Many-option questions: head_max_len per request lifts Banking77 (77 intents) from 0.4885 to 0.5175 without training.
  • Fine-tuning: Banking77 0.4885 → 0.8655 (ECE 0.37 → 0.04) with held-out Emotion and AG News unchanged, thanks to distillation replay. Weights are reproducible with the scripts, not shipped.
  • Multilingual benchmark: MASSIVE intents and sentiment in 12 languages, seeded and stratified.
  • Brand identity: logo, wordmark, icons, social preview.

Added

  • Vulkan GPU backend. --device cpu|gpu|vulkan|<name> (or STATIM_DEVICE) selects the ggml
    backend; weights are copied to VRAM once. Exact f32 by default: ggml-vulkan's f16 matmul paths
    moved logits by up to 0.12, so they are disabled unless --gpu-fast / STATIM_GPU_FAST=1.
    Build with -DSTATIM_VULKAN=ON, which also registers the parity gates on the GPU. RTX 3070 vs.
    Ryzen 7 5800X: 7.7–8.9× HTTP throughput, parity 240/240 at max |Δlogit| 8.8e-5 / 1.6e-4.
  • Option budget per request. max_len / head_max_len request fields (as in Laya's
    predict_batch) and server defaults --max-len / --head-max-len. Many-option questions
    (e.g. 77 intents) otherwise see one subword per option.
  • Playground redesign. Two-pane layout, live server status, model details from /v1/models,
    API key dialog, light/dark theme, examples (including a German support ticket with a separate
    revenue-impact question), question templates, JSON linting, per-type result views with
    confidence meters, JSON and cURL tabs, session history, keyboard shortcut.
  • Fine-tuning toolkit (tools/finetune/):
    • train_banking77.py: Laya's RLCD recipe on one task, frozen token embeddings, typed-decisions
      replay and --distill learning-without-forgetting replay.
    • train_multitask.py: Banking77 + MASSIVE (51 languages) + multilingual sentiment (12
      languages); train rows that also occur in a test split are dropped (14,121 MASSIVE and 452
      sentiment rows, per language in bench/results/multitask_train_test_overlap.json); per-epoch
      task budgets, LR warmup and EMA weights.
    • merge.py: model soup, task arithmetic and TIES merging of checkpoints from one base.
    • build_mixture.py: licence-clean training mixture from tasksource/tasksource-jev-typed-decisions
      (commercial rows only, per-source cap, evaluation and emotion sources excluded, exact-match
      dedup against every reported test split).
    • eval_laya.py (test suites on GPU) and eval_dev.py (model selection on validation data only).
  • Multilingual benchmark bench/eval_multilingual.py: MASSIVE intents (59 labels) and
    multilingual sentiment, 12 languages each, seeded stratified samples with identical MASSIVE rows
    across languages; results for the base checkpoints in bench/results/.
  • CHANGELOG.md and docs/ROADMAP.md.
  • Brand identity (assets/brand/): logo mark (one pass meeting a column of options, one of
    them chosen), custom monoline wordmark whose i-dot repeats the decision point, light and dark
    lockups, favicon, PNG icons (16–512 px) and a 1280×640 social preview. Used in the README header
    and the playground.

Changed

  • /health and the startup log report the actual compute device instead of a fixed "cpu".
  • The engine raises max_len with a raised head_max_len so the state keeps at least 128 tokens.
    This never binds at the checkpoint defaults; CPU and GPU parity are unchanged.
  • The model parity test prints per-state deviations (STATIM_VERBOSE) and can run one graph per
    item (STATIM_BATCH1).

Fixed

  • Mixture training: items sharing a document with the held-out mixture slice are dropped, so the
    slice no longer inflates checkpoint selection (found by a pre-training audit).
  • eval_dev.py refuses models trained with fewer than 500 held-out Banking77 rows.
  • --gpu-fast help text understated the logit drift (up to ~0.12, not ~1e-2).
  • Playground: the mode selector silently overwrote the chosen model. consensus is now a model
    option, offered only when both checkpoints are loaded.

Results (weights are reproducible with the scripts; not shipped)

  • Banking77 fine-tune v3 (--distill 6000 --epochs 5), first 2,000 test rows: 0.4885 → 0.8655,
    ECE 0.37 → 0.04; held-out Emotion 0.532 → 0.528 and AG News 0.938 → 0.9385 (within noise).
  • Multi-task checkpoint (experimental): MASSIVE macro over 12 languages 0.340 → 0.689, multilingual
    sentiment 0.559 → 0.654, but Banking77 0.8275 (−3.8 vs. v3). Not yet recommended.

Links

Verification: CI green (build + tokenizer/model/engine parity gates for both checkpoints); locally CPU ctest 5/5 and Vulkan ctest 9/9.