Skip to content

Statim 0.3.0

Choose a tag to compare

@BEKO2210 BEKO2210 released this 27 Sep 07:51
· 127 commits to main since this release
74de292

Statim

Statim 0.3.0 · better decisions in 12 languages, trained only on data you may use commercially

Statim 0.3.0 vs. the base checkpoint on 54 held-out suites

Highlights

  • Multilingual intents 0.340 → 0.733 over 12 languages (Arabic 0.20 → 0.61, Hindi 0.27 → 0.67); Banking77 0.517 → 0.891; typed decisions 0.351 → 0.691. Zero-shot suites the model never trained on stay within noise or improve.
  • No-harm gate: a new model only replaces the old one if validation improves and none of 54 held-out suites drops by more than two standard errors. This release: 16 significant gains, 0 regressions.
  • Licence-clean training: every training source was audited at its origin (279 checked, 116 excluded); only permissive licences, listed in DATA_LICENSES.md.
  • Licensing model: code Apache-2.0; published model weights free for noncommercial use under PolyForm Noncommercial 1.0.0, commercial use by licence (COMMERCIAL.md).
  • Enterprise security hardening: 12 findings of an independent review fixed with regression tests (fail-closed auth, JSON and work limits, bounded HTTP queues, log-injection fix, cpp-httplib 0.58.0). Details in docs/SECURITY.md.
  • Documentation: OpenAPI 3.1 spec and API guide; zero-shot benchmark over eight unseen suites.

Results

  • The 0.3.0 model (multilingual checkpoint fine-tuned with train_multitask.py --clean on
    licence-audited data only) against the base checkpoint, evaluated by tools/finetune/gate.py on
    54 held-out suites: MASSIVE intents over 12 languages 0.340 → 0.733 (Arabic 0.200 → 0.613,
    Hindi 0.267 → 0.673), Banking77 0.5175 → 0.891, typed-decisions 0.351 → 0.6905, HWU64 (sibling
    of MASSIVE, overlapping rows removed) 0.500 → 0.760; zero-shot suites never trained on stay
    within noise or improve (GoEmotions +6.0, SIB-200 +2.5, SemRel +2.9, Belebele +2.7, FarsTail
    −2.0, DAIR Emotion −1.5, AG News +0.1 points). 16 significant gains, 0 significant
    regressions; calibration error roughly halves. Chart: assets/diagrams/results-0.3.0-*.svg.
    Weights are reproducible with the scripts and not yet published.

Added

  • Release results chart generated from gate evaluations (tools/diagrams/gate_chart.py).
  • Updated docs/API.md and docs/openapi.yaml for the hardened server (Codex, checked against a
    running server).
  • Licensing model: source code stays Apache-2.0; model weights published by Statim are licensed
    under PolyForm Noncommercial 1.0.0 (LICENSE-MODEL.md, verbatim official text), commercial use
    needs a paid licence (COMMERCIAL.md).
  • DATA_LICENSES.md, generated by tools/finetune/data_licenses.py: base models, training data of
    released weights with licences, evaluation-only data, and data excluded from releases.
  • train_multitask.py --clean: commercial-clean training (no tyqiangz sentiment, distillation texts
    from the licence-filtered mixture).

Changed

  • build_mixture.py keeps only rows whose every listed licence is permissive (Apache-2.0, MIT, BSD,
    CC0, CC-BY, ODC-By, AFL-3.0); ShareAlike, copyleft, custom and unknown terms are excluded.

Security

  • Fail closed when any explicitly configured key file/environment value cannot supply valid
    keys. Log unauthenticated local mode explicitly; require the production environment file/key.
  • Preflight JSON with depth, node, object-width, key-length, and duplicate-key checks before
    constructing the ordered DOM, preventing deeply nested crashes and quadratic wide-object parsing.
  • Bound HTTP workers and pending sockets; authenticate and admit inference requests before
    body reception. Reject oversized declared bodies without draining, and enforce absolute
    header/body deadlines even when a peer keeps sending bytes.
  • Bound field sizes and aggregate batch/ensemble/calibration/consensus work, token budgets,
    attention-memory estimates, and response bytes. Add cooperative inference deadlines and
    systemd memory/CPU/task/file-descriptor ceilings without changing graph packing or math.
  • Replace the entry-count-only calibration cache with a byte-bounded LRU keyed on the validated
    question only, so unknown fields (still ignored, as by laya.serve) are never retained; a test
    sends 64 MiB of ignored metadata and checks that memory stays flat.
  • Allowlist request IDs and serialize log records as JSON to prevent log injection.
  • Validate signed/unsigned token and ensemble budgets before integer narrowing, including
    CLI defaults; reject effective sequence lengths beyond model capacity with 422.
  • Upgrade vendored cpp-httplib from 0.26.0 to 0.58.0. Explicitly reject simultaneous
    Content-Length/Transfer-Encoding (including zero lengths) and duplicate Content-Length headers.
  • Require bearer auth for /metrics and /v1/models when configured. Keep /health and
    /ready open; reduce /health to status and version. /v1/models now also reports each
    model's compute device, and the playground reads models and device from there.
  • Use local timestamp storage with gmtime_r (gmtime_s on Windows) for concurrent logs.
  • Install a global HTTP exception handler with fixed client errors and server-only details.
  • Add C++, CPU model-backed, and live HTTP security regressions alongside existing parity gates.

Links

Verification: CI green (build, parity gates, security suites); locally ctest 12/12 including GPU parity; gate PROMOTE with 0 regressions on 54 held-out suites. Model weights are reproducible with the documented commands and not yet published.