Skip to content

Releases: bladedevoff/stuntd

stuntd 0.2.0

Choose a tag to compare

@bladedevoff bladedevoff released this 10 Oct 07:00

What's new

OpenAI Decisions API. stuntd now answers POST /v1/decisions, the typed-question API OpenAI opened in beta on 2026-10-06 (gpt-6-luna). Each predicate, choice and score question becomes a decision site exactly like a Jev question, so recording, training, report, the novelty gate and the modes all work as before:

  • In front of OpenAI, a request is answered locally only when every question has a live head that is sure; otherwise it goes to OpenAI untouched and each answer becomes a label. Refusals are relayed and never recorded.
  • With no upstream, heads answer what they know and zero-shot Laya answers the rest, in the Decisions shape the openai SDK parses (boolean choice values stay booleans, scores are the probability-weighted level).
  • Not learned, relayed whole: images, colliding choice values, more than 10 score levels and the other cases listed in the README.

Not measured against the live API: the shapes were checked against the Decisions guide and the openai 3.27.0 SDK types, and the release check ran the real SDK against stuntd in both modes.

Stock sentence encoders. training.encoder can name intfloat/multilingual-e5-base, intfloat/multilingual-e5-small, sentence-transformers/all-MiniLM-L6-v2 or another sentence encoder, and stuntd train fits a linear head on its frozen pooled vectors (one vector per row, a few seconds per site). Laya stays the default and the only zero-shot model.

  • Encoders load with trust_remote_code=False and only from an allowlist of architectures.
  • Every head records the encoder it was trained on. A mismatched head is never served; it shows as encoder-mismatch until you retrain.

On MASSIVE intent (60 intents, 3,000 training rows, the full official test split), accuracy of every answer:

Laya multilingual-e5-base multilingual-e5-small all-MiniLM-L6-v2
ru 26.2% to 27.2% 81.6% 78.6% 54.7%
de 29.5% to 31.5% 79.6% 75.9% 65.1%
es 34.3% to 34.4% 80.5% 78.0% 63.5%
zh 42.7% to 44.1% 79.8% 78.3% 40.2%
ja 32.1% to 32.2% 79.8% 79.5% 57.7%
en 60.1% to 60.7% 83.7% 81.2% 81.0%
Banking77 65.1% to 66.1% 89.6% 87.5% 90.0%

multilingual-e5-base answers 18.0% to 41.9% of the MASSIVE test rows locally at 97.9% to 99.3% accuracy, and decides in about 14 to 27 ms on a laptop CPU against 340 to 400 ms for Laya. Laya's numbers vary between two runs, so its cells give both. benchmarks/encoders.py reproduces every number, and benchmarks/README.md has the full tables, hardware and dataset revisions.

Jev-compatible teachers. The README now documents Kev, Mica v0.1 4B, Perplexity Decider v1.1 and Strom as jev.upstream: how to run or reach each one, and what differs from Jev. tests/test_teachers.py contract-tests each one's documented request and response through stuntd. Nothing was run against the real models.

The new jev.path setting (default /v1/systemone) covers Perplexity's hosted /v1/decisions path.

Fixes. A body nested too deeply or carrying a lone UTF-16 surrogate no longer gives a 500 on /v1/systemone; it is relayed unlearned in proxy mode and gets a 400 locally.

Behaviour changes to know

  • /v1/decisions used to be relayed to the upstream like any other path; it is now learned. With upstream empty it is answered locally, so a daemon with no top-level upstream now loads the Laya checkpoint at start even when jev.upstream is set.
  • /healthz still counts a site with an encoder mismatch as live; stuntd status shows the real state.

stuntd 0.1.4

Choose a tag to compare

@bladedevoff bladedevoff released this 07 Oct 16:22

What's new

The novelty gate cuts at the 0.99 quantile by default. At 0.95 every head sent about 5% of perfectly ordinary traffic to the provider on purpose, and a decision with several fields paid that once per field. Measured with the new examples/support/check.py:

support demo, 1,000 tickets no gate 0.95 (0.1.3) 0.99 (0.1.4)
familiar tickets answered locally, heads on all 20 templates 72.3% 67.1% 70.6%
familiar tickets answered locally, heads on 15 templates 88.8% 78.2% 85.8%
tickets from the 5 left-out templates answered locally 68.7% 0.0% 0.3%

All three answers stay right on about 97% of the familiar tickets the heads answer, and on the left-out templates, where the heads without a gate were right on only 42.4% of what they answered, the gate still sends almost everything to the provider. "What's the weather in Paris", asdf qwer zxcv and an empty state are stopped at every setting. Heads trained with 0.1.3 keep the cut-off stored with them; retrain to get the new one, or set training.novelty_quantile yourself.

examples/support/check.py. One command trains the demo's heads twice (all templates, and with one template per category left out) and prints the numbers above, the quantile sweep, how often an answer changes when only fields no rule reads change, and whether the junk probes are stopped. Every generalization number in the README comes from it, so you can rerun it instead of trusting it.

auto_retrain works on the Jev path. Jev captures now count toward it exactly like OpenAI and Anthropic ones; before, a Jev-only setup never retrained on its own.

Fixes. stuntd site rm removes a site whose mode file is unreadable, and report --gold in local mode counts a failed zero-shot answer as unanswered instead of stopping.

Behaviour changes to know

  • New heads gate less: on the demo, 1.7 to 3.0 points of familiar traffic come back to the head compared with 0.1.3. Set training.novelty_quantile = 0.95 to keep the old behaviour.

Thanks

To Dipankar Sarkar for asking what 0.99 would do.

stuntd 0.1.3

Choose a tag to compare

@bladedevoff bladedevoff released this 03 Oct 06:53

What's new

Novelty gate. A head trained with 0.1.3 keeps the encoder vectors of its training rows. A request unlike anything it trained on goes to the provider with reason=novel, however confident the head looks. On the support demo with one template per category left out of training, the gate sent 100% of the tickets from the unseen templates to the provider; without it, all three heads were sure on 40% of them and all three right on only 61% of those. "What's the weather in Paris", asdf qwer zxcv and an empty state were stopped too; without the gate the heads answered the weather question at confidence 1.00. The cost: about 6 points fewer requests answered locally on familiar tickets (76.6% to 70.6%), with accuracy unchanged. It is on by default; training.novelty_quantile and serving.novelty_gate change that.

An auto loop you can leave running. auto_retrain counts distinct texts, not repeated captures, and waits training.auto_retrain_min_minutes (30) between runs of a site. In a decision with several fields, a field that is live keeps being compared with the provider while the others are still in shadow, so it can still be demoted. A retrained head that replaced a live one shows as shadow (retrained, was live) instead of switching back silently. stuntd enable PARENT enables every field of a decision.

Honest local mode. Without a Jev provider, a request under the threshold is answered zero-shot by the base checkpoint. serving.local_fallback = "head" answers it with the head instead, report --gold now scores what is actually served under that setting, and enable warns when a head would answer less than half of the traffic.

Honest report. Holdout agreement and the operating point come with a 95% interval and the rows behind them, with a warning when there are too few. "Surest mistakes" show the probability the head gave and whether each one would have been served.

Sites. stuntd import skips rows already recorded (--dry-run to see the counts). stuntd site rm and stuntd site rename exist; a renamed site keeps receiving the requests of its old name. status shows the schema name next to a hash site and groups fields under their decision.

Operations. GET /healthz, stuntd stop, one progress line per epoch during training, a startup line that says what is served, reason=unsupported-schema for schemas stuntd cannot learn, and pip install "stuntd[train,jev]" for the local Jev quickstart.

Behaviour changes to know

  • The novelty gate is on for heads trained with 0.1.3; heads from earlier versions serve exactly as before.
  • auto_retrain retrains less often on traffic with many repeated requests.
  • Importing the same file twice no longer records its rows twice.

Thanks

To Dipankar Sarkar, whose review of the published support heads on Hugging Face showed that the demo's test leaked templates and that a head's confidence says nothing about inputs it never saw, and to the people who tried 0.1.2 and said where it got in the way.

stuntd 0.1.2

Choose a tag to compare

@bladedevoff bladedevoff released this 30 Sep 15:13

What's new

Decisions with several fields. A structured output with 2 to 8 typed fields (enum, boolean or number), like {category, urgency, needs_human}, is now learnt. Each field gets its own head, and the request is answered locally only when every field is sure. Otherwise the provider answers and every field is recorded. Single-field sites keep their names, so heads from 0.1.0 and 0.1.1 keep serving.

Anthropic Messages API. stuntd serve --upstream https://api.anthropic.com learns decisions sent to /v1/messages, through output_config.format or a single tool. The local answer is checked against the Message model of the official anthropic SDK in the tests. It has not been run against a live API key yet, so reports are welcome.

Retraining in the background. training.auto_retrain = N retrains a site once N new captures arrive, one run at a time, as a child process with its own log. With serving.auto_promote on, the loop runs without you: collect, train, shadow, promote, demote, retrain.

Lazy checkpoint load. stuntd serve --lazy (or serving.lazy_load = true) starts at once and loads the base checkpoint on the first request that needs it. A failed load answers that request with the usual model error and the next one tries again. Closes #1.

Upgrading

pip install -U stuntd. Nothing to change in your config: the new settings are off by default. One thing to know: requests with several typed fields used to pass through as free text and are now recorded, so the store grows for them.

Community

Questions and show and tell now live in Discussions.

stuntd 0.1.1

Choose a tag to compare

@bladedevoff bladedevoff released this 25 Sep 18:36

stuntd 0.1.1

Sites with many or long labels now train with room to read every label, the encoder cache sizes
itself by the machine, and stuntd report --gold scores a head against rows a person checked.

What changed

  • Each site gets its own layout. A site with many or long labels widens the option window
    until every label fits whole, up to training.max_option_tokens (default 1024). The room left
    for the request text stays as it was. Underscores in labels are shown to the model as spaces,
    unless that would make two labels equal. Heads trained by 0.1.0 keep serving exactly as before.
  • The encoder cache is sized by the machine. training.cache_max_mb = 0, the new default,
    means half of physical memory instead of a fixed 4 GiB. A site that still does not fit trains
    uncached and says so on stderr, with the GiB it needs and the setting to raise.
  • stuntd report SITE --gold FILE. Scores the head and its teacher against a JSONL file of
    verified rows, in the stuntd import format: head accuracy, head accuracy on rows the store
    has not seen, what stuntd would serve at the threshold, teacher accuracy, and where the two
    disagree. The unseen-rows line is the out-of-sample number. Do not import the gold rows, or
    send them through the proxy, before scoring them.
  • report --curve is short. It prints the threshold at every 5% of coverage plus the
    operating point, about 20 rows, instead of one row per holdout example. --json and the saved
    metadata still carry the whole curve.
  • Docs. The README covers the new settings, report --gold, starting offline, and keeping
    one decision to a question.

Behaviour changes to know

  • training.cache_max_mb = 0 now means half of physical memory. In 0.1.0 it left every site
    uncached. To turn caching off, set training.cache_encoder = false.
  • Half of memory is read from the host, not from a container's cgroup limit. In a
    memory-limited container, set training.cache_max_mb explicitly.
  • report --gold: rows the store has already captured may be training rows. The line for rows
    stuntd has not seen is the out-of-sample number. Never import the gold file before scoring it.
  • On Windows, a training run that exhausts the commit limit can crash with an access violation
    instead of a MemoryError. If that happens, lower training.cache_max_mb or free memory.
  • For code that calls LayaTrainer directly: a call now returns TrainedHead (the holdout
    logits plus the layout the head was trained with) instead of a bare list of logits; use
    .logits for the old value and store .layout with the head. question_for still imports
    from stuntd.train.trainer and keeps its 0.1.0 call.

Benchmark

banking77 (77 intents), a head trained on the frozen Laya encoder from 8,000 of the 10,000
training rows, checked on the full official test split (3,080 rows):

0.1.0, 24 epochs 0.1.1, 24 epochs 0.1.1, 48 epochs
accuracy 67.5% 75.2% 76.0%

The 95% interval on 3,080 rows is about ±1.6 points.

On the 500-row sample dhruvmehra/jevbench uses, 0.1.1 at 48 epochs scores 76.0% and Jev 76.4%, a
tie within that sample's noise. The number of epochs was picked on the holdout, not on the test.
The four demos in the README reproduce their numbers exactly on 0.1.1.

stuntd 0.1.0

Choose a tag to compare

@bladedevoff bladedevoff released this 23 Sep 00:57

First release.

stuntd is a local proxy for typed LLM decisions. It records the decisions your app already asks a provider for (OpenAI Chat Completions with a typed response_format, or Jev-protocol questions), trains a Laya head per decision site, serves it in shadow mode, goes live only at your target agreement, falls back to the provider below the confidence threshold, and demotes itself on drift.

  • Jev protocol: POST /v1/systemone answered locally by the base Laya checkpoint without a key, or proxied to any Jev-compatible server (the paid API, kev, SemIf, LLM2Jev, laya-server) while learning from its answers.
  • OpenAI-compatible proxy for typed decisions, byte-exact passthrough for everything else.
  • stuntd import for labelled JSONL, learn = false for a pure proxy.
  • Training with the frozen encoder's output cached once per site (3.2x faster than re-encoding every epoch).
  • Four demos with measured numbers: Snake, support triage, banking intents and risk, a command gate for coding agents with a Claude Code PreToolUse hook.

Install: pip install "stuntd[train]". Python 3.10+, Apache-2.0.