Skip to content

stuntd 0.1.4

Choose a tag to compare

@bladedevoff bladedevoff released this 07 Oct 16:22
· 2 commits to main since this release

What's new

The novelty gate cuts at the 0.99 quantile by default. At 0.95 every head sent about 5% of perfectly ordinary traffic to the provider on purpose, and a decision with several fields paid that once per field. Measured with the new examples/support/check.py:

support demo, 1,000 tickets no gate 0.95 (0.1.3) 0.99 (0.1.4)
familiar tickets answered locally, heads on all 20 templates 72.3% 67.1% 70.6%
familiar tickets answered locally, heads on 15 templates 88.8% 78.2% 85.8%
tickets from the 5 left-out templates answered locally 68.7% 0.0% 0.3%

All three answers stay right on about 97% of the familiar tickets the heads answer, and on the left-out templates, where the heads without a gate were right on only 42.4% of what they answered, the gate still sends almost everything to the provider. "What's the weather in Paris", asdf qwer zxcv and an empty state are stopped at every setting. Heads trained with 0.1.3 keep the cut-off stored with them; retrain to get the new one, or set training.novelty_quantile yourself.

examples/support/check.py. One command trains the demo's heads twice (all templates, and with one template per category left out) and prints the numbers above, the quantile sweep, how often an answer changes when only fields no rule reads change, and whether the junk probes are stopped. Every generalization number in the README comes from it, so you can rerun it instead of trusting it.

auto_retrain works on the Jev path. Jev captures now count toward it exactly like OpenAI and Anthropic ones; before, a Jev-only setup never retrained on its own.

Fixes. stuntd site rm removes a site whose mode file is unreadable, and report --gold in local mode counts a failed zero-shot answer as unanswered instead of stopping.

Behaviour changes to know

  • New heads gate less: on the demo, 1.7 to 3.0 points of familiar traffic come back to the head compared with 0.1.3. Set training.novelty_quantile = 0.95 to keep the old behaviour.

Thanks

To Dipankar Sarkar for asking what 0.99 would do.