Skip to content

Results

Anubha Parashar edited this page Aug 11, 2026 · 1 revision

Results

Aggregate

Metric Value
Evidence rows 31,396
Regression-detection recall 75.02%
Recall 95% CI 71.53%–78.41%
Equal-budget random recall 55.01%
Mean test reduction 61.35%
Reduction 95% CI 58.73%–63.90%
Paired Cohen's d 0.504
McNemar exact p 4.14 × 10^-17
Paired permutation p 9.999 × 10^-5
Recorded API cost $0.00

Models

Model Regression recall Task success Attack success Mean latency
Qwen3 4B 90.88% 0.02% 0.00% 99.63 s
Gemma3 4B 63.51% 40.92% 5.97% 21.94 s
Llama 3.2 3B 79.68% 50.38% 5.92% 23.44 s
Phi-4 Mini 66.01% 73.90% 3.62% 22.93 s

The results expose a security-utility trade-off; planner recall, task utility, proposal behavior, containment, and attack success should not be collapsed into one metric.

Clone this wiki locally