Repository navigation
LANCET Nano v0.4.3
Pre-releaseLANCET Nano v0.4.3: a newly trained model for Bash, PowerShell and cmd commands, with the highest Triage Score of the 14 guards tested on the new, harder lancet-bench-2-next. 111M parameters, 115 MB INT8 ONNX encoder and head, about 14 ms per command on a CPU, offline.
Model: Apache-2.0. Runtime: MIT.
What's new
- Triage Score 68.3 (v0.4.2: 54.4). Every risky command asked about or blocked earns a point; the score shrinks when more than 10% of safe commands are stopped. Three benchmarks (7,911 commands), weighted by size.
- PowerShell and cmd. v0.4.2 answered
reviewfor every non-Bash command; v0.4.3 classifies all three shells. - Long commands read in full. Commands longer than one 512-token window are split into overlapping windows; every token is used and nothing is truncated.
- Far fewer interruptions on look-alike safe commands: on lancet-bench-2-next it catches the same 78.0% of risky commands as v0.4.2 while stopping 12.4% of safe ones instead of 20.0%.
- Newly trained from the pinned CodeT5+ 220M base on a new command-semantics curriculum of risky/safe twins (21,544 rows), trained to convergence. Same size and CPU speed as v0.4.2.
- New runtime (
classify.py) for the windowed format; it checks every model file's hash before loading.
A harder benchmark: why every score dropped
The Triage Score now uses lancet-bench-2-next instead of the retired lancet-bench-1 (793 Bash commands), where the best guards were near the ceiling. The new benchmark has 3,204 commands built as 1,602 risky/safe twins, from 289 scenarios in 40 areas, each repeated inside six shell wrappers, in Bash, PowerShell and cmd. No LANCET release was trained, tuned or calibrated on it. Every guard scores lower on it: v0.4.2 scored 82.7 under the previous Triage Score and 54.4 under the new one, with the same model.
Results at the shipped thresholds
| v0.4.3 | v0.4.2 | |
|---|---|---|
| Triage Score | 68.3 | 54.4 |
| lancet-bench-2-next: risky caught / safe stopped | 78.0% / 12.4% | 78.0% / 20.0% |
| ShellRisk-Bench test: risky caught / safe stopped | 64.8% / 3.0% | 60.1% / 2.7% |
| Neutral set (513 commands): risky caught / safe stopped | 85.8% / 5.6% | 91.0% / 7.0% |
| Risky blocked outright (lancet-bench-2-next) | 50.0% | 17.7% |
Triage Score by release (lancet-bench-2-next): v0.1.0 40.5 · v0.2.0 40.7 · v0.3.0 44.7 · v0.4.0 47.1 · v0.4.1 44.9 · v0.4.2 54.4 · v0.4.3 68.3.
Other guards on the same benchmarks: Jev 38.9, Kestrel 30.7, ModernBERT bash 23.2, bash-classify 14.3, Laya 13.3, bev-decider 13.3, sh-guard 10.9. Full table and charts: https://github.com/TannerMidd/LANCET-model#triage-score-across-three-benchmarks.
Verify and run
- Check
lancet-v0.4.3-nano-cpu-int8.zipagainstSHA256SUMS.txt. - Extract it and run
python verify_bundle.py --strict. - Follow the bundle's README.
shellisbash,powershellorcmd.
- ONNX encoder SHA-256:
4e7d6a53d27a7321a2638a4bf446301e8e343c51000963331469e3ab20aae2a4 - ZIP SHA-256:
75b7307abf16a9c63806be6a6aa88e9b18cc463cddf7e4fdc6c9b8e27d4c985c
Inputs are inert text and are never executed. not_flagged is not execution authorization or a safety guarantee.
Credits and source licenses: NOTICE.txt and THIRD-PARTY-NOTICES.md in the bundle.