Releases: TannerMidd/LANCET-model
Release list
LANCET Nano v0.4.1
LANCET Nano v0.4.1: the v0.4.0 model with a lower review threshold. It asks about more borderline Bash commands and catches more of the risky ones. 110M parameters, 111 MB INT8 ONNX, about 14 ms per command on a CPU, offline.
Model: Apache-2.0. Runtime: MIT.
What's new
- Same weights, same speed, same files. Only the
reviewthreshold inmodel/model.jsonchanges, from0.819to0.303. Therisky(block) threshold is unchanged, so v0.4.1 never blocks more than v0.4.0. - Triage Score 69.7 (v0.4.0: 69.2), a small lead. Every risky command asked about or blocked earns a point; the score shrinks when more than 10% of safe commands are stopped. Three benchmarks (5,500 commands), weighted by size. The neutral set was expanded to 513 outside-party commands after release; v0.4.1 stops more of its safe commands (16.9% vs 11.0%).
Results at the shipped thresholds
| v0.4.1 | v0.4.0 | |
|---|---|---|
| Triage Score | 69.7 | 69.2 |
| lancet-bench-1: risky caught / safe stopped | 91.7% / 9.4% | 89.0% / 6.2% |
| ShellRisk-Bench test: risky caught / safe stopped | 70.5% / 3.1% | 60.6% / 2.3% |
| Neutral set (513 commands): risky caught / safe stopped | 66.5% / 16.9% | 55.7% / 11.0% |
| lancet-bench-1: risky secrets caught (112) | 77% | 72% |
Expect more review results: about 1 in 11 safe commands on the release benchmark (v0.4.0: 1 in 16).
Verify and run
- Check
lancet-v0.4.1-nano-cpu-int8.zipagainstSHA256SUMS.txt. - Extract it and run
python verify_bundle.py --strict. - Follow the bundle's README.
- ONNX SHA-256 (unchanged from v0.4.0):
f412c91867f769aa2b7b0bd5625b460efeb2018fcc5bddd4b39f09dfd2dc4f32 - ZIP SHA-256:
894c87814d52e919e8342c4ed1b5e07d999b16354dd7c06c222dbe20c81a7ab0
Inputs are inert text and are never executed. not_flagged is not execution authorization or a safety guarantee.
Credits and source licenses: NOTICE.txt and THIRD-PARTY-NOTICES.md in the bundle.
LANCET Nano v0.4.0 (CodeT5+ 220M)
LANCET Nano v0.4.0: a local Bash command-risk classifier. 110M parameters, 111 MB INT8 ONNX, about 14 ms per command on a CPU, offline.
Model: Apache-2.0. Runtime: MIT.
What's new
- New base: the Salesforce CodeT5+ 220M encoder, trained from its pinned upstream weights. Same size and speed as v0.3.0.
- New training data: attack-technique commands from the ShellRisk-Bench training split (Atomic Red Team, InternalAllTheThings), more SWE-smith and Terminal-Bench everyday commands, a project red-team corpus with matching safe look-alikes, and more project-authored evaluation cases.
Results
| Risky caught at ≤10% of safe commands stopped | v0.4.0 | v0.3.0 |
|---|---|---|
| lancet-bench-1 (793 commands) | 92.2% | 93.4% |
| ShellRisk-Bench test split (4,194 commands) | 91.2% | 59.1% |
| Neutral third-party set (66 commands) | 52.4% | 40.5% |
| At the shipped thresholds | v0.4.0 | v0.3.0 |
|---|---|---|
| lancet-bench-1: risky caught / safe stopped | 89.0% / 6.2% | 85.8% / 5.5% |
| lancet-bench-1: risky secrets caught (112) | 72% | 66% |
| ShellRisk-Bench test: risky caught / safe stopped | 60.6% / 2.3% | 46.6% / 6.1% |
| Neutral set: risky caught / safe stopped | 52.4% / 12.5% | 50.0% / 20.8% |
| AUROC (lancet-bench-1 / ShellRisk / neutral) | 0.974 / 0.952 / 0.746 | 0.962 / 0.827 / 0.701 |
Network/remote-execution commands remain its weakest area on lancet-bench-1 (59% caught, 22 cases).
Verify and run
- Check
lancet-v0.4.0-nano-cpu-int8.zipagainstSHA256SUMS.txt. - Extract it and run
python verify_bundle.py --strict. - Follow the bundle's README.
- ONNX SHA-256:
f412c91867f769aa2b7b0bd5625b460efeb2018fcc5bddd4b39f09dfd2dc4f32 - ZIP SHA-256:
b0dde0b755908025221f485163a747bef332174dbfe2f215439f02c5d3123053
Inputs are inert text and are never executed. not_flagged is not execution authorization or a safety guarantee.
Credits and source licenses: NOTICE.txt and THIRD-PARTY-NOTICES.md in the bundle.
LANCET Nano v0.3.0
LANCET Nano v0.3.0 is a larger, more accurate local Bash command-risk classifier. It uses a CodeT5-base encoder, trained from the pinned upstream weights rather than an earlier LANCET checkpoint. It remains experimental.
Model: Apache-2.0. Runtime: MIT.
What's new
- Larger model: 109.6M parameters, a 111 MB INT8 ONNX file, and about 22 ms per command on a CPU (median, measured on a busy Ryzen 9 3900X). v0.2.0 is 35.3M parameters and 36 MB.
- More training data:
- Azure CLI, GitHub CLI, kubectl and Docker CLI reference examples.
- A documentation rule that labels commands which print or mint secrets as risky, using AWS's own
sensitiveoutput-field markings. - A project-authored secrets family.
- LANCET's earlier agent-authored evaluation suites, retired into training with their original labels.
- Labels: no language model, hosted API or human labeler produced any training label.
Results on the 793-command release benchmark (one pass each)
| v0.3.0 | v0.2.0 | v0.1.0 | Jev (hosted) | |
|---|---|---|---|---|
| Risky caught | 85.8% | 73.6% | 65.5% | 96.8% |
| Safe wrongly stopped | 5.5% | 6.2% | 5.5% | 7.8% |
| Risky secrets caught | 66% | 24% | 14% | 98% |
- Benchmark caveat: the benchmark's labels were written by the developer, an AI agent. This is diagnostic evidence, not independent acceptance.
- Upstream ShellRisk sets (not agent-authored, never trained on):
- source-external: v0.3.0 caught 48.9% vs 48.3% for v0.2.0 and stopped 6.3% vs 9.7% of safe commands.
- source-holdout: v0.3.0 caught 41.9% vs 45.9%.
- Weaker areas: network/remote execution and infrastructure-as-code.
Release by owner exception
- Failed check: one of four preregistered release checks. It required the 95% lower bound of the ShellRisk catch-rate difference to be at least −5 points. The observed difference was +0.6 points [−6.5, +7.8].
- Why: with only 176 risky commands, even an equal model would usually fail that check. This design error is disclosed.
- Decision: the owner approved the release as a documented exception, recorded in
provenance/release-gate.json.
Verify and run
certutil -hashfile lancet-v0.3.0-nano-cpu-int8.zip SHA256 # compare with SHA256SUMS.txt
python verify_bundle.py --strict
python -m pip install -r requirements.txt
echo {"command": "kubectl delete namespace prod", "shell": "bash"} | python classify.py --model model
- Checksum: ZIP SHA-256
d7ab46ab6477a7803cfed230c3d6fcc27d7bb96864d7ceb83e1bffb28132d3e5. - ONNX SHA-256:
b635571b0c55fba5224a00e2c9132918db816b76c433e8c8dc75d151ef03de1c. - Parity: the release runtime matched the research model on all 7,527 development rows (identical bands, zero score difference).
- Also available: a browser demo at https://huggingface.co/spaces/fingerthief/lancet-nano and the model repo at https://huggingface.co/fingerthief/lancet-nano (tag
v0.3.0).
not_flagged is not execution authorization or a safety guarantee. v0.2.0 (smaller and faster) remains available unchanged.
LANCET Nano v0.2.0
LANCET Nano: a local classifier for Bash command risk. It has 35.3M parameters, ships as a 38 MB CPU INT8 ONNX model, takes about 3 ms per command and needs no network connection.
Model: Apache-2.0. Runtime: MIT. The datasets used for training are not included.
Results on 400 fresh commands (sealed diagnostic suite, one scoring pass)
| LANCET Nano | V5 (v0.1.0) | |
|---|---|---|
| Risky caught | 83.9% (177/211) | 76.3% |
| Safe commands wrongly stopped | 5.8% (11/189) | 8.5% |
| AUROC | 0.935 | 0.902 |
For comparison on the same suite:
| Model | Risky caught | Safe commands wrongly stopped |
|---|---|---|
| Hosted Jev (about 200× larger by estimate) | 97.6% | 4.8% |
| Laya (421M parameters, GPU) | 77.7% | 42.9% |
The suite's labels were written by the developer, an AI agent. These results are diagnostic, not independent acceptance.
What changed from V5
- The new training labels come from documentation. The wording of the human-written example descriptions in tldr-pages (CC BY 4.0) sets each example's label, and so do the AWS operation verbs in the botocore service models (Apache-2.0).
- No language model, hosted API or human labeler produced any label.
- Nano now returns three bands:
risky,reviewandnot_flagged.
Install and verify
- Download
lancet-v0.2.0-nano-cpu-int8.zip. - Check it against
SHA256SUMS.txt. - Extract it and run
python verify_bundle.py --strict. - Follow
README.mdinside the ZIP.
Tested
- The ZIP builds reproducibly.
- Strict verification passes on all 30 files.
- In a clean virtual environment with the pinned dependencies (numpy 2.5.3, tokenizers 0.23.2, onnxruntime 1.30.0), the release runtime exactly matched the research runtime on 5,262 development rows: identical scores, identical bands and identical input-contract handling.
Limits
- Bash only.
not_flaggedis not execution authorization or a safety guarantee.- Independent acceptance and human label adjudication remain unavailable.
Model ONNX SHA-256: 183ae5051fffe8578267c524efb78ef82c2f668f1e575690923ced8c5f1827f5
LANCET v0.1.0 - V5 CPU INT8
LANCET v0.1.0 — V5 CPU INT8
Experimental prerelease. Model: Apache-2.0. Runtime: MIT. The included licenses permit use, modification and redistribution, including commercial use, subject to their terms. Training data is not included or relicensed.
This packages the existing V5 codet5-balanced / recall95 classifier—not V6 or an in-progress research candidate. Model weights, tokenizer, runtime, calibration and threshold are unchanged. This is a licensing/documentation/packaging revision, not a model promotion.
Download and verify
lancet-v0.1.0-v5-cpu-int8.zip— 37,939,527 bytes; real ONNX model, tokenizer, CPU runtime, standalone model card, full licenses/notices and hash verifier.SHA256SUMS.txt— ZIP checksum.
Check the ZIP checksum, extract it, then run python verify_bundle.py --strict before creating an environment inside the extracted directory. The verifier checks all 29 payload files against the manifest without model inference. The ZIP has 30 entries, including that manifest.
Use Python 3.12, install requirements.txt in a virtual environment, set OPENBLAS_NUM_THREADS=1, and run python classify.py --model model. The process classifies JSONL Bash text on stdin; it never executes inputs. No API key or network access is needed for inference after installation. PowerShell and unsupported/invalid/oversized inputs require review. not_flagged is not execution authorization.
Credits and license scope
- Salesforce CodeT5-small, pinned revision
b1ee9570c289f21b5922b9c768a1ce12957bf968. Credit to Yue Wang, Weishi Wang, Shafiq Joty and Steven C. H. Hoi for CodeT5 (2021). - The original model card declares Apache-2.0. Its text, including Hugging Face/nielsr credit, and the full license are preserved. The separate CodeT5 source-project BSD-3-Clause notice—copyright (c) 2021 Salesforce.com, Inc.—is also retained, not substituted for the model license.
- LANCET-specific contributions and runtime: Tanner Middleton, copyright (c) 2026.
MODEL-LICENSE.mdexplicitly grants Apache-2.0 for the V5 model assets; runtime code remains MIT. - Downstream data credits: Kontext Security / ShellRisk-Bench, Kwai-Klear / SWE-smith trajectories, yoonholee / Terminal-Bench trajectories, and TellinaTool / NL2Bash. Actual V4/V5 fitting provenance was audited: the external fitting pools use only these MIT/Apache-declared sources, not the collection's GTFOBins or undeclared-license source rows. NL2Bash's
data/bashhas its own MIT license despite the repository's GPL code badge; the actual dataset notice is included.
See LICENSE.md, MODEL-LICENSE.md, NOTICE.txt, MODIFICATIONS.md, THIRD-PARTY-NOTICES.md and licenses/ in the archive. No upstream endorsement is claimed. Dataset declarations do not independently certify rights in every underlying third-party snippet/trajectory; the technical audit is not legal certification.
Provenance and limits
This public distribution starts from a new, standalone history, commit cdc44d7c0c3f39a4583fef8a62c7b76e169fdb8e, and does not mirror the private research repository. The frozen model/runtime originated at commit c9a571fa40452ad26d3b8bf75399af91c41416c0. The original archive is preserved byte-for-byte: its manifest records the licensing/documentation overlay as uncommitted at package-build time; that overlay is now included in this public distribution. Historical pre-release wording is distinguished from the current model grant.
No raw datasets, private logs, evaluation corpora, fitting checkpoints, research models or vendored dependency distributions are attached. No new fitting or benchmark scoring was performed for release preparation. V5 remains unchanged.
Historical classifier-only eval150 recorded 138/150 required interventions caught and 14/128 controls interrupted. These adaptive research results are not independent test accuracy. Independent acceptance and human label adjudication remain unavailable and unmet.
Development and programmatic evidence; not independently validated. Safety guidance is not an additional license restriction.
Model SHA-256: e8f2e75580a435eba9744b4a8f14dc851f7c71e579c344bde75dee047f7c07a3
ZIP SHA-256: fa2784231bdf54ad0aaed35d097724ee2e26c99dd264b14c2ba1db6e8656a0f1