Releases: SecFathy/xss-specialist
Release list
XSS Decision 0.8B — Kev XSS Specialist v2 (Experimental)
license: apache-2.0
base_model: Qwen/Qwen3.5-0.8B-Base
library_name: peft
pipeline_tag: text-classification
tags:
- xss
- security
- decision-model
- research-preview
XSS Decision 0.8B — Kev XSS Specialist v2 (Experimental)
An experimental three-way XSS triage checkpoint built by fine-tuning the Kev-0.8B
decision model. It returns probability distributions for vulnerable, safe, or
unknown, plus context and defense classifications. It does not generate exploit
payloads or execute code. Model predictions are not browser-execution evidence and
must not be used as the sole basis for declaring code safe.
Status
Research preview only. Not production-ready. It passed this repository's frozen
synthetic development gate and was scored once on a separate synthetic locked split.
It has not been evaluated on independent real repositories, live applications, or
PortSwigger labs. The very high benchmark scores should not be interpreted as expected
real-world accuracy: the dataset consists of programmatically generated JavaScript/HTML
fixtures with repeated sink primitives and limited structural variation.
Lineage and files
- Base:
Qwen/Qwen3.5-0.8B-Base, pinned at
dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68. - Warm-start parent:
jaredpalmer/kev-0.8b, pinned at
9a45d25eb2ab761841196625383fa1dff0e56c1e. - LoRA rank 16, all supported Kev projection targets, pointer head dimension 256.
- The release contains the adapter, pointer head, tokenizer/configuration, calibration,
training metrics, sanitized provenance, and this card. It does not include Qwen base
weights; those are fetched separately by the serving code. - Adapter SHA-256:
84eea1185e6e4b776839f525f0fee763cc6ee5149ca66c9ccb2170725abb5cfa. - Pointer-head SHA-256:
b3158258853284ede8bb37468917086de6c1311b9df853c15231b386c8fcea68. - Training-suite manifest SHA-256:
bffa4bca7f5ff8f6a1425f38aed88a7682f064132251a89f080ccf2a8574b45d.
Training
One epoch over 1,512 browser-checked training fixtures, 189 optimizer steps, seed 0,
learning rate 2e-5, batch 8, accumulation 1, fp32 on Apple MPS; elapsed time was
3,111.56 seconds. All examples were generated locally by deterministic code; no public
repository snippets, real target responses, or model-generated labels were used. The
full suite had 2,160 JavaScript/HTML fixtures split into train/calibration/development/
locked-test partitions by template group. The locked split was not used for training,
calibration, or candidate selection.
Temperature 0.9267400610196151 was fitted only on the 216-case calibration partition
and is bound to this exact checkpoint's hashes.
Results
Each split has 216 balanced cases (72 per verdict). Recall includes every gold vulnerable
case; a safe/unknown prediction counts as a miss.
| Split | Verdict accuracy | Vulnerable recall | False-safe rate | False-alarm rate | Unknown recall | Unknown overclaim | Context accuracy | Defense accuracy |
|---|---|---|---|---|---|---|---|---|
| Development | 97.22% | 100% | 0% | 8.33% | 100% | 0% | 97.69% | 95.83% |
| Locked test (one read, after selection) | 99.07% | 100% | 0% | 2.78% | 100% | 0% | 99.54% | 98.15% |
The development-only gate marked the candidate RESEARCH_ELIGIBLE. Paired template-group
bootstrap versus the unchanged Kev parent showed a verdict-accuracy delta of +72.69
percentage points (95% CI +67.13 to +78.70) on this synthetic development set. On the
locked test, two safe fixtures were predicted vulnerable; no vulnerable fixture was
missed. This is not a real-world comparative claim.
Intended use and limitations
Intended only for defensive research and code-triage experiments by users who can review
the supplied code and verify findings independently. A safe answer does not prove the
application safe. A vulnerable answer is not a confirmed exploit. Use a browser or other
execution oracle and human review before reporting a finding. unknown is appropriate
when source, sanitizer implementation, framework behavior, or sink context is missing.
Coverage is limited to generated JavaScript/HTML examples involving selected DOM/JS sinks,
four synthetic source types, and a small set of wrappers. It does not establish behavior
for PHP, server-side template engines, full frameworks, URL/CSS contexts, CSP, Trusted Types,
stored workflows in real applications, minified/obfuscated programs, or live systems. It
may over- or under-classify code outside this narrow distribution.
Out of scope: automated exploitation, unsupervised vulnerability claims, production release
gating, or replacing static analysis, browser verification, and expert review.
Local usage
From a compatible checkout of this repository (Python 3.12+):
uv sync --extra decision
uv run --extra decision xss-decision-serve \
--run models/xss-decision-0.8b-kev-specialist-v2 \
--calibration models/xss-decision-0.8b-kev-specialist-v2/temperature.json \
--backend mlx --port 8009Use the versioned questions XSS_QUESTIONS_V2 from xss_decision/questions.py. Keep
relevant imports, helper definitions, and data flow in the request. The first run downloads
the separately licensed Qwen base model. See docs/DECISION_MODEL_V2.md in the source
repository for the evaluation protocol and API client example.
License and attribution
The adapter and pointer head are released under Apache-2.0. The Qwen base is separately
licensed under Apache-2.0. The Kev engine and parent checkpoint are Apache-2.0; this release
is an independent fine-tune and is not endorsed by Qwen or the Kev authors. See the included
LICENSE and upstream model repositories for their notices. The synthetic suite was authored
for this project; it contains no copied application code.
XSS Decision 0.8B — Experimental Pipeline Checkpoint
Experimental pipeline checkpoint
This is the first serialized Kev-style XSS decision checkpoint. It validates the training, LoRA,
pointer-head, MLX loading, and typed-probability inference pipeline.
This model is rejected for security use. It is not production-ready and must not be used to
decide that code is safe. Near-miss vulnerability accuracy is only 2.7%.
Contents
- Qwen3.5-0.8B rank-16 LoRA adapter
- 256-dimensional pointer head
- Tokenizer and adapter configuration
- Training configuration and metrics
- XSS-specific model card
The Qwen/Qwen3.5-0.8B-Base weights are not included and must be downloaded separately.
Measured results
| Split | All decisions | Vulnerability |
|---|---|---|
| Development | 41.2% | 54.4% |
| Generalization | 40.5% | 50.0% |
| Near-miss | 21.1% | 2.7% |
| Adversarial | 35.0% | 50.0% |
The locked test split was not read. See the tracked release manifest for artifact hashes and full
provenance.