LLMFuzz Red Team v0.2.0
LLMFuzz Red Team uses GPT-5.6 Sol through the OpenAI Responses API to generate a bounded, structured adversarial corpus for AI-agent security regression testing.
The accepted corpus contains 16 cases across four demonstrated attack classes:
- prompt injection;
- secret exfiltration;
- forbidden tool use;
- approval bypass.
The corpus is validated and identified with SHA-256. Deterministic machine-event invariants—not an LLM judge—assign the final verdicts.
Verified control-fixture results
Using the same accepted corpus:
- Vulnerable fixture: 16 FAIL, 16 critical failures, 4 failure clusters.
- Fixed fixture: 16 PASS, 0 critical failures, 0 failure clusters.
- Comparison: 16 critical failures resolved, 0 introduced, and 4 failure signatures removed.
Runtime
Supported runtime:
- Linux;
- CPython 3.11 or 3.12.
After wheel download and installation, the judge flow requires no OpenAI API key, GPU, local model, or private LAB infrastructure.
Scope and limitations
This release:
- uses synthetic vulnerable and fixed control fixtures;
- demonstrates exactly four attack classes;
- evaluates command-driven targets that emit the documented machine-event contract;
- is not production-ready;
- is not comprehensive agent-security proof;
- does not claim support for arbitrary agents;
- uses SHA-256 identity and canonical conflict checks, not digitally signed evidence custody.
Release identity
- Tag:
v0.2.0 - Wheel:
plutonus_llmfuzz-0.2.0-py3-none-any.whl - Wheel SHA-256:
35c3c9e6cf6ff8d6501a1de66c7a5871247773a95715d0e8a1a7d33f163bfd47