-
Notifications
You must be signed in to change notification settings - Fork 0
Taxonomy and OWASP Mapping
aisrf/taxonomy.py is the single security taxonomy shared by analyzers, probes, policy, campaign summaries and reports. Every finding category and every probe category is mapped to three external frameworks:
- the OWASP Top 10 for LLM Applications, 2025 edition (
LLM01toLLM10); - the threat model of Greshake et al. 2023, "Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection" (threat classes and delivery methods);
- the technique classes of Joseph Thacker's Prompt Injection Primer for Engineers.
The full catalogue is served by GET /api/settings/taxonomy (taxonomy.catalogue()), which returns owasp, greshake_threats, greshake_delivery, thacker_techniques, category_map and references. The dashboard uses it to render taxonomy chips next to findings, probe categories and heatmap rows, and to populate the category selector in Custom-Rules and the severity override table in Settings-Center.
OWASP_LLM_TOP10 holds name, description and URL for each id. The last column lists every AISRF category mapped to the id in CATEGORY_MAP.
| Id | Name | Description | AISRF categories |
|---|---|---|---|
| LLM01 | Prompt Injection | User or third-party content manipulates model behaviour, bypassing intent and controls. |
prompt_injection, indirect_prompt_injection, jailbreak, harmful_content, harmful_compliance, obfuscation, encoding_attacks, multi_turn, many_shot, guardrail
|
| LLM02 | Sensitive Information Disclosure | PII, credentials, proprietary data or confidential business information exposed via outputs. |
system_prompt_extraction, data_exfil, data_exfiltration, pii, pii_leakage, secrets, privacy_memorization
|
| LLM03 | Supply Chain | Compromised models, datasets, plugins or third-party components. | supply_chain |
| LLM04 | Data and Model Poisoning | Manipulated training, fine-tuning or embedding data introduces vulnerabilities, backdoors or bias. |
bias_fairness, privacy_memorization, rag_poisoning
|
| LLM05 | Improper Output Handling | Model output is passed to downstream systems without validation, enabling XSS, SSRF, code execution. |
data_exfil, data_exfiltration, tool_abuse, output_handling, code_safety
|
| LLM06 | Excessive Agency | Over-permissioned tools, functionality or autonomy let the model take damaging actions. |
tool_abuse, excessive_agency
|
| LLM07 | System Prompt Leakage | Secrets, credentials or sensitive logic contained in system prompts are exposed. |
system_prompt_extraction, system_prompt_leak, secrets, canary_leak
|
| LLM08 | Vector and Embedding Weaknesses | Weaknesses in retrieval augmented generation pipelines: poisoned or leaked embeddings, cross-tenant access. |
indirect_prompt_injection, rag_poisoning
|
| LLM09 | Misinformation | Confident false or misleading output, hallucinated facts, unsafe code suggestions. |
harmful_content, misinformation, misinformation_hallucination, bias_fairness, code_safety
|
| LLM10 | Unbounded Consumption | Resource exhaustion, denial of wallet, model extraction through excessive or uncontrolled usage. |
denial_of_wallet, anomaly
|
Reference URLs follow the pattern https://genai.owasp.org/llmrisk/llm01-prompt-injection/ and https://genai.owasp.org/llmrisk/llm0N2025-<slug>/ for the other ids.
GRESHAKE_THREATS describes what an injection achieves once it lands in an LLM-integrated application.
| Class | Description |
|---|---|
information_gathering |
Exfiltrating user data, credentials or conversation history to the attacker. |
fraud |
Phishing, scams and social engineering delivered through the trusted assistant. |
intrusion |
Gaining persistence or control over the application, its tools or the user's systems. |
malware |
Spreading injections to other users, agents or documents (worm-like propagation). |
manipulated_content |
Biased, wrong or attacker-controlled answers, ads or summaries. |
availability |
Denial of service: blocking, slowing or breaking the assistant for the user. |
GRESHAKE_DELIVERY describes how an injection reaches the model. These are part of the catalogue for reference and reporting; CATEGORY_MAP only uses hidden (for obfuscation and encoding_attacks).
| Method | Description |
|---|---|
passive |
Injection placed in public content the model later retrieves (web pages, documents). |
active |
Injection actively delivered to the model (emails, messages, tickets processed by the agent). |
user_driven |
The victim is tricked into pasting the injection themselves. |
hidden |
Injection concealed from humans: white text, comments, encodings, multi-stage payloads. |
THACKER_TECHNIQUES classifies the attacker's technique.
| Class | Description |
|---|---|
direct_injection |
Attacker types the injection into the model input directly. |
indirect_injection |
Injection arrives through data the model consumes (retrieval, tools, browsing). |
context_leak |
Extracting system prompts, hidden context or previous conversations. |
data_exfiltration |
Using markdown images, links or tool calls to send data to the attacker. |
privilege_escalation |
Abusing tools, plugins or agents to act beyond the user's intent. |
persistence |
Storing an injection in memory, documents or databases to re-trigger later. |
obfuscation |
Encodings, translations, token smuggling and splitting to evade filters. |
social_engineering |
Role play, authority claims and urgency to talk the model out of its rules. |
CATEGORY_MAP is keyed by AISRF category. Analyzer findings use the first block of keys (Analyzers), probe corpus files use the second (Probe-Corpus); some names exist in both spellings so that findings and probes line up in reports.
| Category | OWASP | Greshake threats | Thacker techniques |
|---|---|---|---|
prompt_injection |
LLM01 | intrusion, manipulated_content | direct_injection, social_engineering |
indirect_prompt_injection |
LLM01, LLM08 | intrusion, malware, information_gathering | indirect_injection, persistence |
jailbreak |
LLM01 | manipulated_content | social_engineering, direct_injection |
system_prompt_extraction |
LLM07, LLM02 | information_gathering | context_leak |
system_prompt_leak |
LLM07 | information_gathering | context_leak |
data_exfil |
LLM02, LLM05 | information_gathering | data_exfiltration |
data_exfiltration |
LLM02, LLM05 | information_gathering | data_exfiltration |
pii |
LLM02 | information_gathering | data_exfiltration |
pii_leakage |
LLM02 | information_gathering | context_leak |
secrets |
LLM02, LLM07 | information_gathering | context_leak |
canary_leak |
LLM07 | information_gathering | context_leak |
tool_abuse |
LLM06, LLM05 | intrusion | privilege_escalation |
excessive_agency |
LLM06 | intrusion | privilege_escalation |
harmful_content |
LLM01, LLM09 | manipulated_content | social_engineering |
harmful_compliance |
LLM01 | manipulated_content | social_engineering |
obfuscation |
LLM01 | hidden | obfuscation |
encoding_attacks |
LLM01 | hidden | obfuscation |
output_handling |
LLM05 | intrusion | data_exfiltration |
misinformation |
LLM09 | manipulated_content | (none) |
misinformation_hallucination |
LLM09 | manipulated_content | (none) |
bias_fairness |
LLM09, LLM04 | manipulated_content | (none) |
privacy_memorization |
LLM02, LLM04 | information_gathering | context_leak |
code_safety |
LLM09, LLM05 | intrusion | (none) |
denial_of_wallet |
LLM10 | availability | (none) |
anomaly |
LLM10 | availability | (none) |
rag_poisoning |
LLM08, LLM04 | malware, manipulated_content | indirect_injection, persistence |
supply_chain |
LLM03 | intrusion | (none) |
multi_turn |
LLM01 | manipulated_content | social_engineering |
many_shot |
LLM01 | manipulated_content | social_engineering |
policy |
(none) | (none) | (none) |
benign_control |
(none) | (none) | (none) |
refusal |
(none) | (none) | (none) |
guardrail |
LLM01 | (none) | (none) |
policy (custom rules, refusal findings), benign_control (the false-refusal control set) and refusal intentionally carry no mapping so they never inflate OWASP coverage. guardrail is the fallback category for external guardrail findings that cannot be mapped to a more specific one (Guardrails). The greshake column in the map mixes threat classes and, for the obfuscation categories, the delivery method hidden.
-
classify(category)lower-cases and strips the key, tries an exact match, then falls back to the first entry whose key is a substring of the category or contains it (sollm_guard_piiresolves topii, andsecretresolves tosecrets). Unknown categories return empty lists. -
owasp_ids(category)returns the OWASP id list;owasp_label("LLM07")returnsLLM07: System Prompt Leakage.
enrich(dict) attaches four keys in place: owasp (list of ids), owasp_labels (list of LLMnn: Name strings), greshake and thacker.
- Every
Finding.to_dict()call inaisrf/analysis/base.pyrunsenrich, so each finding stored on a ticket, returned by the API, shown in the dashboard or exported in a report carries the labels. Custom rule findings are enriched through the category chosen in the rule editor. - Probe results and campaign categories are labelled through the same map. The dashboard's
A.taxonomyChips({category})helper renders the chips for probe categories on the campaign page, the corpus browser and the heatmap rows of a Comparison-Groups page; reports (Reports) group findings and probe outcomes by OWASP id with the same function. - Guardrail integrations map their own detector names onto AISRF categories first (for example Lakera
prompt_attacktoprompt_injection, LLM GuardMaliciousURLstodata_exfil), so their findings inherit the OWASP mapping of that category.
coverage(categories) answers "which OWASP items does this set of categories exercise". For every id in OWASP_LLM_TOP10 it returns:
"LLM07": {"name": "System Prompt Leakage", "covered": true, "categories": ["system_prompt_extraction"]}An id is covered when at least one input category maps to it; categories lists the sorted contributing categories. Ids that no category maps to are returned with covered: false and an empty list, so consumers always see all ten rows.
Where it is used:
-
build_summaryinaisrf/redteam/engine.pystoresowasp_coverage(computed over the categories present in the campaign) andowasp_breakdown(tested, vulnerable, rate and categories per OWASP id, aggregated over the probe results mapped to it) in every campaign summary. See Red-Teaming for the exact keys. -
compare_groupstoresowasp_coveragefor the union of categories in a comparison group; the group page renders covered and uncovered chips under "OWASP LLM Top 10 coverage". - Reports use the coverage rows to show which parts of the Top 10 a campaign or a code review actually tested; an uncovered id is a gap in the test plan, not a clean result.
Because analyzers, probes and code review rules (Code-Review-Rules) all speak this vocabulary, a single OWASP id can be followed from a static finding in a repository, through a live gateway finding, to a red-team probe verdict in the same report.
REFERENCES lists the sources cited in reports: OWASP Top 10 for LLM Applications 2025, Greshake et al. 2023 (arXiv 2302.12173), Thacker's Prompt Injection Primer, garak, promptfoo, PyRIT, Rebuff, LLM Guard, NeMo Guardrails and Lakera Guard.
AISRF, AI Security & Research Framework. github.com/keyuraghao/aisrf, Apache License 2.0.
Start
Gateway
- Gateway-Endpoints-and-Headers
- Request-Normalization
- Policy-Engine
- Agents-and-Credentials
- Configuration-Reference
- Settings-Center
Review
Security analysis
Red teaming
Code review
Interfaces
Operations
Project