-
Notifications
You must be signed in to change notification settings - Fork 1
research corpus
Issue: #73 ·
Built by: make research-index · Guard: tests/test_research_corpus.c
agents/AGENTS.md § Discipline says every detector names its basis. The
citations exist — in issue bodies, wiki pages and the rule table — but
they are prose scattered across three places, so an operator looking at
a CRIT cannot ask "what backs this" without a browser and the git log.
This is the machine-readable version of that habit: one curated markdown document per source, indexed into SQLite FTS5, queryable at runtime.
research/
cert/vu-871675.md CERT/CC advisories
cve/2023-52160.md NVD records
mitre/T1557.md ATT&CK techniques
ieee/802.11-2020-9.6.14.md clause-annotated summaries
papers/bl0ck-2302.05899.md academic sources
tools/ tool documentation
Each document opens with frontmatter:
---
source_url: https://www.kb.cert.org/vuls/id/871675
retrieved: 2026-08-31
topics: [wpa3, dragonblood, pmf]
alert_kinds: [ALERT_TYPE_WPA_DOWNGRADE]
citation: CERT/CC VU#871675
---source_url and retrieved are required. Without them a search hit
says "something backs this" and cannot say what or when, which is not a
citation. topics, alert_kinds and citation are optional.
The parser accepts scalars and single-line bracketed lists, nothing else. A full YAML parser is a large dependency and a large attack surface for a format we control at both ends — and the failure mode of a permissive parser here is specific and bad: a document that silently indexes under the wrong alert kind, or under none, with nothing downstream to notice.
So the parser refuses rather than half-accepts. A line inside the
frontmatter with no colon is an error, not an ignored line — because a
typo'd alert_kinds key means a document that never surfaces for any
detector, and that failure is invisible. Likewise an over-long list is
refused rather than truncated: a silently dropped alert kind is a
document that covers fewer detectors than it claims.
research_ingest exits non-zero on any unparseable document, so
make research-index fails loudly rather than quietly shipping a corpus
missing the file you just added.
Documents are split on ## headings, one FTS5 row per section, with
anything before the first heading attributed to the # title. A BM25
hit on a 400-line advisory should point at the paragraph that matched,
not at the file.
research.db is in the repository. That is unusual for a generated
artifact and it is deliberate: it is small, it is byte-for-byte
reproducible (documents are visited in sorted path order, so repeated
builds are identical — verified, not assumed), and committing it means
the runtime query works from a fresh clone with no build step.
Regenerate with make research-index after editing anything under
research/. .gitattributes marks it binary so git does not attempt
line diffs or CRLF conversion on it.
tests/test_research_corpus.c checks two directions. Both are
enforced.
No document cites an alert kind that does not exist. Frontmatter
names kinds as strings, so a renamed or deleted ALERT_TYPE_* leaves
documents pointing at nothing and the runtime query returns zero hits
with no indication why.
Every alert kind that can be cited is. Enforced since the content
pass. A new detector arrives uncited and turns the suite red until
someone writes down what it detects from — the rule
agents/AGENTS.md § Discipline states, with a mechanism behind it.
"Can be cited" is doing work. alert_technique() returns "" for a
rule reporting sloth's own operational state rather than an adversary,
and there is no CVE, advisory or clause to cite for one. Those are
excluded rather than counted as gaps — see below.
The suite prints coverage every run:
corpus coverage: 59/59 citable alert kinds cited (1 have no external basis)
and names the offenders when it fails, so the fix is obvious from the output rather than requiring a query.
This check used alert_kinds MATCH <kind> from slice 1 until the
content pass, and that was wrong for three slices.
FTS5's unicode61 tokenizer splits on underscores, and a bare sequence
of terms is a phrase query. So MATCH 'ALERT_TYPE_EVIL_TWIN' is
satisfied by a document naming only ALERT_TYPE_EVIL_TWIN_PROXIMITY —
the shorter kind's tokens are a consecutive prefix of the longer one's.
It is the same trap rq_for_alert hit in slice 2 and was fixed for; the
guard was simply never updated to match. It happened to report the same
number as an exact query, because every kind that is a token-prefix of
another is also independently cited — but that is luck, not correctness,
and a guard that is right by luck is not a guard.
Now delimiter-wrapped LIKE, identical to the query layer's, with
test_coverage_query_is_exact_not_fts_match asserting both that the
exact form rejects the prefix case and that MATCH accepts it. The
second half matters: without it the test documents a preference rather
than a bug.
- Write
research/<class>/<slug>.mdwith the frontmatter above. -
make research-index. -
make test— the guard will reject an alert kind that does not exist, and the coverage line will tick up. - Commit both the document and the regenerated
research.db.
research/query.c implements the four operations #73 specifies —
search, for_alert, cite, recent — as a library, not only as
an MCP server. sloth links it directly:
sloth --with-research research.db
That is a departure from the issue, which has sloth spawn a subprocess and speak JSON-RPC to ask about its own data file. One implementation, two front doors: sloth links it, and the MCP server below is a transport wrapper for the external consumers MCP is actually for — Claude sessions and scheduled tasks querying a corpus they did not build.
Additive by construction. A corpus that is missing, unreadable or carrying the wrong schema logs one line and leaves sloth running with no research context:
sloth: research corpus unavailable: cannot open /nope/x.db — continuing without it
Every entry point tolerates a NULL handle and returns zero results, and the no-SQLite build stubs the whole layer to no-ops. Nothing in the capture path may depend on research context.
Alert-kind lookup cannot use FTS5 MATCH. unicode61 splits on
underscores, so alert_kinds MATCH 'ALERT_TYPE_EVIL_TWIN' needs only
the tokens alert, type, evil, twin to be present — and a
document naming only ALERT_TYPE_EVIL_TWIN_PROXIMITY contains every
one. The References block for one alert would cite a document about a
different one, which is worse than citing nothing because it looks
right. rq_for_alert and rq_cite use delimiter-wrapped LIKE
against the stored list instead, which is exact.
Free-text search still goes through MATCH — fuzzy is the point there.
Only its filter is exact.
for_alert returns documents, not sections. The index stores a row
per heading, so a document with four matching sections would otherwise
appear four times in a References block and read as four citations.
for_alert and cite order by retrieved date and path, never by BM25.
A --report regenerated tomorrow must be byte-identical to today's, and
relevance scores shift as the corpus grows.
sloth-research-mcp exposes the same four functions over MCP, for
consumers that are not sloth. It is not part of all:
make research-mcp
./sloth-research-mcp --db research.db
Register it with an MCP client:
{ "mcpServers": {
"sloth-research": {
"command": "/path/to/sloth/sloth-research-mcp",
"args": ["--db", "/path/to/sloth/research.db"]
} } }Four tools: research_search, research_for_alert, research_cite,
research_recent. Their descriptions say explicitly that for_alert
wants the enum name and not the display title, because that distinction
already cost three bugs on sloth's own side of the same query layer.
A missing corpus is not a startup failure. The server still answers
initialize and tools/list, and every tool call reports why it has
nothing. A client that cannot start its server sees a connection error,
which says far less than "the corpus is not built".
| file | what it owns |
|---|---|
research/mcp/json.c |
reading JSON |
research/mcp/mcp.c |
one request → one response |
research/mcp/main.c |
the pipe, and nothing else |
mcp_handle() is a pure function of (corpus, request, clock), so the
protocol has real tests. A dispatcher that only existed inside a read
loop could only be tested by spawning a process and talking to it —
slower, flakier, and it catches less.
This tree had no JSON parser: jsonl.c writes and cannot read. The
two options were vendoring one into a codebase that has carried no
third-party source, or scanning the raw text for "key":.
The scan is wrong in a way that matters here. A request whose
arguments contain the string "name" — a search for alert "name" field, say — would have its tool name read out of the user's own query.
Input arriving over a pipe from something other than us is exactly where
that stops being hypothetical. So: a bounded recursive-descent parser,
no allocation, ~300 lines, with the subset MCP needs.
What it deliberately refuses rather than accepts loosely:
-
\uescapes above U+00FF — refused, not folded to?or split into a broken byte pair. - Raw control characters inside strings. The transport is newline-delimited, so a raw newline would let one request masquerade as two.
- Leading zeros,
+1,nan,inf— all of whichstrtodaccepts and JSON does not. - Trailing content after the root value.
{...} {...}on one line means the framing already went wrong upstream. - Anything past the node, text or depth caps — an error, never a truncation.
A half-parsed request answered as if it were whole is the failure mode worth engineering against.
| situation | shape |
|---|---|
| unparseable, no method, unknown method, unknown tool | JSON-RPC error
|
| missing argument, corpus unavailable |
result with isError: true
|
| nothing matched |
result with isError: false
|
The middle row is the MCP convention and it is the useful one: the model sees the reason and can correct itself, instead of the transport failing. The last row matters just as much — reporting "no results" as a failure would train a client to retry a query that will never succeed.
A message with no id is a notification and draws no reply at all.
With a corpus loaded, the report gains a References section listing the sources behind each alert that fired:
## References
**BTM_ABUSE**
- [IEEE 802.11-2020 §9.6.14 — BSS Transition Management](https://…) — retrieved 2026-08-31The lookup key is alert_type_name(), not the alert's display
title. Titles are capped at ALERT_TITLE_LEN and abbreviated to fit —
BLOCKACK_ATK for ALERT_TYPE_BLOCKACK_ATTACK, PEAP_NO_CERT for
ALERT_TYPE_PEAP_NO_SERVER_CERT — so keying on them silently loses
citations for every alert whose title was shortened, and a partly-empty
References block looks exactly like a complete one.
An alert with no documents emits nothing rather than an empty heading. That is most of them: 59 of 60 kinds are cited.
Slice 3. One row per alert kind that has fired, with the sources behind
it — a companion to [v] Alerts, answering the question that view
raises: why should I believe this?
The uncited rows are the point. A CRIT with no document behind it is a
behavioural threshold with no cited basis, which agents/AGENTS.md says
is indistinguishable from a guess — and the operator deciding whether to
act needs to know which they are looking at. A view that showed only the
covered kinds would answer the easy half of the question.
research/coverage.c builds the table after alerts_update() on each
poll, through the same rq_for_alert() the References block uses. With
no corpus loaded it still lists every fired kind with no sources, and
the header says not loaded rather than leaving the reader to infer it:
"nothing is cited" and "no corpus is loaded" look identical if the table
is simply empty.
Full write-up in docs/views/research.md.
ALERT_TYPE_NO_MONITOR_MODE reports that sloth has no monitor-mode
radio. It is sloth's own operational state, not a claim about an
adversary, and there is no CVE, advisory, technique or clause to cite
for it.
alert_technique() already said so — it returns "" for exactly this
case — so that is where the fact lives, and research/coverage.c reads
it rather than keeping a second list the two could disagree about.
The distinction matters to the number. Counted as a gap, coverage reads
59 of 60 forever and a target that cannot be met stops being read. The
[f] view therefore shows n/a for such a kind rather than -,
and excludes it from both halves of the ratio: nothing to cite and
nothing cited are different claims and must not be coloured the same.
Any future detector that reports sloth's own state rather than the network's gets this treatment automatically, by having no technique.
- Coverage. 59 of 60 alert kinds have a document, and the sixtieth never will — see below. The guard can now be flipped from warning-only whenever you want it enforcing.
The view key is [f], not the [q] the issue proposed: q is the
quit key, checked before the view switch as an absolute global. c and
f are the only free letters.
Mirrored from docs/wiki/ on main by .github/scripts/wiki_sync.sh. Edit there, not here — hand edits to this wiki are overwritten on the next push.
Read this first — the complete reference
- what-sloth-does
- how-wifi-works
- monitor-mode
- where-exploits-happen
- wifi-sigint-techniques
- cli-reference
- wifi-state-of-the-art
Start here
Engines
WiFi SIGINT
- wifi-sigint
- non-ip-sensors
- mac-randomisation
- evil-twin-reproducer
- btm-abuse
- action-frames
- research-corpus
- captive-portal
- fragattacks
- tool-fingerprints
- enterprise-rogue
- ipv6-ndp
- smb-snoop
- kerberos-snoop
- ldap-snoop
- bgp-snoop
- ssh-snoop
- rdp-snoop
- snmp-snoop
- mqtt-snoop
UI and infrastructure
- ip-palette
- platform-vtable
- version-checkin
- manifest-format
- pcap-export
- jsonl-schema
- data-socket-exposure
- sqlite-schema
- ring-buffers
Factory infrastructure
Reference
Source material
Maintenance