Skip to content

research corpus

sloth wiki-sync edited this page Sep 30, 2026 · 1 revision

The research corpus

Issue: #73 · Built by: make research-index · Guard: tests/test_research_corpus.c

agents/AGENTS.md § Discipline says every detector names its basis. The citations exist — in issue bodies, wiki pages and the rule table — but they are prose scattered across three places, so an operator looking at a CRIT cannot ask "what backs this" without a browser and the git log.

This is the machine-readable version of that habit: one curated markdown document per source, indexed into SQLite FTS5, queryable at runtime.

Layout

research/
  cert/vu-871675.md              CERT/CC advisories
  cve/2023-52160.md              NVD records
  mitre/T1557.md                 ATT&CK techniques
  ieee/802.11-2020-9.6.14.md     clause-annotated summaries
  papers/bl0ck-2302.05899.md     academic sources
  tools/                         tool documentation

Each document opens with frontmatter:

---
source_url: https://www.kb.cert.org/vuls/id/871675
retrieved: 2026-08-31
topics: [wpa3, dragonblood, pmf]
alert_kinds: [ALERT_TYPE_WPA_DOWNGRADE]
citation: CERT/CC VU#871675
---

source_url and retrieved are required. Without them a search hit says "something backs this" and cannot say what or when, which is not a citation. topics, alert_kinds and citation are optional.

Why a YAML subset, not YAML

The parser accepts scalars and single-line bracketed lists, nothing else. A full YAML parser is a large dependency and a large attack surface for a format we control at both ends — and the failure mode of a permissive parser here is specific and bad: a document that silently indexes under the wrong alert kind, or under none, with nothing downstream to notice.

So the parser refuses rather than half-accepts. A line inside the frontmatter with no colon is an error, not an ignored line — because a typo'd alert_kinds key means a document that never surfaces for any detector, and that failure is invisible. Likewise an over-long list is refused rather than truncated: a silently dropped alert kind is a document that covers fewer detectors than it claims.

research_ingest exits non-zero on any unparseable document, so make research-index fails loudly rather than quietly shipping a corpus missing the file you just added.

Sections, not files

Documents are split on ## headings, one FTS5 row per section, with anything before the first heading attributed to the # title. A BM25 hit on a 400-line advisory should point at the paragraph that matched, not at the file.

Why the index is committed

research.db is in the repository. That is unusual for a generated artifact and it is deliberate: it is small, it is byte-for-byte reproducible (documents are visited in sorted path order, so repeated builds are identical — verified, not assumed), and committing it means the runtime query works from a fresh clone with no build step.

Regenerate with make research-index after editing anything under research/. .gitattributes marks it binary so git does not attempt line diffs or CRLF conversion on it.

The guard

tests/test_research_corpus.c checks two directions. Both are enforced.

No document cites an alert kind that does not exist. Frontmatter names kinds as strings, so a renamed or deleted ALERT_TYPE_* leaves documents pointing at nothing and the runtime query returns zero hits with no indication why.

Every alert kind that can be cited is. Enforced since the content pass. A new detector arrives uncited and turns the suite red until someone writes down what it detects from — the rule agents/AGENTS.md § Discipline states, with a mechanism behind it.

"Can be cited" is doing work. alert_technique() returns "" for a rule reporting sloth's own operational state rather than an adversary, and there is no CVE, advisory or clause to cite for one. Those are excluded rather than counted as gaps — see below.

The suite prints coverage every run:

    corpus coverage: 59/59 citable alert kinds cited (1 have no external basis)

and names the offenders when it fails, so the fix is obvious from the output rather than requiring a query.

The tokenizer trap, again

This check used alert_kinds MATCH <kind> from slice 1 until the content pass, and that was wrong for three slices.

FTS5's unicode61 tokenizer splits on underscores, and a bare sequence of terms is a phrase query. So MATCH 'ALERT_TYPE_EVIL_TWIN' is satisfied by a document naming only ALERT_TYPE_EVIL_TWIN_PROXIMITY — the shorter kind's tokens are a consecutive prefix of the longer one's.

It is the same trap rq_for_alert hit in slice 2 and was fixed for; the guard was simply never updated to match. It happened to report the same number as an exact query, because every kind that is a token-prefix of another is also independently cited — but that is luck, not correctness, and a guard that is right by luck is not a guard.

Now delimiter-wrapped LIKE, identical to the query layer's, with test_coverage_query_is_exact_not_fts_match asserting both that the exact form rejects the prefix case and that MATCH accepts it. The second half matters: without it the test documents a preference rather than a bug.

Adding a document

  1. Write research/<class>/<slug>.md with the frontmatter above.
  2. make research-index.
  3. make test — the guard will reject an alert kind that does not exist, and the coverage line will tick up.
  4. Commit both the document and the regenerated research.db.

Querying it

research/query.c implements the four operations #73 specifies — search, for_alert, cite, recent — as a library, not only as an MCP server. sloth links it directly:

sloth --with-research research.db

That is a departure from the issue, which has sloth spawn a subprocess and speak JSON-RPC to ask about its own data file. One implementation, two front doors: sloth links it, and the MCP server below is a transport wrapper for the external consumers MCP is actually for — Claude sessions and scheduled tasks querying a corpus they did not build.

Additive by construction. A corpus that is missing, unreadable or carrying the wrong schema logs one line and leaves sloth running with no research context:

sloth: research corpus unavailable: cannot open /nope/x.db — continuing without it

Every entry point tolerates a NULL handle and returns zero results, and the no-SQLite build stubs the whole layer to no-ops. Nothing in the capture path may depend on research context.

Two things the tokenizer forces

Alert-kind lookup cannot use FTS5 MATCH. unicode61 splits on underscores, so alert_kinds MATCH 'ALERT_TYPE_EVIL_TWIN' needs only the tokens alert, type, evil, twin to be present — and a document naming only ALERT_TYPE_EVIL_TWIN_PROXIMITY contains every one. The References block for one alert would cite a document about a different one, which is worse than citing nothing because it looks right. rq_for_alert and rq_cite use delimiter-wrapped LIKE against the stored list instead, which is exact.

Free-text search still goes through MATCH — fuzzy is the point there. Only its filter is exact.

for_alert returns documents, not sections. The index stores a row per heading, so a document with four matching sections would otherwise appear four times in a References block and read as four citations.

Ordering is stable on purpose

for_alert and cite order by retrieved date and path, never by BM25. A --report regenerated tomorrow must be byte-identical to today's, and relevance scores shift as the corpus grows.

The MCP server

sloth-research-mcp exposes the same four functions over MCP, for consumers that are not sloth. It is not part of all:

make research-mcp
./sloth-research-mcp --db research.db

Register it with an MCP client:

{ "mcpServers": {
    "sloth-research": {
      "command": "/path/to/sloth/sloth-research-mcp",
      "args": ["--db", "/path/to/sloth/research.db"]
    } } }

Four tools: research_search, research_for_alert, research_cite, research_recent. Their descriptions say explicitly that for_alert wants the enum name and not the display title, because that distinction already cost three bugs on sloth's own side of the same query layer.

A missing corpus is not a startup failure. The server still answers initialize and tools/list, and every tool call reports why it has nothing. A client that cannot start its server sees a connection error, which says far less than "the corpus is not built".

Where the code lives, and why it is split

file what it owns
research/mcp/json.c reading JSON
research/mcp/mcp.c one request → one response
research/mcp/main.c the pipe, and nothing else

mcp_handle() is a pure function of (corpus, request, clock), so the protocol has real tests. A dispatcher that only existed inside a read loop could only be tested by spawning a process and talking to it — slower, flakier, and it catches less.

The JSON reader

This tree had no JSON parser: jsonl.c writes and cannot read. The two options were vendoring one into a codebase that has carried no third-party source, or scanning the raw text for "key":.

The scan is wrong in a way that matters here. A request whose arguments contain the string "name" — a search for alert "name" field, say — would have its tool name read out of the user's own query. Input arriving over a pipe from something other than us is exactly where that stops being hypothetical. So: a bounded recursive-descent parser, no allocation, ~300 lines, with the subset MCP needs.

What it deliberately refuses rather than accepts loosely:

  • \u escapes above U+00FF — refused, not folded to ? or split into a broken byte pair.
  • Raw control characters inside strings. The transport is newline-delimited, so a raw newline would let one request masquerade as two.
  • Leading zeros, +1, nan, inf — all of which strtod accepts and JSON does not.
  • Trailing content after the root value. {...} {...} on one line means the framing already went wrong upstream.
  • Anything past the node, text or depth caps — an error, never a truncation.

A half-parsed request answered as if it were whole is the failure mode worth engineering against.

Errors: protocol vs tool

situation shape
unparseable, no method, unknown method, unknown tool JSON-RPC error
missing argument, corpus unavailable result with isError: true
nothing matched result with isError: false

The middle row is the MCP convention and it is the useful one: the model sees the reason and can correct itself, instead of the transport failing. The last row matters just as much — reporting "no results" as a failure would train a client to retry a query that will never succeed.

A message with no id is a notification and draws no reply at all.

--report References

With a corpus loaded, the report gains a References section listing the sources behind each alert that fired:

## References

**BTM_ABUSE**

- [IEEE 802.11-2020 §9.6.14 — BSS Transition Management](https://…) — retrieved 2026-08-31

The lookup key is alert_type_name(), not the alert's display title. Titles are capped at ALERT_TITLE_LEN and abbreviated to fit — BLOCKACK_ATK for ALERT_TYPE_BLOCKACK_ATTACK, PEAP_NO_CERT for ALERT_TYPE_PEAP_NO_SERVER_CERT — so keying on them silently loses citations for every alert whose title was shortened, and a partly-empty References block looks exactly like a complete one.

An alert with no documents emits nothing rather than an empty heading. That is most of them: 59 of 60 kinds are cited.

The [f] Research view

Slice 3. One row per alert kind that has fired, with the sources behind it — a companion to [v] Alerts, answering the question that view raises: why should I believe this?

The uncited rows are the point. A CRIT with no document behind it is a behavioural threshold with no cited basis, which agents/AGENTS.md says is indistinguishable from a guess — and the operator deciding whether to act needs to know which they are looking at. A view that showed only the covered kinds would answer the easy half of the question.

research/coverage.c builds the table after alerts_update() on each poll, through the same rq_for_alert() the References block uses. With no corpus loaded it still lists every fired kind with no sources, and the header says not loaded rather than leaving the reader to infer it: "nothing is cited" and "no corpus is loaded" look identical if the table is simply empty.

Full write-up in docs/views/research.md.

The kind that will never be cited

ALERT_TYPE_NO_MONITOR_MODE reports that sloth has no monitor-mode radio. It is sloth's own operational state, not a claim about an adversary, and there is no CVE, advisory, technique or clause to cite for it.

alert_technique() already said so — it returns "" for exactly this case — so that is where the fact lives, and research/coverage.c reads it rather than keeping a second list the two could disagree about.

The distinction matters to the number. Counted as a gap, coverage reads 59 of 60 forever and a target that cannot be met stops being read. The [f] view therefore shows n/a for such a kind rather than -, and excludes it from both halves of the ratio: nothing to cite and nothing cited are different claims and must not be coloured the same.

Any future detector that reports sloth's own state rather than the network's gets this treatment automatically, by having no technique.

What is not here yet

  • Coverage. 59 of 60 alert kinds have a document, and the sixtieth never will — see below. The guard can now be flipped from warning-only whenever you want it enforcing.

The view key is [f], not the [q] the issue proposed: q is the quit key, checked before the view switch as an absolute global. c and f are the only free letters.

Clone this wiki locally