v1.4.0 — the ruler's subject now lives here
The foreign ruler graded a repository on one machine. It is now a copy of Flask 3.1.3 carried in this repository — and repeating the measurements against a corpus that holds still reversed two conclusions this project had been publishing.
benchmarks/corpus/flask/ is Flask 3.1.3 at commit 22d9247, BSD-3-Clause, docs/ and .git/ omitted and the rest as published: 1,572 units of code nobody here wrote, indexed with no descriptions. rag-your-code.toml keeps it out of this project's own index — without that, eighteen hundred units of a web framework enter two rulers that measure retrieval over this repository and falsify a third asserting nothing here answers thirty questions. Neither the wheel nor the sdist carries it.
The local model is worse or identical on every ruler. 1.1.0 read the comparison as "better on every positive ruler". Its largest gain was on ruler A, whose two arms had been taken against two states of a repository being edited while the script ran. Repeated with one fingerprint per row: A 0.200/0.286/0.238 against MiniLM's 0.171/0.257/0.214, B identical, C 0.443/0.614/0.507 against 0.429/0.600/0.500. The B and C gains were one or two questions — inside the noise of a two-unit change to the corpus, and they reverse sign under it. The extra stays shipped for the cross-language and paraphrase pairs the hash scores zero on, which these rulers cannot see.
Concentration does not subsume coverage. On the retired subject the two bars were interchangeable. Here they are not: together they silence 0.833 of the foreign absent questions, against 0.800 for concentration alone and 0.733 for coverage alone. One question, and the first measurable reason in four releases to keep both.
Which side wins against Grep depends on the repository. The README said flatly that a Grep loop wins on an undescribed one; that was a single subject generalised. On Flask it reverses — 37.1% right file first against 22.9%, no descriptions added — because Flask has docstrings to retrieve against. What does not depend on the subject is what descriptions buy: 58.6% against 22.9% here.
Foreign silence is 0.833, down from 0.933, and the cause is a limit rather than a defect: a word counts as evidence unless it reaches 5% of units, so how, when and are are ubiquitous in a repository with 304 written descriptions and discriminating across 1,572 short undocumented methods. A sweep says the shipped 0.28 stays — every step up buys that silence out of the answers.
CI now runs the foreign ruler, the foreign absent ruler and the foreign head-to-head, none of which was possible before. On a different OS it reproduced 5fd51169eacc / 0.200 / 0.286 / 0.238 / 0.833 exactly.
354 tests (+1, 0 removed, by node-id set diff against v1.3.0). Full detail in CHANGELOG.md.