Added
-
tests/rss_ceiling.rs: a guard on the memory non-negotiable, honest about
what it cannot check. It does not run the 512 MiB measurement — the tree
that ceiling is stated against is not in the corpus repository, so CI never
executes it, and the file's own comment says so rather than implying
otherwise. Nor does any Go corpus stand in for it: three runs of each build on
go/caddymeasure 55,752–55,952 kB before the change and 55,808–56,320 kB
after, and ongo/codeiq46,184–46,208 kB against 46,168–46,296 kB — one
build's own spread covers the difference, so no threshold on a Go tree could
separate them.Three non-Go corpora can, which is recorded rather than glossed: the change
movespython/djangofrom 68,908–69,020 kB to 59,160–59,288 kB,
javascript/fastifyfrom 26,112–26,368 kB to 19,688–20,080 kB and
javascript/expressfrom 16,172–16,336 kB to 14,592 kB flat — gaps of
9,620 kB, 6,032 kB and 1,580 kB, where no build's own spread on any of them
exceeded 512 kB. What they cannot do is stand in for the ceiling — they
separate the two builds by ~10,000 kB where the tree the ceiling is stated
against separates them by 545,596 kB — and an RSS threshold is a number
about the hardware and the allocator, which is why it is not asserted on CI
runners that are not the reference hardware.So it pins the mechanism, which is what a future change would break: counting
extractor invocations through the publicscanentry, a Go file is extracted
exactly twice and a Java file exactly three times. A count of one is the
regression. It fails with one on the parent commit, verified by running it
there, and takes 0.01 s with no corpus. What it does not catch is named in
its own comment: re-adding the references to the retained record while keeping
the re-read leaves both counts where they are. -
arthron pin, and a committed target pin per tier-1 corpus: the gate can
now see a wrong edge.arthron gatecompares four integers —resolved,
external,local_binding,unresolved— and a reference that resolves to
the wrong definition moves none of them. It is still oneResolvedrow and
still one edge; only the far end changed. The rate cannot see it, the
denominator_shrankcheck cannot see it, and neither drift check can, because
not one of them reads a target. That blind spot is the whole of the standing
verdict that a wrong edge is worse than a miss, because a miss is counted and
a wrong edge is not.pins/<corpus>.pinsrecords the target of every resolved reference row, per
corpus, for the fifteen tier-1 corpora — 80,272 rows over 18,687 distinct
targets.tests/edge_pins.rsscans each corpus cold and compares. The rule:- a pinned row whose target changed fails, by name —
target_moved—
printing the file, the line, the reference kind, the enclosing FQN, the
site text, the arity, the old target and the new one; - a row that appeared is coverage growth and passes;
- a row that vanished is flagged in the output and does not fail: the
counting gate owns that half, and a re-pin that drops rows shows the drop
as deleted lines in the pin file's own diff.
The format, and why. Written out in full those rows are 14.5 MB of
committed text (measured). Stored as a 64-bit hash of the store's own
canonical row key plus an index into a plaintext, deduplicated dictionary of
target names, the fifteen files are 3.1 MB. The target names stay in
plaintext deliberately: a check whose failure printed
0x8f3a… became 0x1c07…would tell a reviewer that something moved and
nothing about what, so the old target is recoverable by name from the file
and the new one — with the file, line and site text — is re-derived by
joining the failing hash against the scan already in hand. Rows are grouped
under their file so a re-pin's diff says which files' edges moved.Regeneration is one command per corpus, and every pin file carries it in
its own header:arthron pin corpus/go/codeiq --pins pins/go-codeiq.pins --write --commit <sha>Verified by mutation, not by assertion. Making the Java member walk prefer a
superclass's declaration over the receiver's own override left
java-commons-langatresolved 34217 / unresolved 16279 / external 63385 / local_binding 15162andjava-gsonat12885 / 6105 / 16737 / 6706—
byte-identical to their baselines, both gates green — while the pin check
failed both, naming 275 and 266 moved edges with 0 appeared and 0 vanished
(ImmutableTriple.ofresolving toTriple.of,new JsonTreeReader
resolving toJsonReader.<init>). Reverting made all 21 pin tests green
again.No workflow change:
.github/workflows/gate.ymlalready ends by running the
whole suite withARTHRON_REQUIRE_CORPUS=1in the one job that fetches the
corpus, so the fifteen comparisons run and block a merge there. Fifteen cold
scans, 7.4 s wall. - a pinned row whose target changed fails, by name —
Changed
-
A cold scan no longer holds the whole tree's references, and peak RSS drops
from 158.4% of the 512 MiB ceiling to 54.7% of it. Teaching Go to emit type
uses, non-call selector reads and composite-literal keys took a 5,353,211-line
Go tree from 595,892 references to 1,678,021, and peak RSS with it: 830,612 kB
against a hard 524,288 kB ceiling on the 2 vCPU reference hardware.Live-byte accounting that closed to 97.7% of measured VmRSS put 89.8% of the
813.2 MiB peak — 729.9 MiB — in one place: the changed set's references,
extracted by the walk and kept alive until phase 2 consumed them. RSS climbed
monotonically to 729.9 MiB with the database still untouched, and minor faults
stopped at t≈38 s, so phase 2 never asked the kernel for another page. The
peak was the whole parsed tree, held.The driver now keeps a file's path, hash, header and declarations, and not
its references. Declarations are 72,362 against 1,678,021 references and cost
13.3 MiB in total; phase 1 needs every changed file's before it names
anything, and nothing needs every file's references at once. Each later phase
reads the file again — twice per file for a language with no link kinds, three
times for one that declares them and so runs a supertype phase. What the walk
ends up holding is larger than those 13.3 MiB — 110,896 kB, measured, against
a 10,240 kB fixed cost — because a path, a hash and a language header ride
beside every file's declarations; the rest of that retained set has not been
decomposed, anddocs/decisions.mdsays so rather than leaving the
declaration figure to be read as the whole of it.peak RSS wall of ceiling before 830,612 kB 70.59 s 158.4% shrink_to_fiton the extractor's vectors778,328 kB 70.57 s 148.5% and no retained references 286,872 kB 109.62 s 54.7% The last row is the worst of nine runs of the shipped build, which spanned
284,612–286,872 kB and 106.47–109.62 s; the first is reproduced at
832,468 kB / 69.54 s on a re-measurement of the same commit.Re-extraction cannot change a resolution, and the signature is the
enforcement:Extractor::extracttakes a path and a string — no probe, no
config, no other file — so the same bytes give the same facts. The bytes are
re-hashed on the second read, and a file that moved under the scan goes to the
existingstalepath rather than being resolved against declarations its
source no longer makes.Nothing about the graph moved. All 29 corpus gates and all 15 pin
comparisons exit 0 against the committed baselines and pin files, and the
complete stdout of all 44 runs — every rate, every reason tally, every
held/appeared/vanished/movedcount — is byte-identical to the same 44 runs
on the parent commit.The cost is one extra parse per file: 39.0 s more wall clock on that tree,
20.5 s per 1M lines against a 60 s target — and it is not the same margin
in every language. Median of three runs each, per 1M lines:go/caddy19.4 →
32.6,java/commons-lang24.5 → 44.4,javascript/fastify19.2 → 37.1,
python/django18.4 → 42.7, all still inside the target, against
typescript/vue-core36.7 → 70.1 andtypescript/zod45.3 → 82.5,
both now outside it. TypeScript cold indexing misses the timing target on
this build. Timing is a target and the ceiling is hard, so the trade stands;
the miss is recorded here and inREADME.mdrather than inferred from the Go
number. Warm scans are not affected: 11.73 s and 12.37 s on the unchanged
5.35M-line tree against 12.35 s and 12.63 s before the change, with warm peak
RSS equal to within 0.2%. -
The EcmaScript universe scope has a second half, and it un-pollutes
NoMatchingDefinition. That reason's contract says the reference was
understood, the lookup table was complete, and the name is absent — "in a
corpus that compiles this should mean our bug, and should sit near zero."
Both halves were false. All 1,728 occurrences on express were five names —
it1,111,describe554,before59,after3,XMLHttpRequest1 — and
13,833 of vue-core's 15,276 were eight more:expect9,930,test2,860,
it609,describe373,beforeEach37,afterEach21,afterAll2,
beforeAll1. They are what a test runner puts in the global scope of the
files it runs, reaching the file with no import because the runner injects
them.The universe scope now models both provenances. The host's half is
unchanged: a name ECMA-262, Node or the web platform declares isExternal,
because the thing on the other end genuinely exists. The new half is a
package's: a name a declared dependency injects is
Unresolved(UnknownPackage), filed against the package the definition is
actually in. Six environments are recognised (mocha, jasmine, jest,
vitest, cypress, qunit), each by its documented global set in full rather
than by the subset that looked unlikely to collide: what makesitmocha's
is not its spelling but the project declaring mocha, checked per file against
package.jsonandtsconfig.json'scompilerOptions.types. A repository
that declares no runner still getsNoMatchingDefinitionfordescribe, and
any declaration or import of the name wins, because the universe scope is
consulted last.No rate moves and nothing is reclassified into
External, which is the
point:UnknownPackageisUnresolved, so every one of these references
stays in both terms. expressNoMatchingDefinition1,728 → 0 with the rate
at 28.99% before and after; vue-core 15,276 → 1,443 with the rate at 48.48%
before and after (13,833 fromNoMatchingDefinitionand 785 from
NeedsTypeInference, the latter beingvi.fnandexpect.any, where the
head decides exactly as it already does forconsole.log). fastify does not
move at all — it injects nothing. The one external that did move is
XMLHttpRequest, absent from the host list whileWebSocket,
AbortControllerandfetchwere all in it: expressexternal701 → 702,
the same shape of omission asError's and worth one row.The written-out import of the same name does not always agree, and the two
channels differ. An environment turns on through eitherpackage.jsonor
tsconfig.json'scompilerOptions.types. Throughtypesthe two spellings
match exactly: zod states"types": ["vitest"]and declares no dependencies
at all, so its 168from "vitest"specifiers, the names bound by them and
the injected spellings all reportUnknownPackage. Throughpackage.json
they do not: a declared dependency is the dependency boundary, so the import
and every name reached through it answerExternal("npm:<pkg>"). vue-core
declares vitest in its rootdevDependenciesand carries both — 95
occurrences of vitest's names underExternal, written as imports, against
14,618 injected underUnresolved(UnknownPackage). The injected half is not
the one that may move:Unresolvedkeeps it in both terms of the rate, where
Externalwould take it out of both and raise the gate without linking
anything. Closing the asymmetry means moving the imported side, which is a
change to whatExternalmeans at the dependency boundary for every package
and is not made here. Both halves are pinned by a test so neither drifts in
silence. -
TypeScript's
compilerOptions.customConditionsis read, and it is the
largest single miss on zod. The condition set handed to NODE
PACKAGE_TARGET_RESOLVEwas hardcoded per dialect and module kind, so a
monorepo that publishes built artefacts and points its own compilation at the
sources instead —"@zod/source"written ahead of"types"in the same
exportsentry, and named inpackages/zod/tsconfig.json— took the
"types"branch. That branch names anindex.d.ctsbeside the manifest that
no scan of the sources can see, so every self-import missed, and every name
reached through one missed with it. The option is now read (flattened through
extends, folded into the config fence) and added to the set for the nearest
TypeScript project, for bothexportsandimportsmaps. It is a set, not
a priority list: NODE matches conditions in the map's own key order, so a
custom condition can only make a branch reachable that was unreachable, and
which branch wins stays the package author's decision.zod:
ModuleNotFound7,822 → 1,resolved10,043 → 17,080, rate 27.24% →
46.33% (+19.09 points), every one of the 7,037 new edges landing in
packages/zod/src/.NoMatchingDefinition524 → 1,123 and
NeedsTypeInference1,576 → 1,761 in the same movement, and neither is a new
miss: a module that could not be found gave every name reached through it one
answer, and a module that is found gives each of them its own — including,
for a namespace re-exported by name, an honest miss this resolver does not
yet follow. No other corpus states acustomConditionsand none of the other
three moved a row.extendsis nearest-wins on stated, not on non-empty."types": []is the
documented way to say "no ambient type packages" and"customConditions": []
says the same about conditions; both were read as "unstated" and silently
given the base's value back, turning an ambient environment on under a config
that turned it off and sending an import down a branchtscwould not take.
Presence of the key now decides. No corpus states either as an empty list, so
no gated number moves — verified by a whole-row join against the previous
scan on all four, which changed nothing.Every changed row on all four corpora was joined whole-row against the
previous scan. The row-key set is byte-identical — nothing added, nothing
removed — and no reference that already resolved changed its target or its
outcome, on any corpus. The 2,194 rows that changed on zod all came out of
ModuleNotFound, the 187 on express and 800 on vue-core all out of the
ambient class above, and fastify changed nothing. -
Go reads a member as well as calling one — the last two un-emitted Go
reference sites, and a re-base of two baselines. The extractor emitted a
reference for a call through a selector and for a written type name, and
nothing for the two other places the Go grammar names a member: a selector
read (pkg.Name,t.field,T.Method,x.y) and a struct literal's
field keys (theFieldinT{Field: v}). Both are now
RefKind::FieldAccessrows — 8,200 selector-read occurrences oncodeiq
and 9,947 oncaddy, 3,134 and 3,776 literal keys — which is every one of
those sites in both corpora, counted at the grammar and matched exactly by
what the store holds.A read resolves the way a call of the same shape already did:
pkg.Name
through the import table, a receiver root asthis(soc.Namereaches
Conn.Name), a root some enclosing block binds asLocalBinding. What is
new is the owner written at the site —T.Method,T{Field: v},
pkg.T{Field: v}— where the member is probed under that owner and a miss is
answered by what the owner is:NeedsReceiverTypewhen the owner is a type
in this repository,NeedsTypeInferencewhen it is not. Go struct fields are
not nodes in this build, so an honest field read lands in the first of those,
exactly as a receiver-rooted call already did. A map or array literal's key
is an expression rather than a member name and is not a reference; an
anonymousstruct{…}has no canonical name, so neither it nor its fields are
nodes and its keys name nothing.Both of those read the type as written, which is the whole of what one
file says.map[K]V{k: v}is rejected because the site writesmap;
type Registry map[K]Vused asRegistry{k: v}is not, because the site
writes a name and the declaration is usually in a sibling file. So a named
map, slice or array type keyed by an identifier is reported as a member of
it — 9 rows / 90 occurrences oncodeiq(CapabilityMatrixin
internal/intelligence/query), 0 oncaddy— and lands in
NeedsReceiverType, inside the denominator, understating the rate rather
than flattering it: 69.5% as measured against 70.0% without those rows.
Closing it needs a fact no single file holds, so it is stated where it is
made (named_type_pathinextract_go.rs, andref-litkeyin
rules/go.yml) rather than fixed by guessing. What is closed is the harm:
a literal key's member is never probed, so no such site can link. A named
non-struct type may carry a method, and a method name and a map-key
constant do not collide the way a method name and a field name do, so the
probe could only ever have found the wrong node — and a wrong edge is
strictly worse than an unresolved reference. Nothing a compiling corpus had
earned is lost by skipping it: a Go struct field is not a node in this
build, so every literal key on both corpora was already unresolved, and all
three Go gates hold their counts to the row.Separately,
NoMatchingDefinitionis now empty on both Go corpora, from
123 rows oncodeiqand 269 oncaddy. Every one of them was a predeclared
type name at a conversion —string(b),int64(n),byte(c)— which Go
writes exactly as it writes a call, so the grammar filed it as a
call_expressionand the resolver checked it against the predeclared
functions. The name was never absent; it was in the universe block, one
list over. That bucket's contract is that the lookup table was complete and
the name missing, which in a corpus that compiles means arthron's own bug, so
a row that does not mean that does not belong in it. A type cannot be called,
so a one-argument call naming one is a conversion and nothing else.Measured, release build, cold store:
baseline resolved unresolved external local_binding rate go-codeiq8,016 → 9,794 884 → 4,295 12,210 → 12,595 4,113 → 9,873 90.1% → 69.5% go-caddy10,208 → 10,585 2,700 → 9,014 19,201 → 21,304 8,252 → 13,181 79.1% → 54.0% go-probes17 0 26 1 100.0% go-probesis byte-identical: it writes no selector read and no keyed
literal. No other baseline is touched — this is a Go rule file and a Go
resolver.A rate that falls here is not a regression. It is the same argument the
LocalBindingunification made in the other direction: what is in the
rate's terms changed, so the two numbers are not measurements of the same
thing. Go's denominator —resolved + unresolved— grew from 8,900 to 14,089
oncodeiqand from 12,908 to 19,599 oncaddy, which is 5,312 and 6,960
new references inside the rate's terms less the 123 and 269 the conversion
fix moved out of it. 1,778 and 377 of the new occurrences resolved to a
definition, and nothing that resolved before stopped resolving.Attributed per row, not inferred from the totals. A whole-row join
between a binary built from the previous commit and this one, keyed
`file + kind + declaration space + enclosing FQN + site text + argument count- locally-bound`, over both corpora:
-
The conversion fix moves 89 pre-existing rows on
codeiq(123
occurrences) and 199 oncaddy(269), every one of them
NoMatchingDefinition → External("go:builtin"), and nothing else: no row
added, none removed, andresolved,local_bindingand every other reason
identical on both sides. -
The two new constructs then change zero pre-existing rows. Every
movement is a new row, all of kindfield-access: 7,436 rows / 11,334
occurrences oncodeiqand 8,332 / 13,723 oncaddy. Split by construct
and outcome —corpus construct resolved external local_binding NeedsReceiverType NeedsTypeInference NeedsExpressionType codeiqselector reads 1,778 51 5,745 205 12 409 codeiqliteral keys 0 211 15 2,908 0 0 caddyselector reads 377 1,184 4,634 2,904 541 307 caddyliteral keys 0 650 287 2,825 6 0 — where each row of the table sums to that construct's whole site count as
counted at the grammar: 8,200 and 3,134 oncodeiq, 9,947 and 3,776 on
caddy.caddy's literal keys sum to 3,768 there and not 3,776 because
the last eight land on two rows the first construct created rather than
on rows of their own:TestBufferingdeclares a function-local
type args, and the readargs.bodyand the four literal keys
args{body: …}are one target,LocalBindingeither way, so each of the
two keys they share carries five occurrences instead of one.
The reference census in
tests/corpus.rsis new and is what makes this
observable next time: a rule that stops being emitted moves no baseline —
the gate compares four occurrence totals and another rule can supply them —
and moves no reason bucket either. Rows and occurrences are now pinned per
kind on both corpora, because a rule that stops deduplicating moves one and
not the other. -
A probe corpus for each of the other four tier-1 languages, and the four
gates that pin them. Go had a synthetic corpus that states a resolver
outcome per named site; Java, JavaScript, TypeScript and Python did not, so
every method-call outcome in four of the five languages was observable only
as its contribution to a real corpus's totals — where a fix and a regression
of the same size cancel.corpus/java/probes,corpus/javascript/probes,
corpus/typescript/probesandcorpus/python/probesare hand-written truth
tables: every call site is asserted by name, hit or miss, with the reason a
miss carries. Baselines,tests/baselines.rsGATEDrows and steps in
.github/workflows/gate.ymlland with them — twenty-five gates become
twenty-nine.baseline resolved external local_binding unresolved rate java-probes13 7 1 1 92.9% javascript-probes6 0 1 2 75.0% typescript-probes12 0 1 3 80.0% python-probes5 0 2 1 83.3% These four and
go-probesare pins, not ratchets. The corpora are
hand-written, so their rates are properties of the fixtures and are not
evidence of a capability; re-basing one to claim a better number would be
claiming something nobody measured. They are in the README's tier-1 table
because they are gated exactly like the others, marked†so no reader takes
them for a sample of real code.The misses are pinned as exactly as the hits, including two that are filed
under the wrong reason.super.greetandthis.greetare the same call one
keyword apart and get different answers:super.walks the writtenextends
into the other module,this.does not. The failure is then reported as
UnindexedSupertype, whose definition insrc/lib.rsrequires the receiver
type to be in-repository, the member to be in no indexed supertype, and some
supertype to be external or unindexed — and on this row all three conjuncts
are false, provably, because the same scan resolvessuper.greetthrough
that supertype. One missing branch causes it:walk_members
(src/track_ecma/resolve.rs) probes the base under a module-local id, so an
imported base always misses, and onlyresolve_supercarries the
import-following fallback. Both ECMAScript dialects assert it, because the
defect belongs to the shared track rather than to either language. A probe is
a truth table, so this is recorded as what is rather than what should be —
recording the mislabel is what stops it surfacing later as an unattributable
movement. Its fix is measured, not guessed: givingresolve_thisthe
fallback movesjavascript-probesto 87.5% (resolved 6 → 7) and
typescript-probesto 86.7% (resolved 12 → 13) and moves nothing on
express, fastify, vue-core or zod, so it is a deliberate re-base of two pins
plus adocs/decisions.mdentry, and no ratchet is touched.TypeScript's probe adds a third row that keeps a different reason on
purpose:this.inner.greetisNoMatchingDefinition, the bucket that in a
corpus which compiles means arthron's own bug — and this corpus compiles
(tsc --noEmit,strict, clean). Three annotations naming a class two lines
of import away land in three different buckets while every one of them
resolves as aTypeUse: the type is read and simply not used to type a
receiver. -
The Python census walks named trees instead of a whole language
directory.tests/python_corpus.rswalkedcorpus/pythonentire — the one
whole-language walk intests/, where every other corpus test names its tree
— which made its constants a function of the corpus repository's contents
rather than of this repository's extractor.gate.ymlchecks the corpus out
atref: main, unpinned, so any commit adding a.pyfile anywhere under
corpus/pythonwould have turned the test red with nothing here changed;
addingcorpus/python/probeswould have done it immediately. It now names
corpus/python/djangoandcorpus/python/flask, restoring the property that
a census moves only when the extractor does. Paths stay relative to
corpus/python, so every module name derived from one is byte-identical and
no constant moves.corpus/python/probesis deliberately outside it: a probe
is pinned row by row intests/python_probes.rs, which is a stronger check
than a total that would blur it. -
One
LocalBindingrule in every tier-1 track, and Go emits type uses — a
deliberate re-base of seven baselines. The ratified rule is that a
reference whose root is a parameter or a local variable names a thing that is
not a node by decision, so it is reported besideexternaland excluded from
both terms of the resolution rate. Go, TypeScript and JavaScript already
read it that way; Java and Python applied it only when the whole target was
the bound name, sof.m()sat outside both rate terms in Go and inside them
in Java, and the two rates were computed over differently-sized denominators.A receiver is not a local, and that half of the rule went the other way.
Go has nothiskeyword — the receiver is the name a method uses to reach
its own value — and Go alone filed a member selected through it as a local
binding. Java, Python, JavaScript and TypeScript all resolvethis.m()/
self.m()by declared-type lookup and count it in both rate terms, so the
commonest shape in object-oriented code sat outside Go's denominator and
inside everyone else's. Go now resolvest.m()against the receiver type its
own signature states, which is the strongest declared-type evidence any of
the five gives; a member the receiver type does not itself declare is
NeedsReceiverTyperather thanNoMatchingDefinition, because this track
indexes neither Go embedding nor struct fields and the lookup table is
therefore not complete.Separately, the Go extractor emitted only calls and imports, so "tier 1: call
sites, imports and type uses" was not true of Go;ref-typenow emits a
reference for every written type position. Baselines are re-based, not
compared, because what is in the rate's terms changed. Measured, release
build, cold store:baseline resolved unresolved external local_binding rate go-codeiq4,467 → 8,016 799 → 884 6,085 → 12,210 4,276 → 4,113 84.8% → 90.1% go-caddy3,006 → 10,208 1,815 → 2,700 9,571 → 19,201 9,425 → 8,252 62.4% → 79.1% go-probes17 0 0 → 26 1 100.0% java-commons-lang39,591 → 34,217 19,093 → 16,279 68,297 → 63,385 2,062 → 15,162 67.5% → 67.8% java-gson16,074 → 12,885 7,215 → 6,105 18,187 → 16,737 957 → 6,706 69.0% → 67.9% python-django19,103 13,764 → 6,185 13,326 826 → 8,405 58.1% → 75.5% python-flask1,192 → 1,185 2,847 → 877 2,336 → 2,317 150 → 2,146 29.5% → 57.5% The other eighteen baselines are byte-identical, including both TypeScript
and both JavaScript corpora and all fourteen tier-2 baselines, whose
local_bindingis still zero.Attributed per reference, not inferred from the totals. In Java and
Python not one reference was added or removed and every reference that
changed outcome moved intolocal_binding— 13,100 on commons-lang (5,374
fromresolved, 4,912 fromexternal, 2,814 from an unresolved reason),
5,749 on gson, 7,579 on django, 1,996 on flask — and nothing moved in any
other direction. In Go two things moved rows and they are separable. Every
added occurrence is a type use, 9,596 on codeiq and 16,544 on caddy, all of
kindtype-use, of which 3,439 and 6,732 resolve to a definition that had no
row at all before. Every pre-existing occurrence that changed its answer is
rooted at a method receiver — 195 on codeiq and 1,349 on caddy, counted at
the extractor by re-rooting — and every one of them leftlocal_binding, 110
and 470 of them forresolvedand the rest forNeedsTypeInference(the
t.a.b()shape, whose real receiver is the fielda) orNeedsReceiverType
(3 and 123). No other Go occurrence changed its outcome or its count. The
changes do not touch each other's languages, measured by re-running each
corpus with the other reverted.A rate that rises here is not an improvement. Excluding a class from both
terms is exactly how a rate rises with nothing linked better, and Python's
does: django'sNeedsTypeInferencefalls 10,256 → 2,677 and flask's 2,119 →
186 because those references are nowlocal_binding, not because any of them
reached a definition. Thelocal_bindingcolumn is gated for drift for this
reason and a re-base has to state it.And what it costs is edges. The 5,374 commons-lang and 3,189 gson
occurrences that moved fromresolvedtolocal_bindingare 13.6% and 19.8%
of those corpora's resolved edges, and they are gone from the graph
arthron queryreads and the MCP server serves — for many of them the
resolver had already produced theNodeId, and the store still holds the
node.LocalBindingdoes not claim the target is unnameable; it claims the
reference is not evidence about cross-file linking, because reaching its
target needs the type of a binding no other file can see. Python loses
almost nothing this way: django 0 and flask 7. -
arthron scanprints the rate's denominator under every language line.
rate denominator 14089 of 36557 references (38.5%):(resolved + unresolved)over every reference the language emitted. Excludingexternal
andlocal_bindingfrom both of the rate's terms is correct and it also
makes the denominator a fraction of the surface — codeiq's Go rate of 69.5%
covers 38.5% of Go's references, fastify's 63.0% covers 14.2% of
JavaScript's — and a rate published without its share reads as a claim about
the whole. The codeiq figures quoted here are the ones this release ships,
not the ones the feature was written against: the Go field-access entry above
moved them inside this same unreleased section. Text report only, onscanand ongate.--jsonis
unchanged and itsschemadoes not move: the document already carries all
four counts, so a consumer derives the share exactly, and a field for
arithmetic is not a field. -
The tier-1 claim is retracted to what is measured. The README, the
changelog and the report line called tier 1 "call-graph resolution". It is
not: method dispatch mostly does not resolve, and the buckets that need a
type environment —NeedsExpressionType,NeedsReceiverType,
NeedsTypeInference,AmbiguousOverload,UnindexedSupertype,
DynamicDispatch— are the majority of what tier 1 leaves unlinked on nine
of the ten real corpora, 51.9% on flask to 100.0% on both Go corpora, the
tenth being vue-core at 42.1%. This entry first said the single reason
NeedsTypeInferencewas "most of what tier 1 leaves unlinked in all five
languages", citing 758 of codeiq's 884 unresolved rows. That was an
over-attribution and is corrected here: it holds for no language once the
reasons are counted — commons-lang is led byAmbiguousOverload(9,218 of
16,279) and gson byNeedsExpressionType(4,713 of 6,105), where
NeedsTypeInferenceis 342 and 72 rows — and codeiq's own leading bucket is
nowNeedsReceiverType, 3,116 of 4,295. What the leading buckets share is
the type environment, which is the work; no one of them stands for it. A call
through a receiver whose type its own signature states does resolve, in all
five, since the locals re-base above. Tier 1 now
reads "definitions, references, and cross-file import and function-call
resolution", and the scan line printstier 1: call, import and type-use resolution— which is what the denominator holds. Nothing measured changed;
no baseline moved.
Fixed
-
arthron pincompared the tree a pin file names against the tree it scanned
with the platform's own path semantics, while the header can only ever hold a
/-separated path. On Windows the two forms differ, so whether a pin file
matched its own tree was decided by the separator rather than by the tree,
and the refusal printed one path each way — reading as if the separator were
the difference. Both sides are now normalised before the comparison and in
the message. Found by the Windows CI job, which had never reached the
positive half of the check. -
Every number the README published was stale, and nothing could have caught
it. Three re-bases moved the tier-1 counts — the locals unification, Go's
field-access surface, and the ECMAScript config and globals work — and four
probe baselines landed, while the tables, thearthron scansample and the
prose around them still stated the pre-wave figures. A gate compares a scan
against a baseline and has no opinion about prose, so all twenty-nine gates
were green over a README that was wrong in every tier-1 row. Both tables are
now re-rendered frombaselines/*.toml: fifteen tier-1 rows including the
five probe pins, fourteen tier-2 rows unchanged, and both derived columns
recomputed rather than carried forward.every_readme_table_row_matches_its_baselineand
every_published_rate_carries_its_denominator_shareintests/baselines.rs
are what make the next drift fail instead of ship. The first re-derives all
eight cells of every row — the four gated counts, the commit pin, the rate
and the denominator share — and asserts one row per committed baseline, so a
baseline with no row and a row with no baseline both fail. The second checks
the commitment the README makes in prose, that no rate is published without
its share, as a shape rather than a sentence. Both readREADME.mdand
baselines/only, so they run in CI where the corpus is absent and every
ratchet skips. -
The retraction over-attributed the gap to one reason. The README, this
changelog andLang::tier's doc comment all saidNeedsTypeInferencewas
most of what tier 1 leaves unresolved, citing 758 of codeiq's 884 unresolved
rows. Counted per reason on all ten real corpora it holds for none of them —
commons-lang is led byAmbiguousOverload(9,218 of 16,279), gson by
NeedsExpressionType(4,713 of 6,105) whereNeedsTypeInferenceis 72, and
codeiq's own leading bucket is nowNeedsReceiverTypeat 3,116 of 4,295.
What is true, and is what the three now say, is that the reasons needing a
type environment are together the majority on nine of the ten, 51.9% on
flask to 100.0% on both Go corpora, with vue-core the tenth at 42.1% because
the injected test-runner globals outnumber them there. Replacing one reason
with the family it belongs to is the same retraction the tier-1 claim already
made, one level down. -
A mixed
.js/.tstree reported a higher JavaScript rate the second time
it was scanned. The track runs two passes over one store, and the wake set
each computes is filtered to the files that pass owns. So when the TypeScript
pass declared an identity a JavaScript row had already probed and missed — a
workspace member whose entry point is a.tsfile is the shape that does it
— applying it withdrew that file's currency claim and no pass in that scan
could give it back. The scan ended with the claim outstanding, the next
scan re-read exactly those files, and they resolved against a store that by
then held the TypeScript definitions. On a two-package fixture the JavaScript
rate is 0% cold and 100% warm, for a tree nobody touched. A rate that depends
on how many times it has been measured is not a measurement, and the cold
number is the one every baseline is taken from. The track now runs JavaScript
once more to converge; that pass's changed set is exactly the files whose
claims are outstanding, which is empty in the ordinary case, and it
terminates because a module's identity here is its path. Measured cost, best
of three interleaved cold runs: express +0.03 s, fastify +0.01 s, vue-core
+0.09 s, zod within noise. The returned report'sfile_errorsare now the
union of all three passes' rather than the last one's — a tally is
whole-store, but a file error belongs to the pass that tried to read the
file. None of the four gated corpora is mixed, so no committed number moves;
the fixture is the gate. -
Two Go definition defects the new type-use surface exposed.
def-type
read onlytype_spec, so a package-leveltype X = Ydeclared no node —
free while Go emitted no type uses, and 57 codeiq / 7 caddy
NoMatchingDefinitionrows the moment it did; it now readstype_aliastoo,
which is the whole of theDefKind::Typecensus moving 229 → 232 on codeiq
and 507 → 511 on caddy. Andcase nil:in a type switch is a
type_identifierin this grammar, now answered from the predeclared block
rather than left unmatched. After both,NoMatchingDefinitionis 123 on
codeiq and 269 on caddy — unchanged from before the wave. -
Generic[T](x)was reported as a type use naming a function, and as no
call at all. An explicit instantiation is unambiguously a type, so the Go
grammar files the whole call as atype_conversion_expressionover a
generic_typeand thecall_expressionrule never saw it; the only row the
site produced was aTypeUsewhose target was aDefKind::Function. The
callee position of a call written in call syntax is now reported as the
Callit is — still exactly one row for one site, and the type arguments
are unaffected. No published number moves: there is no explicit
instantiation incodeiq,caddyorprobes, counted syntactically over
all 728 files. -
Two Java external nodes that claimed a package which does not exist.
Outer.NonStaticInnerandEnclosing<T>.Innerin gson'sTypeTokenTest
name method-local classes (JLS §14.3); their two-segment targets escaped the
narrow local rule and were filed asExternal("Outer")and
External("Enclosing"). gson's stored external census is 36 → 34. -
A repository's
dbmay no longer name a store outside the tree through a
link with nothing on the other end. The containment check canonicalises the
deepest existing component of the resolveddbpath and asks whether it is
under the root. A dangling symbolic link answerslstatand fails
canonicalize, and that failure read as "not there yet", so the walk stepped
past the link to its parent — inside the root — and called the whole path
contained. The store was then created through the link: one arbitrary file,
anywhere the process could write, from a scanned repository's own
arthron.toml, at exit 0 with nothing said. A component that exists and does
not resolve is now refused, because it cannot be shown to stay inside the
root. Reachable fromarthron scanand from the MCPscantool. -
A scan of a root that is not there answers 2 whatever
--dbsays.
Creating the store's directory ran first, and the default store lives at
<root>/.arthron/graph.redb— socreate_dir_allmade the missing root, the
walk found the empty tree it had just made, and the run answered 0 with a
report of zeros. With--dbelsewhere the same invocation already answered 2,
so the exit code depended on where the store happened to sit. The root is now
checked before anything is created. -
A track whose project layout it cannot read reports no tally for its
language. The rows an earlier scan wrote stay in the store — a track that
cannot read the layout is in no position to say which files are gone — but
scanno longer prints their tally beside the line saying the track measured
nothing, andgate --db <persistent store>no longer re-bases a baseline onto
numbers this run did not produce. -
A symbolic link out of the tree is now named as such whatever is on the other
end of it; the message no longer says "definitions" about a directory. -
Documented exit code 2 was narrower than the code. Every place that
described it — the README table,--json's help, the module docs,gate --help— said "nothing was measured: usage, I/O or the environment, safe to
retry".gatealso returns 2 forGateVerdict::Error, a baseline or a run
whoseresolved + unresolvedis zero: measured, deterministic, and not
worth retrying. The exit code is unchanged and the documentation now says
both halves and which is which. -
gate --helpcontradicted itself aboutdb. It refused the config's
dbkey on the grounds that "a gate is only meaningful against a cold
store", and then documented a--dbflag that is honoured as given —
including at a store that already holds a graph, which is re-scanned warm.
The real reason the config key is ignored is that where the run writes is
not the scanned repository's decision;--dbis yours, warm store and all,
which is why the default is a fresh temporary one. -
arthron mcp --helpstated the wrong default forscan_repo'sdb. It
said<path>/.arthron/graph.redb, omitting that the scanned repository's
arthron.tomldbwins first — which the tool's own JSON schema already
said correctly. The two now agree. -
The
--dbcwd-versus-config-root asymmetry is written down. A config's
dbis resolved against the repository it sits in;--dbis resolved
against the current working directory, soarthron scan ./repo --db graph.redbwrites./graph.redb. Documented onscan --db, ongate --db, and in the README'sarthron.tomlsection. Behaviour unchanged. -
CONTEXT.mddefined an edge as a resolved reference only. AnExternal
reference produces an edge too, to the dependency node it reached — that is
what makesquery impactsee a call into a dependency instead of a dead
end. The glossary entry now saysResolvedorExternal, and that
Unresolvedproduces none. -
Two comments still described the tier-2 tracks as disabled —
src/lib.rs
and theREGISTRYlist — from before all fourteen went live. Comments only. -
The 0.0.1 changelog omitted the Windows baseline round-trip fix; it is now
recorded under that release, where it shipped. -
The kubernetes cold-scan RSS that failed the hard gate is quoted as its
measured value, 729.1 MiB, everywhere it appears. It was rounded to
729.0 in the summaries downstream of the benchmark that measured it.