gotreesitter v0.49.0
Removed
- One member of the Go result-compatibility arm is retired: the member
that widened the root span across a trailing end-of-file newline.
extendNodeToTrailingWhitespaceruns unconditionally after the
compatibility pass and accepts a superset of the byte set the Go member
tested, so the Go member could never reach a span the later pass did not
already reach.
A census recorded 35,244 gate entries and zero rewrites.
The four pinned canonical Go deep-tree digests are byte-identical, and
the exhaustive C-oracle fresh and incremental parity sweep stays green.
Added
-
FactProgramcompiles selected definition, call, heritage, and import work
into a dense 16-bit instruction stream. One program can process compatible
trees with one traversal while the individual extraction APIs remain stable.
The 20-seed combined benchmark reduced Go tree inspection time by 71.19%
for definitions and calls. The all-facts path reduced time by 84.14%, bytes
by 1.88%, and allocations by 49.41%. Parse-plus-extraction stayed unchanged. -
Opt-in Lean 4 support now provides a grammar blob, scanner, highlights,
outline tags, and focused corpus tests ingrammars/lean. The default
206-language registry remains unchanged. -
The V10 fleet harness now runs bounded Google Cloud spot workers across all
registered languages. It enforces wall, memory, disk, and automatic deletion
limits. Accepted epoch20260808T202958Z-v10-full-5003ffbacompleted all
1,435 measurements for 206 languages. -
scripts/run_randomized_benchmarks.shnow runs one process for each shuffle
seed and randomizes benchmark order. The standard comparison uses 20 seeds,
GOMAXPROCS=1, a 750 millisecond benchmark time, and memory reporting. -
Merge-event census instrumentation now records merge decisions and refusal
gates on the production and C-oracle paths. Its first constructed receipt
reports 15 Go merges against 191 C merges across 104 sources. It records no
source where Go merges more often than C. -
A regression guard covers issue #660.
It checks anonymous comma nodes in Python imports and subscripts on both
parser routes. -
The included-ranges route now has committed test coverage.
Parser.SetIncludedRangeshad no test for any language, and injection
uses that call for every injected child.
A Markdown document with two or more Go fences reaches it in production.
Four new gates cover the route: a root-symbol gate, a positive control
that proves the Go arm still rewrites the tree there, a C-oracle table
that pins the measured root of both parsers across four range
geometries, and a deletion guard.
The route is not at parity with C, and the new tests do not claim it is.
The root span matches C only when the first range starts at byte 0 and
the last ends at end of file, which is a shape an injection child never
receives.
One pinned geometry publishes anERRORroot where C publishes
source_file.
That case is an open defect in included-range clipping, recorded so the
fix moves the pin. -
The Go result-compatibility arm's three members are now registered as
named census subpasses, so a census receipt names the member that
rewrote the tree instead of only the arm.
Census receipts are recorded on the production route only; the default
candidate route leavesParseRuntime().NormalizationPassesnil.
Fixed
-
C enum lists with three or more enumerators no longer publish
ERRORor
MISSINGnodes. A clean forest result can replace a recovered tree only
after it covers the full source and contains no recovery nodes. This fixes
issue #667. -
Node.HasErrorOrMissingreports both recovery node forms. The
grammargen parse -strictcommand now rejects either form. -
JavaScript, TypeScript, and TSX scanners now bind external results through
each language's positional symbol table. Regenerated blobs no longer mistype
shifted external symbols. -
The full-parse retry selector no longer releases an incumbent when a
candidate aliases it. This restores selected roots across retry and compact
fallback paths. -
The accepted-error retry ladder now honors explicit stack and merge caps.
It also keeps bounded pass counts and configured wall budgets across retries. -
Kotlin published an
ERRORroot instead ofsource_filewhen a parse used
Parser.SetIncludedRangeswith more than one range and the parse entered
recovery.
The recovered-root normalization that owns this result was removed on
2026-08-02 as dead code.
The census behind that removal measured the fresh, over-64-KiB, incremental
and pinned routes.
It never measured the included-ranges route, and the member is live there.
The member is restored.
On the committed witness,testdata/included_ranges/kotlin_work_queue_test.kt
with three ranges, the root returns tosource_file, which is the kind the
locked Kotlin C reference runtime publishes for the same input.
Production reaches this route through injection.
The member retags the root.
Its downstream consequence is not always toward the reference runtime: on one
measured file the retag lets a later stage flatten a clean
class_declarationinto root-level members.
The route now has committed test coverage for Kotlin, in the root package and
in the C-parity lane. -
The C-recovery missing-token search (
cHandleError/
cDoAllPotentialReductions,parser_recover_c.go) cloned the whole GSS
stack without a work limit.
A 4-byte erlang input and a 56-byte jsdoc input drove heap use past 2 GB
in seconds.
Two new loop ceilings now bound the search directly.
cRecoverMaxReductionCandidateAttemptscaps candidate attempts within one
cDoAllPotentialReductionscall.
cRecoverMaxMissingTokenTrialscaps total trials across one
cHandleErrorsearch.
cRecoverMaxReductionCandidateAttemptsis the active mechanism.
It is what stops every known witness and every corpus file measured so
far.
cRecoverMaxMissingTokenTrialshas not fired once on any of them.
It stays in place as an unexercised backstop for input this codebase has
not sampled yet, not as a mechanism this fix currently relies on.
A newParser.budgetScratchpointer feeds GSS-scratch allocation into
the existing 512 MB soft memory budget.
The check already in the main parse loop now covers the C-recovery
candidate search too.Measured through the shipped regression test on the
erlang_pfx_017_71b
witness (testdata/recovery_memory_bound_witnesses/): heap growth after
this fix is 144.1 MB on the production route and 138.9 MB on the compact
route, in well under half a second.
Before this fix, the same input reached 695 MB and 52.0 seconds.
Both post-fix numbers are the real, reproducible figures.
An earlier draft of this fix reported smaller ones.parser_memory_budget_runtime.go'sruntimeMemoryHardCeilingEnabled
function keeps its exact prior behavior.
This fix adds a comment there recording why an earlier draft's
source-length-independent hard ceiling was tried and dropped.
It reopened issue #454's determinism symptom class on sub-64 KiB input.
It also cost 3.6-9x more time on ordinary small parses, with no
offsetting protection over the two loop ceilings above. -
The GSS-forest link cap could silently drop the widest hidden-symbol
alternative when it tied a narrower one on score and error cost.
The cap kept the earlier arrival by default in every tie.
json5's flat-array grammar shape hits this tie constantly on ordinary
input.
Other forest-default languages hit it rarely or not at all on their
current tables.
The cap now keeps the wider alternative when two links tie on the same
symbol and the same end byte.
A narrower-tie or cross-symbol tie keeps the prior behavior unchanged.
Parser.ForestCapTieStats()is a new method.
It reports how often the tie fires and how often the fix changes the
outcome.
SetGOT_FOREST_CAP_TIE_DUMP=1to also record a bounded per-decision
receipt list. -
Corrected the root-cause comment on the javascript declared-conflict
election witness test.
grammargen/lr.goalready retains thelabeled_statement/_property_name
GLR fork that the real tree-sitter-javascript grammar declares.
That fix landed 2026-03-16.
The shippedjavascript.binblob predates the fix by two weeks.
Nothing resynced the blob afterward, so the raw parse still diverges from
the C oracle today.
A new unit test pins the retention rule directly against the generator.
Regressions now surface without a blob rebuild. -
The Swift optional-binding vs trailing-closure fix (#542) added a
shift/reduce precedence branch togrammargen/lr.go.
It ran before the declared-conflict retention check.
It also matched ordinary undeclared conflicts in other grammars and
picked the wrong side.
JavaScript'supdate_expressionvsbinary_expressionconflict is one
example.
The branch now runs last, after declared-conflict retention and the
ordinary precedence ladder get a chance to resolve the conflict first.
The Swift case from #542 still resolves correctly.
A new test now guards javascript and typescript regeneration against the
C reference parser on real source files.
Changed
-
Parser stop checks now skip inactive callbacks and keep the common callback
direct. Result materialization reads the wall clock every 64 checkpoints.
Cancellation and sticky stop checks still run at every checkpoint. -
GLR recovery now computes C-compatible error cost and visible counts in one
tree walk. Memo indexing uses pointer-bit folds and checks the primary way
first. Graph-structured stack (GSS) nodes store clean-zero merge results
without a larger node layout. Extra-link mutations invalidate the result. -
C-recovery promotes an error stack to the graph-structured stack before
reduction forks. Deep recovery branches now share their immutable prefix.
The Swift recovery witness reduced time by 9.96%, bytes by 59.65%, and mean
peak resident memory by 22.09%. The 20-seed combined suite reduced KDL
recovery time by 1.20%, bytes by 13.14%, and allocations by 1.58%.
Other parser timings stayed neutral. -
The randomized benchmark suite now accepts an exact recovery corpus file and
language. The 20-seed comparison against the release boundary reduced the
timing geomean by 1.77%. Elixir recovery improved by 15.21%, KDL recovery by
9.12%, full parse by 1.16%, and incremental no-edit by 6.51%.
FactProgramparse and extraction improved by 1.23%.
The parser-core control stayed neutral. No timing, byte, or allocation metric
had a significant regression. -
The guarded parser-core bytecode experiment now supports
REDUCE_CHAINand
REDUCE_SHIFT. The corridor remains off by default. Each superinstruction
also requires its own experiment gate. -
Synthetic-root replay now hashes frames and gap cursors, memoizes gap tokens
and advance transitions, reuses scratch, and pools external token sources.
Paged advance and close streams bound retained memo storage.
Advance memoization cut the hard Elixir target latency by 22.62 percent.
Close paging cut bytes per operation by 5.15 percent and allocations by
60.48 percent. Its latency remained neutral across 20 balanced pairs. All
16 combined-suite latency rows also remained neutral. -
Exact C and V runtime profiles now avoid certified duplicate retry work.
Other grammars keep the conservative retry ladder. -
Performance counters now expose maximum resident memory, replay closure
distribution, memo capacity skips, and parser stop attribution. -
The compact fresh-path route now skips two tail steps for a language with
no live result-compatibility entry: the C-recovery-swallow resolver and
the final-tree compaction pass.
The eligible set is computed from
testdata/result_compat_ownership_v1.json.
It is not a maintained list.
A future dispatcher arm cannot silently escape it.
163 of 206 registered languages are eligible today.
Go is not one of them.
dispatch.gostays live, sogrammargen_lrand the other three canonical
Go fixtures still take the full tail.
A deep-tree digest comparison (elided against unelided) is exact across
every eligible language's smoke sample.
It is also exact across every real-corpus file this campaign measured, up
to 484 KB.
This is a correctness-neutral simplification, not a measured performance
win: the result-compatibility dispatch and its error-summary walk already
run once, during materialization, for every language; the tail's own copy
of that work was already unreachable in the common case before this
change, eligible language or not.
Two eligible-language timing probes (OCaml, Zig) and the Go warm-route
benchmark all read within this shared host's noise floor, consistent with
that finding.
See the PR for the full reading anddocs/compat-tail-elision.mdfor the
corrected performance and correctness analysis. -
The condense-candidate dispatch path no longer passes a closure through
two wrapper layers per event.
Each shift, cohort, and reduction entry point now validates the scheduler
owner and calls the uncheckpointed operation directly.
Behavior is unchanged; every identity, work-count, and allocation check
still passes.
Local timing on a shared host showed no significant change across four
fixtures.
The host carried heavy background load throughout the run, so the result
is not a sealed measurement. -
Removed the 64 KiB source-length eligibility decline from the compact
admission switch.
A fresh full parse of any size now attempts the compact route first.
The scheduler's stop-control poll bounds a large or pathological input:
it compares the compact core's own real retained-memory footprint
against the same soft memory budget production honors, and falls back to
production with a matchingParseStopMemoryBudgetstop reason when the
budget is exceeded.
Every decline path now releases the compact core's retained capacity
before returning, not only its logical record count, so a production
fallback does not run alongside megabytes of memory an earlier declined
attempt on the same parser left allocated.
An operator watchingAdmissionCandidateCounters()sees this directly:
a large input that declines bumps the fallback count exactly as a small
one always did.
This bounds retained footprint, not the compact scheduler's own transient
per-token allocation during a declined attempt; closing that remaining
gap trades against how large an input the compact route can still serve,
and is an open follow-up, not resolved by this change.
Routing only changes; every canonical tree digest stays identical.