Skip to content

gotreesitter v0.54.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 22:05
· 99 commits to main since this release
90a9c27

Added

  • Add FactProgram.ExtractInto to reuse caller-owned fact storage across trees.
  • Add a per-language admission allowlist (empty by default) so a later change
    can graduate one language's compact route without flipping the global
    default.
  • Add the admission_route_performance_sanity CI gate, comparing the compact
    route against the production route in-process, interleaved, on a fixed
    9-language corpus.

Changed

  • Default the compact ("candidate") admission route to off. Buildbox's
    tamarack harness measured it 1.1x to 2.2x slower than production on typical
    files across Python, Rust, Markdown, Lua, CSS, Bash, and Go, and
    TypeScript/YAML paid for both routes on every parse (compact declines, then
    production reparses). Set GTS_ADMISSION_CANDIDATE=1 (or true/on/yes)
    to opt back in.
  • Throttle the compact scheduler's memory-footprint poll: recompute the exact
    footprint on the first poll of a parse attempt, every 64th poll after it, or
    whenever a cheap, O(1) capacity-based growth proxy shows growth covering
    1/16th of the configured budget since the last exact check. This bounds
    worst-case overshoot by construction, not just by dispatch count. Measured
    peak-footprint overshoot: 1.25-2.16% over budget on an adversarial witness.
  • Make ExternalLexer's read-frontier tracking lazy: skip the frontier
    recompute once the recorded values already provably cover the current
    position. This benefits every route that uses an external scanner (YAML,
    Python, Markdown, Bash, and others), not just the compact route.

Fixed

  • Preserve explicit end-token acceptance when importing and minimizing C lexer
    states. Regenerate CSV from its locked C source. Older blobs retain their
    end-of-input fallback.
  • grammargen gives a named token.immediate() terminal with a string body the
    same specificity bonus that an inline one gets. The lexer again accepts the
    immediate terminal instead of a plain string literal with the same text.
    Swift var opt: Int? parsed with an ERROR after a table rebuild, because the
    anonymous ? won the same-span tie over _immediate_quest. Pattern-bodied
    named immediate terminals keep their authored precedence.
  • Fix parser_result_yaml.go silently clearing HasError() for a lone
    unmatched YAML flow-collection opener (for example a bare [ at end of
    input), which diverged from the C reference's (ERROR ...) shape. The
    production route now declines recovery for a bare [/{ and promotes an
    all-ERROR reduction to the ERROR root itself. This bug predates this
    release; the compact-route default flip is what exposed it.
  • The scanner fuzz allocation check now measures the minimum across 5 samples
    instead of requiring every re-run to exceed budget, removing a transient
    GC-timing false failure. The byte-budget assertion is skipped under the
    race detector, where instrumentation noise swung a measured TotalAlloc delta
    for the same input from about 66KB to about 29MB (more than 400x) and could
    no longer separate noise from a real finding. The elapsed-time check and the
    parser memory budget (WithParserPoolMemoryBudgetBytes) remain the hard
    bounds under -race.
  • Invalidate the GLR shape-prefix cache only when a stack merge rewrites a
    link 0, not on every successful merge. A fixed nesting depth that forks and
    merges on every token used to rewalk the whole spine on the next head hash,
    which turned the parse superlinear past depth 1600 (issue #454). An
    extra-link-only merge now keeps its exact cached prefixes.
  • Resume C and Java TokenSource scanning from the still-valid queued
    literal-tail tokens, instead of resetting and relexing, when incremental
    reuse resumes inside an already-queued string or character literal. The
    reset path used to drop the queue and misread the remaining bytes, so the
    incremental tree could differ from a fresh parse.
  • Track reused bytes for plain ParseIncremental, not only for
    ParseIncrementalProfiled, so the reuse-budget stop arms on both entry
    points. ParseIncremental allocates no timing record, so a reuse-hostile
    edit taken through that entry point never armed the stop before this fix.
  • Exempt extra (comment and whitespace) leaves from the compact incremental
    reuse proof. One unproven extra leaf used to disable incremental reuse for
    the whole compact tree; an ordinary leaf still requires a proof.
  • Fall back to a fresh full parse when ParseIncremental receives an old
    tree with no recorded Tree.Edit call and a new source of a different
    length. The incremental path used to trust stale byte positions against
    the new source in that case, with no error or ParseStoppedEarly signal.
  • Route every Groovy incremental parse through a fresh full parse. The
    Groovy grammar table derives a function-call juxtaposition only for a
    block's first statement, so spliced incremental reuse could produce a tree
    that a fresh parse of the same bytes never produces. Groovy's incremental
    performance is unchanged; parity with a fresh parse is now guaranteed
    instead of usually holding.

Known gaps

  • COBOL's column-dependency edit-invalidation over-invalidates a later-line
    token after an earlier-line, non-crossing edit. This is fail-safe (extra
    reparse work, not a wrong tree); COBOL does not support incremental reuse
    today, so it has no current observable effect. Root cause is out of scope
    for this release.

Breaking Changes

  • GTS_ADMISSION_CANDIDATE now defaults off; its meaning is inverted from the
    prior release. Set it to 1, true, on, or yes to keep the compact
    route on.

Performance evidence

Buildbox's tamarack harness measured the default route's in-process Go/C
ratio across three independent interleaved runs on a fixed 9-language
typical-file corpus, comparing this release's candidate code against the
prior default:

lang before after speedup
go 5.97x 5.42x 1.10x
python 2.87x 1.54x 1.86x
typescript 3.17x (100% fallback) 2.31x 1.37x
rust 5.06x 2.53x 2.00x
yaml 3.56x (100% fallback) 2.34x 1.52x
bash 3.65x 2.08x 1.75x
markdown 6.36x 2.87x 2.22x
lua 4.66x 2.33x 2.00x
css 4.60x 2.36x 1.95x

The lever-2 overshoot-bound refinement (dominant-capacity growth trigger) cost
nothing measurable; the ratio table is unchanged within run-to-run noise after
it landed. Source: #1264. A downstream #454 report attributes its resolved
regressions to the production route becoming the default.