v0.17.0
Java corpus parity and parser-performance release.
Added
- Java corpus Docker harnesses for seeded Apache Lucene stress testing,
including largest/random corpus selection, timeout sweeps, cgo comparison
benchmarks, UAX generated-file stress runs, materialization profiles, runtime
diagnostics, and ambiguity profiling. Parser.ParseNoTreeBenchmarkOnlyfor diagnostic parser-loop benchmarks that
suppress full public tree materialization while keeping lexing and parse
actions active.- Language-family full-parse benchmark matrix controls, warm parser reuse
benchmarks, and parser scratch/reset regression coverage. - Top-50
grammargenparity coverage checks and focused fixtures for Java,
Bash, Python, Swift, comment, CPON, git config, gomod, ini, and related
imported-grammar edge cases.
Changed
- Java parsing now handles contextual keyword/token selection, compact generic
close-angle splitting, switch rule labels versus lambdas, shift expressions
before calls, array initializer commas, repetition shifts, and downstream
recovery cases much closer to the C runtime. - Parser hot paths cache language traits on DFA token sources, preserve scratch
buffers across pooled token-source resets, clear GLR/GSS scratch by written
range and epoch, and reduce parser clearing/lookup overhead. - Initial Java full parses defer parent-link wiring until the tree API needs it,
avoiding public tree bookkeeping during the parse-time materialization hot
path. - Edited trees now reuse the old primary arena directly where possible, and
borrowed arenas are deduplicated to reduce incremental parse retention churn. - HTML-family scanner deserialization reuses tag snapshots and shared ASCII
lookup construction while preserving first-match behavior.
Fixed
- Bash generated-parser parity issues around command names, statement
boundaries, broad DFA relexing, and arithmetic expansion token normalization. - Comment tag parsing, parser compatibility normalization, parser-valid
zero-width token preference, broad relex candidate matching, and string
whitespace recovery behavior. grammargennormalization and conflict-resolution gaps for lexical choices,
aliased inline precedence, long Unicode escapes, augmented start symbols,
terminal collisions, Python/Swift parity regressions, Julia assignment
conflicts, D binary repeat, PowerShell binary repeat, and gomod grouped
retract intervals.- Parser reset paths now avoid stale node-equivalence and GSS cache hits after
reuse.
Performance
- Main-branch Go/editor benchmark median on the standard generated Go workload:
full DFA parse~1.98 ms, incremental single-byte edit~666 ns, no-edit
incremental reparse~2.84 ns, with full parse at5 allocs/op. - Java Lucene largest top-10 Docker benchmark: Go full DFA
~537 ms, Go
no-tree diagnostic~402 ms, cgo full~394 ms; full/cgo is about1.36x. - Java generated UAX file Docker benchmark: Go full DFA
~306 ms, Go no-tree
diagnostic~235 ms, cgo full~213 ms; full/cgo is about1.44x.
Testing
- CI for the release commit includes green build, freshness, cgo parity smoke,
and perf-regression gates on PR #80. - Java real-corpus parity and large-file timeout diagnostics are now
reproducible through bounded Docker lanes rather than ad-hoc local runs.