You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A freshly loaded index reported absolute display paths (and the UI could
not round-trip them) until the next finalize, because the configured roots
were registered after the per-file metadata pass cached the paths.
Persisted index corruption after a file removal: save_index compacted the
file table over tombstoned ids but wrote trigram bitmaps, symbols and dependency
edges keyed by live id, so every file after a removed one was misattributed or
lost on reload. Ids are now remapped at save time.
Every match on line 1 of every file received the 3× symbol-definition boost
(the synthetic filename symbol at line 0 counted as a definition line).
Files with a line over 100 KB or extreme nesting were dropped from the whole
index; they are now text-searchable and only skip symbol extraction.
The persisted mtime/size described the on-disk state at save time, so a
file edited during a long build was never detected as stale. Both are now
captured when the content is read.
Symbols: .tsx is parsed with the TSX grammar; JS/TS class methods, const f = () => …, generator functions, abstract classes and namespaces are
extracted; symbol line/column now come from the name node (Java @Override,
decorators, attributes and multi-line C signatures reported the wrong line).
Directory rename/delete events were silent no-ops in the watcher; every file
under the old path stayed indexed. The watcher now also applies the same
eligibility rules as the initial build (include extensions, excludes, size).
A panic on the watcher path poisoned the engine lock and turned every search
into a 500 until restart; the watcher path is now panic-guarded, runs on an
8 MB stack, and the REST/gRPC readers recover from a poisoned lock.
/api/context no longer overflows on a huge context value and is capped at
200 lines each side; user regexes have an explicit compile-size limit.
DependencyIndex::remove_file never pruned the filename index and every
watcher update pushed a duplicate path entry.
Added
A real-corpus benchmark (examples/corpus_bench.rs, run in CI over pinned
tokio and Django checkouts, nightly over the Rust compiler tree) reporting
build throughput, resident memory, query latency percentiles, incremental
update cost and index save/load.
/api/search paging and budgets: offset (deterministic ordering, so pages
are stable), timeout_ms, and total_matches / truncated_by_budget in the
response; has_more is now derived from the real total. Regex and symbol
searches report ranking info too.
Every result carries line_match_start / line_match_end (byte offsets into
the full line) and match_column (0-based character column) on REST and gRPC.
Added
Symbol references: /api/search?references=true (gRPC SearchRequest.references) returns the call sites, type mentions and
implemented traits of an identifier as SYMBOL_REFERENCE results, from the
grammars' tags queries (the Rust query is supplemented with path and generic
calls). References are stored compactly (12 bytes each, names interned) and
persisted.
Multi-line regex: a pattern that mentions a newline (\n, \r, \x0a) or
sets the s flag ((?s)begin.*?end) is matched across lines and reported on
the line where each match starts. Other patterns stay line-oriented.
Query syntax for plain-text searches: "quoted phrases", several AND-ed
terms, -term, file: / -file:, lang: / -lang:, case:yes, word:yes (REST case / word parameters and gRPC case_sensitive / whole_word override the in-query switches).
/api/ready (readiness, distinct from /api/health liveness), /metrics
(Prometheus text: search counters by outcome, latency histogram, index gauges)
and the standard grpc.health.v1 service.
Build: the semantic engine is behind the semantic Cargo feature (off by
default; ml-models implies it), release binaries use thin LTO and are
stripped, OpenTelemetry is on 0.32 (one tonic/axum stack), md5 and glob are
gone (the config fingerprint format changed, so the first start after upgrading
rebuilds the index once), and cargo-deny + Dependabot are wired into CI.
Trigram extraction folds case per byte and dedupes through a bitset instead of
lowercasing a copy of every file and hashing every byte.
Import resolution no longer probes candidate paths through realpath
(17.6 million failing readlink calls while building a 6k-file corpus);
candidates are resolved lexically against the indexed path table. Full
build of tokio + Django: 25 s -> 3 s.
Updating one file from the watcher went from 55 ms to 4 ms on that corpus:
posting-list removal is parallel with a membership fast path, and compiled .gitignore matchers are cached per directory (keyed by mtime) instead of
rebuilt per event.
Memory diet: each indexed file stores its path once (shared with the
path lookup map) and keeps at most 64 bytes of inline state; only files
above the mmap threshold allocate mapping state. The evictable fallback
cache for unmapped large files is gone (they are read per access), and
the all-documents bitmap is borrowed rather than cloned per short or
unconstrained query.
Symbols are extracted with each grammar's own tags.scm query (name-node
positions, upstream-maintained coverage), supplemented by the previous walker;
new symbol kinds Module, Macro, Field and Property; C++ .h headers are
detected; JSON/TOML/YAML/HTML/CSS/Markdown are no longer parsed; parses are
capped at 2 s.
Unknown config keys are rejected and the configuration is validated at
startup (addresses, limits, index path directory); OTEL_SDK_DISABLED=true can
no longer be overridden.
REST requests time out (default 30 s), bodies are capped, and searches beyond max_concurrent_searches get 503 + Retry-After instead of queueing; gRPC has
a per-request timeout and per-connection concurrency limit.
gRPC max_results = 0 now means the default page of 50 (was clamped to 1); the Index RPC applies the indexer's eligibility rules and no longer holds the
engine lock for its whole walk.
/api/diagnostics reports the real indexer configuration and caches its
extension breakdown per index generation; malformed query parameters return the
JSON error envelope; display paths resolve through their root instead of a
per-file suffix scan.
Every query's work is bounded by a match budget (8× the page, minimum 512) and
an optional deadline; rank=full no longer materializes every match of every
candidate before truncating.
Regex searches are accelerated for (?i) literals and repeated suffixes
(abc+), and compiled patterns are cached; plain-text verification scans the
whole buffer with memchr and builds symbol maps only on the first hit.
Symbol search consults the symbol cache before reading any file, never drops
low-scoring files before matching, ranks exact > prefix > substring, and
returns one row per line.
Retrieval no longer reads through a live memory map for files up to 1 MiB:
a searcher touching a mapped file that an editor truncated in place was an
uncatchable SIGBUS that killed the server. Small files (essentially all source
files) are read into an owned buffer per access; only larger files are mapped.
This also keeps the mapping count far below vm.max_map_count.
.gitignore / .ignore files under the indexed paths are honoured by the
initial build and by the watcher (indexer.respect_gitignore, default true).
Persisted index format v7 (FCSIDX05): a sectioned layout (header with
version, CRC-32 and section lengths; metadata; fixed-width sorted trigram
directory; bitmap region). Saving streams the posting lists without an
intermediate copy and loading memory-maps the file, validates it before
decoding, and deserializes bitmaps in parallel from the mapping. The format
also stores unresolved imports (so a checkpoint restore still gains the edge
when the target file is indexed later), nanosecond mtimes, byte paths,
run-optimised bitmaps and symbol references. Index files written by earlier
versions are rebuilt automatically; a golden fixture pins the format.
Import resolution is now per language (Rust crate/module paths, Python
relative and package imports, JS/TS extension and index probing, @/
aliases) and no longer guesses a same-named file anywhere in the repo for bare
package names; unresolved imports are retried only when a file with a matching
name appears instead of after every batch.
Watcher events are gathered for 200 ms and applied under one write lock with a
single posting-list pass; ranking metadata and the all-documents cache stay
current across incremental updates and after a persisted load.
Graceful shutdown: SIGINT/SIGTERM now stop both servers, stop the indexer
(persisting a checkpoint), save pending watcher updates, and flush telemetry.
Security defaults: both listeners bind to 127.0.0.1; CORS is off unless server.cors_origins is set; the gRPC Index RPC only accepts paths under the
configured indexer.paths and no longer follows symlinks; a web-UI bind
failure is fatal instead of silently leaving the server without a REST API.
Minimum supported Rust is 1.89 (was documented as 1.70); the toolchain is
pinned via rust-toolchain.toml; reqwest uses rustls so builds and tests no
longer need system OpenSSL.
CI lints all targets and features, checks docs, has an MSRV job, and releases
are gated on tests; the VS Code extension publishes on ext-v* tags only.
Failing test names are surfaced as workflow annotations.
Releases no longer ship the -debug archives (an unoptimised binary per
platform); release binaries are built with thin LTO and stripped.
Repository
.gitattributes normalises line endings; the corrupted .gitignore is
repaired; onnxruntime/ and the test_corpus gitlinks are no longer tracked.