Releases: adikeshri/tachyon
Release list
Bumped runtime base image alpine:3.21 → 3.24
Bumped runtime base image alpine:3.21 → 3.24
What's Changed
- Add Dependabot config for cargo, github-actions, and docker by @adikeshri in #11
- Add missing labeler config for label.yml workflow by @adikeshri in #19
- Bump actions/labeler from 4 to 7 by @dependabot[bot] in #18
- Bump docker/setup-buildx-action from 3 to 4 by @dependabot[bot] in #16
- Bump alpine from 3.21 to 3.24 by @dependabot[bot] in #12
- Bump docker/metadata-action from 5 to 6 by @dependabot[bot] in #13
- Fix docker run command in README by @adikeshri in #20
- Add community health files: CoC, security policy, changelog, templates by @adikeshri in #21
New Contributors
- @dependabot[bot] made their first contribution in #18
Full Changelog: v4.2.0...v4.2.1
Off-lock segment flush
Flushes no longer hold the write lock for the encode — only for a brief seal and commit, mirroring merges. Worst-case concurrent search latency during ingest: 233 ms → 106 ms, and much less variable run-to-run
What's Changed
- Feature/disk segment writer by @adikeshri in #10
Full Changelog: v4.1.0...v4.2.0
Run segment merges off the collection write lock
Merging no longer blocks the collection's write lock for its full duration - only a brief snapshot and commit. Worst-case concurrent search stall during indexing: 1.92s → 233ms (8x).
- Merge splits into snapshot → build → swap; only snapshot/swap hold the lock.
- merge_gate caps one merge in flight at a time.
- Fixed a resurrection bug found pre-release: tombstone pruning now uses each segment's presence bitmap instead of its declared range, since off-lock timing can make segment ranges overlap.
- Merge cost itself is unchanged — only when the lock is held moved.
- Known gap: a large flush still holds the lock; off-lock flushing is next.
Stream segment writes and merges to bounded memory (format v4)
Segment merging no longer holds everything it's folding together decoded in memory. Peak RSS at 5M documents: 5.2 GiB → 1.5 GiB. Indexing throughput and search latency unchanged.
Format v4 (breaking, reindex required): .post/.col/.doc move their directory to a footer at EOF so writes can stream forward in bounded memory. .terms/.ids stream via fst::MapBuilder. Widened offsets past the old 4 GiB limit.
Streaming merge: no longer rebuilds a MemTable by re-parsing/re-tokenizing — streams a k-way merge of already-encoded segments directly.
Memory: 1M docs 999→646 MiB peak; 5M docs 5.2→1.5 GiB peak. Throughput and latency flat.
Bug fix: a block-skip regression in postings decode, caught by new regression tests before release.
8 new merge correctness tests; tachyon-bench gains --merge-scenario for isolated merge timing/RSS.
Query engine: 2–2.7× faster broad-query search, autocomplete now clears its latency target at every scale
This release eliminates the dominant sources of per-document overhead in the search path — allocation churn, dynamic-dispatch fan-out, and repeated segment-column decoding — cutting broad-query search latency by 2–2.7× and autocomplete latency by up to 31× at scale, with zero behavioral or API changes.
What changed:
- WAND scoring reuses a scratch buffer across documents instead of allocating fresh per-document structures, and now defers decoding match positions until after a document is confirmed worth fully scoring, rather than before.
- The query executor caches each posting frontier's current document and each source lookup's owning segment, removing redundant scans that previously ran on every document visited.
- On-disk segment columns (used by filters and facets) are now decoded once per segment and cached, instead of being fully re-decoded on every query — this was the specific bug making filtered searches slower than unfiltered ones despite touching fewer documents.
- Autocomplete's per-term suggestion counting now reads a stored count directly when nothing has been deleted, instead of decoding and filtering a term's entire posting list.
- The benchmark harness gained --vocab-scale, --max-memtable-docs, and --concurrency flags, the last producing a sustained-throughput number directly comparable to competitor benchmarks (315 concurrent queries/sec measured, vs. a 104 reference point).
Measured impact (same corpus, same machine, before → after):
- Search p95 @ 1M docs: 51.3ms → 31.6ms; @ 5M docs: 259.5ms → 160.3ms
- Search + filter mean @ 5M docs: 232.7ms → 116.5ms
- Autocomplete p95 @ 5M docs: 18.3ms → 1.4ms — now passes its 5ms target at every scale tested
- Concurrent throughput: previously unmeasured; now 315 queries/sec sustained
block-max WAND query pruning- #5
Merge pull request #5 from adikeshri/feature/disk-segment-writer Add true block-max WAND query pruning
Disk Segment Write, Lazy segment reads, Tiered merges
Merge pull request #4 from adikeshri/feature/disk-segment-writer Implement tiered merges
Fix vulnerabilities and optimize search
Drop wget from runtime image, add native healthcheck wget was only used for the Docker HEALTHCHECK, but pulled in CVE-2026-58469/58470/58471/58472, CVE-2025-69194 (GNU Wget), and CVE-2025-60876 (BusyBox wget) — none patched upstream/in Alpine yet. Replace it with a `tachyon --healthcheck` subcommand that does the /health check itself over a raw TCP connection, so the image no longer needs wget at all.