Skip to content

Releases: adikeshri/tachyon

Bumped runtime base image alpine:3.21 → 3.24

Choose a tag to compare

@adikeshri adikeshri released this 23 Aug 14:02
8da9c5d

Bumped runtime base image alpine:3.21 → 3.24

What's Changed

  • Add Dependabot config for cargo, github-actions, and docker by @adikeshri in #11
  • Add missing labeler config for label.yml workflow by @adikeshri in #19
  • Bump actions/labeler from 4 to 7 by @dependabot[bot] in #18
  • Bump docker/setup-buildx-action from 3 to 4 by @dependabot[bot] in #16
  • Bump alpine from 3.21 to 3.24 by @dependabot[bot] in #12
  • Bump docker/metadata-action from 5 to 6 by @dependabot[bot] in #13
  • Fix docker run command in README by @adikeshri in #20
  • Add community health files: CoC, security policy, changelog, templates by @adikeshri in #21

New Contributors

Full Changelog: v4.2.0...v4.2.1

Off-lock segment flush

Choose a tag to compare

@adikeshri adikeshri released this 21 Aug 23:48
899130f

Flushes no longer hold the write lock for the encode — only for a brief seal and commit, mirroring merges. Worst-case concurrent search latency during ingest: 233 ms → 106 ms, and much less variable run-to-run

What's Changed

Full Changelog: v4.1.0...v4.2.0

Run segment merges off the collection write lock

Choose a tag to compare

@adikeshri adikeshri released this 20 Aug 10:48
890d0c7

Merging no longer blocks the collection's write lock for its full duration - only a brief snapshot and commit. Worst-case concurrent search stall during indexing: 1.92s → 233ms (8x).

  • Merge splits into snapshot → build → swap; only snapshot/swap hold the lock.
  • merge_gate caps one merge in flight at a time.
  • Fixed a resurrection bug found pre-release: tombstone pruning now uses each segment's presence bitmap instead of its declared range, since off-lock timing can make segment ranges overlap.
  • Merge cost itself is unchanged — only when the lock is held moved.
  • Known gap: a large flush still holds the lock; off-lock flushing is next.

Stream segment writes and merges to bounded memory (format v4)

Choose a tag to compare

@adikeshri adikeshri released this 19 Aug 20:46
9bd5b32

Segment merging no longer holds everything it's folding together decoded in memory. Peak RSS at 5M documents: 5.2 GiB → 1.5 GiB. Indexing throughput and search latency unchanged.

Format v4 (breaking, reindex required): .post/.col/.doc move their directory to a footer at EOF so writes can stream forward in bounded memory. .terms/.ids stream via fst::MapBuilder. Widened offsets past the old 4 GiB limit.
Streaming merge: no longer rebuilds a MemTable by re-parsing/re-tokenizing — streams a k-way merge of already-encoded segments directly.
Memory: 1M docs 999→646 MiB peak; 5M docs 5.2→1.5 GiB peak. Throughput and latency flat.
Bug fix: a block-skip regression in postings decode, caught by new regression tests before release.
8 new merge correctness tests; tachyon-bench gains --merge-scenario for isolated merge timing/RSS.

Query engine: 2–2.7× faster broad-query search, autocomplete now clears its latency target at every scale

Choose a tag to compare

@adikeshri adikeshri released this 18 Aug 13:58
dbba9d6

This release eliminates the dominant sources of per-document overhead in the search path — allocation churn, dynamic-dispatch fan-out, and repeated segment-column decoding — cutting broad-query search latency by 2–2.7× and autocomplete latency by up to 31× at scale, with zero behavioral or API changes.

What changed:

  • WAND scoring reuses a scratch buffer across documents instead of allocating fresh per-document structures, and now defers decoding match positions until after a document is confirmed worth fully scoring, rather than before.
  • The query executor caches each posting frontier's current document and each source lookup's owning segment, removing redundant scans that previously ran on every document visited.
  • On-disk segment columns (used by filters and facets) are now decoded once per segment and cached, instead of being fully re-decoded on every query — this was the specific bug making filtered searches slower than unfiltered ones despite touching fewer documents.
  • Autocomplete's per-term suggestion counting now reads a stored count directly when nothing has been deleted, instead of decoding and filtering a term's entire posting list.
  • The benchmark harness gained --vocab-scale, --max-memtable-docs, and --concurrency flags, the last producing a sustained-throughput number directly comparable to competitor benchmarks (315 concurrent queries/sec measured, vs. a 104 reference point).

Measured impact (same corpus, same machine, before → after):

  • Search p95 @ 1M docs: 51.3ms → 31.6ms; @ 5M docs: 259.5ms → 160.3ms
  • Search + filter mean @ 5M docs: 232.7ms → 116.5ms
  • Autocomplete p95 @ 5M docs: 18.3ms → 1.4ms — now passes its 5ms target at every scale tested
  • Concurrent throughput: previously unmeasured; now 315 queries/sec sustained

block-max WAND query pruning- #5

Choose a tag to compare

@adikeshri adikeshri released this 17 Aug 17:12
c8a077e
Merge pull request #5 from adikeshri/feature/disk-segment-writer

Add true block-max WAND query pruning

Disk Segment Write, Lazy segment reads, Tiered merges

Choose a tag to compare

@adikeshri adikeshri released this 16 Aug 00:52
792a2e3
Merge pull request #4 from adikeshri/feature/disk-segment-writer

Implement tiered merges

Fix vulnerabilities and optimize search

Choose a tag to compare

@adikeshri adikeshri released this 14 Aug 14:00
Drop wget from runtime image, add native healthcheck

wget was only used for the Docker HEALTHCHECK, but pulled in CVE-2026-58469/58470/58471/58472, CVE-2025-69194 (GNU Wget), and CVE-2025-60876 (BusyBox wget) — none patched upstream/in Alpine yet. Replace it with a `tachyon --healthcheck` subcommand that does the /health check itself over a raw TCP connection, so the image no longer needs wget at all.