Repository navigation
v0.8.2 — faster mining on compressed tables, corpus-wide stats
Two rungs in one release. 0.8.1 was merged but never tagged, so it ships here too; the changelog records each under its own version.
Mining a compressed table is no longer several times slower (0.8.1)
Measured on a real 462 MiB KRvkbn (70.8 MiB compressed), taskset -c 0-3, warm:
| workload | before | after |
|---|---|---|
| full plane scan | 14.2x raw | 2.3x raw |
solution enumeration (--theme model) |
11.8x raw | 1.14x raw |
single warm probe |
1.09x raw | unchanged |
Two independent causes, neither of them block size — which the v0.7.5 experiment had suspected and correctly acquitted, then stopped at.
- The plane scan read one byte per block-cache call. Each
get()on a compressed table takes the cache mutex and copies a single byte; over a whole plane that, not the decompression, was the cost — decompressing both planes is only ~0.2 s of the 9.06 s the scan took. NewTableReader::read_values()pays one lock and one memcpy per block touched. - The block cache was smaller than the enumeration working set. Theme and shape queries probe at effectively random indices, and with the 4 MB budget nearly every probe paid a full block decompression for one byte. The budget is now 64 MB, capped per table at its own logical size. It is a ceiling, not a reservation — that run peaked at 62 MB RSS in total.
The ~6.5x figure documented since v0.7.5 is withdrawn: its benchmark exits as soon as it has 20000 hits and never scans the plane, understating the real cost by 3x.
helpmate stats summarises a whole table directory (0.8.2)
helpmate stats --tables DIR, with no MATERIAL, reports what is on disk and what is in it — formats, sizes, the whole-corpus compression ratio, the cell breakdown, the deepest mate, the largest tables. Covers the roadmap's helpmate list <dir> item.
tools/compress-corpus.sh
Converts a directory of raw tables into compressed ones in a different directory; compact --compress only works in place. One table at a time, source only read, nothing deleted, and it will not even open a table written in the last hour so an active generation run is never disturbed.
Verification
ctest 244/244, Catch2 3,866,730 assertions / 183 cases, lint / typecheck / format-check clean, api 83, bindings 15, web 15, repo 14, JS 36. mine output is byte-identical between raw and compressed tables for --dtm/--count, --starts and --theme, and row counts match the generator's own uniqueness histogram.
Upgrade note: no format change. Existing v1/v2/v3 tables are read exactly as before, and no corpus reconversion is needed.
What's Changed
- perf: mining a compressed table is no longer several times slower by @osick in #4
- feat: corpus-wide
helpmate stats, plus a raw→compressed conversion script by @osick in #5
Full Changelog: v0.8.0...v0.8.2