v13.0.0-rc.1
Pre-release
Pre-release
·
12 commits
to main
since this release
What's Changed
Breaking Changes 馃洜
- perf(encoding)!: initialize only the page metadata a read will touch by @Ali2Arslan in #7465
- feat(fts): add BM25F cross-field search by @sbrunk in #7905
- feat(format): separate index keys from covering fields by @Ali2Arslan in #9159
- feat!: expose file writer page options in dataset APIs by @ddupg in #9192
- refactor: read JSON through one accessor in the JSON semantic module by @Xuanwo in #9479
- fix!: derive merge_columns field ids from the dataset manifest by @zhangyue19921010 in #9547
- perf(vector)!: score duplicate pairs in cache-blocked native tiles by @BubbleCal in #9551
New Features 馃帀
- feat(java): add combined_fields (BM25F) full-text query by @sbrunk in #8549
- feat(python): add combined_fields (BM25F) full-text query by @sbrunk in #8550
- feat: expose exact V2 versions in bindings by @Xuanwo in #8585
- feat(overlay): add a writer for data overlay files by @wjones127 in #8761
- feat(encoding): implement plan_decoded_bytes for fixed-width structural decoder by @westonpace in #8792
- feat(format): specify carried-column storage and allow carrying a keyed column by @vivek-bharathan in #8856
- feat(io): apply byte-sized batch budget to blob materialization by @lichuang in #8927
- feat(mem-wal): let a backpressure controller hold back memtable freezes by @hamersaw in #8994
- feat(format): add a fragment metadata tree by @geruh in #9060
- feat(index): implement stable partition mapping readers by @LuQQiu in #9064
- feat(format): add the MinHash LSH scalar index specification (experimental) by @zhangyue19921010 in #9076
- feat(core): define mapping readers with ordered compaction support by @LuQQiu in #9106
- feat(index): apply shared FRI remapping to scalar and vector queries by @LuQQiu in #9107
- feat(index): support MinHash LSH scalar index by @zhangyue19921010 in #9114
- feat(format): feature flag for a fragment reuse index on stable row id tables by @brendanclement in #9119
- feat: enable the fragment reuse index on stable row id datasets by @brendanclement in #9120
- feat(index): accept a vector index a table is too small to train by @xuanyu-z in #9134
- feat(format): define a unified tagged fragment reuse history by @LuQQiu in #9136
- feat(mem_wal): resolve a schema change through one plan by @xuanyu-z in #9143
- feat(mem_wal): evaluate compound FTS queries on the active memtable by @hamersaw in #9160
- feat(mem_wal): cross-column FTS on the fresh tier by @hamersaw in #9172
- feat: make commit retry timeout configurable by @mackrorysd in #9177
- feat(python): expose registered dataset base paths by @ddupg in #9191
- feat(io): expose Hugging Face resolve cache option by @Xuanwo in #9236
- feat(format): define hidden row lineage columns by @BubbleCal in #9253
- feat(knn): expose ANNIvfBatchExec's query, width and inputs by @hamersaw in #9262
- feat(io): compile GooseFS metadata and page caches by @XuQianJin-Stars in #9263
- feat(mem_wal): evaluate cross-column predicates on the fresh tier by @hamersaw in #9292
- feat(knn): let a caller rebuild a batch KNN node with a new k by @hamersaw in #9334
- feat(dataset): read and write spilled row lineage columns by @BubbleCal in #9336
- feat(dataset): spill row lineage at compaction by @BubbleCal in #9337
- feat(table): load spilled row lineage ahead of a commit by @BubbleCal in #9338
- feat(dataset): spill row lineage when updating rows by @BubbleCal in #9339
- feat(dataset): write compaction's spilled lineage into the fragment's data file by @BubbleCal in #9347
- feat(io): add dataset label mode for object store metrics by @dshepelev15 in #9350
- feat(mem_wal): serve single-column and nested multi-match on the fresh tier by @hamersaw in #9363
- feat(java): expose add columns from stream api to java by @zhangyue19921010 in #9369
- feat(io): accept block_size as a storage option by @dshepelev15 in #9380
- feat(index): configure IVF shuffle offset preload budget by @jackye1995 in #9434
- feat(fts): score unindexed fragments in combined_fields by @sbrunk in #9444
- feat(blob): support mixed v1 and v2 files by @Xuanwo in #9459
- feat(search): calibrate auto ivf probing for dot products by @BubbleCal in #9467
- feat(java): add batch vector search binding by @sezruby in #9475
- feat: check write inputs against the data file schema at a single boundary by @Xuanwo in #9478
- feat(vector): add streaming duplicate pair enumeration by @BubbleCal in #9493
- feat: prewrite MemWAL blob payloads by @jackye1995 in #9589
- feat: expose ANN search stage timings by @Gabriel39 in #9602
Bug Fixes 馃悰
- fix: reserve main branch name by @majin1102 in #7180
- fix(namespace): reload rotated mTLS certificates in the REST client by @maswin in #8456
- fix(fts): validate JSON index queries and targets by @app/lance-gatefixer in #8813
- fix(label_list): bound index build memory with a sorted stream by @vivek-bharathan in #8840
- fix(datafusion): coerce numeric literals to and from Float16 by @LuciferYang in #8847
- fix(linalg): point length-contract panics at the distance function by @LuciferYang in #8864
- fix(index): reject RabitQ indices whose dimension is not a multiple of 8 by @LuciferYang in #8869
- fix(linalg): validate hamming batch layout by @app/lance-gatefixer in #8873
- fix(linalg): preserve half-kernel fallback on AVX-512 hosts by @app/lance-gatefixer in #8877
- fix(linalg): reject a null Int8 query instead of panicking by @LuciferYang in #8884
- fix(index): use writer-side IVF shuffle offsets instead of re-decoding by @XuQianJin-Stars in #8941
- fix: collect unreferenced data newer than retained manifests by @app/lance-gatefixer in #8944
- fix(core): parse fixed-offset timezones in timestamp logical types by @jackylee-ch in #8950
- fix(python): complete some straightforward API annotations by @jonasdedden in #8954
- fix: honor nested projections in the mem-wal LSM scanner by @hamersaw in #8970
- fix(index): keep the shuffler scratch directory alive for the whole build by @LuciferYang in #8993
- fix(index): narrow f16/f64 training chunks to f32 in streaming IVF trainers by @LuciferYang in #8996
- fix(index): reject a supplied PQ codebook that does not match the column by @LuciferYang in #8999
- fix(python): accept numpy ivf_centroids without num_partitions by @LuciferYang in #9001
- fix(index): report corrupt IVF metadata instead of aborting by @LuciferYang in #9006
- fix(index): keep the requested index name after a failed uncommitted build by @LuciferYang in #9011
- fix(fts): prevent overflow for unordered token positions by @app/lance-gatefixer in #9017
- fix(encoding): emit NIL control words when rep/def levels are non-empty but zero-width by @lichuang in #9018
- fix(table): surface a malformed manifest size from DynamoDB by @jackylee-ch in #9023
- fix(index): reject RabitQ metadata without a code dimension by @LuciferYang in #9025
- fix(merge_insert): surface a splitter panic as a stream error by @LuciferYang in #9085
- fix(scanner): do not lower a negated InList into a take by @LuciferYang in #9097
- fix(dataset): stop the fragment-reuse index blocking row id migration by @LuciferYang in #9099
- fix: avoid index cache capacity fragmentation by @kamronis in #9115
- fix(scanner): materialize list columns late as a unit by @zhangyue19921010 in #9148
- fix(mem_wal): honor the query's distance type and ef in fresh-tier vector search by @hamersaw in #9163
- fix(dataset): keep carried index bases through chained shallow clones by @LuQQiu in #9176
- fix(datafusion): size the memory pool by the effective partition count by @LuQQiu in #9183
- fix: preserve adaptive IVF probe bounds by @jackye1995 in #9184
- fix: preserve deleted row zero in compaction remaps by @app/lance-gatefixer in #9203
- fix: validate fully deleted row remaps by @app/lance-gatefixer in #9204
- fix(deps): update rustls for RUSTSEC-2026-0285 by @Xuanwo in #9212
- fix(index): stream prepared partitions through global top-k scoring by @BubbleCal in #9213
- fix: preserve overlays across rebased deletes by @app/lance-gatefixer in #9218
- fix: preserve projected files across rebased writes by @app/lance-gatefixer in #9219
- fix(encoding): split legacy decode batches before i32 offset overflow by @app/lance-gatefixer in #9225
- fix(dataset): handle evolved V2.0 struct projections by @app/lance-gatefixer in #9226
- fix(dataset): fix open branch URI with tag to non-latest version by @dentiny in #9227
- fix(dataset): reject a zero add_columns batch size instead of panicking by @zhangyue19921010 in #9233
- fix: reject unsupported nulls in the
add_columnsstream path by @zhangyue19921010 in #9239 - fix(python): support batched UInt8 Hamming queries by @ddupg in #9244
- fix(rowids): reject duplicate values in a SortedArray segment by @jackylee-ch in #9245
- fix: correct reserved metadata column error message typo by @ziwenzhang in #9248
- fix(encoding): derive full-zip max_visible_def like the writer by @westonpace in #9254
- fix(encoding): preserve special slots when recording validity by @Xuanwo in #9268
- fix(linalg): gate L2 SimdSupport import on x86_64 by @Xuanwo in #9279
- fix(namespace): avoid write starvation under heavy reads by @amunra in #9285
- fix(dataset): reuse the prefetch window in BlobFile::read by @geruh in #9330
- fix(index): prune stale label list entries during updates by @terapyon in #9352
- fix(dataset): correct BlobFile seek semantics by @geruh in #9358
- fix: a metadata update must not rebase over a restore by @wkalt in #9361
- fix(io): allow goosefs:/// URLs to use site.properties by @XuQianJin-Stars in #9364
- fix(mem_wal): build the transient FTS index with positions by @hamersaw in #9386
- fix(python): normalize Arrow JSON in LanceFileWriter by @app/lance-gatefixer in #9394
- fix: return actionable errors for binary/string offset overflow by @app/copilot-swe-agent in #9403
- fix(datafusion): report the scan range in scan statistics by @vivek-bharathan in #9411
- fix(format): move FLAG_UNKNOWN off the tagged FRI bit by @LuQQiu in #9420
- fix(index): merge segments across a compaction without emptying the index by @xuanyu-z in #9421
- fix: conflict an in-place column rewrite with a projection that dropped its field by @wkalt in #9436
- fix(python): expose clone transactions by @everySympathy in #9443
- fix(linalg): build on architectures without SIMD kernels by @sunyuechi in #9446
- fix(commit): re-parent children when fix_schema renumbers a nested field by @shoemoney in #9447
- fix(index): normalize dot k-means centroids by @BubbleCal in #9451
- fix(index): make RaBitQ dot scoring query-scale invariant by @BubbleCal in #9452
- fix(exec): preserve exact KNN ordering through late materialization by @BubbleCal in #9469
- fix(core): reject dictionary types whose logical type cannot be parsed back by @LuciferYang in #9492
- fix(core): release SharedStream state when one half is dropped by @LuciferYang in #9510
- fix(core): make Dictionary equality reflexive for metadata-derived dictionaries by @LuciferYang in #9517
- fix(merge_insert): build partial-schema fill as one projection by @timsaucer in #9523
- fix: preserve carried external blobs in merge insert by @app/lance-gatefixer in #9532
- fix: validate blob thresholds before schema commits by @app/lance-gatefixer in #9533
- fix(index): preserve 4-bit PQ scores in bulk search by @Gabriel39 in #9537
- fix(core): keep typed error signals across CloneableError::clone by @LuciferYang in #9542
- fix: reject deep
Clonecommits that bypassdeep_cloneby @zhangyue19921010 in #9555 - fix: tolerate temporary AWS STS errors in AwsCredentialAdapter by @cmccabe in #9561
- fix: read schema-only nested columns added under a non-nullable parent by @LuQQiu in #9562
- fix(python): preserve pyarrow expressions in mutation and fragment filters by @app/lance-gatefixer in #9573
- fix: bound raw vector reads when splitting IVF partitions by @Xuanwo in #9574
- fix(namespace): build the mTLS REST client with rustls by @Xuanwo in #9582
- fix(label_list): split unnested batches to one sort chunk each by @vivek-bharathan in #9585
- fix(mem_wal): dedup repeated row positions when training the flushed PK sidecar by @beinan in #9595
- fix(encoding): skip full-width block bitpacking by @app/lance-gatefixer in #9598
- fix(table): preserve spilled row lineage during updates by @majin1102 in #9623
Documentation 馃摎
- docs: add experimental features to the governance docs by @westonpace in #8304
- docs: correct nested vector index test description by @Xuanwo in #9210
- docs: clarify sort spilling memory constraints by @Xuanwo in #9211
- docs(knn): drop private intra-doc link that fails rustdoc on main by @hamersaw in #9410
- docs: correct Blob v2 default storage thresholds by @ningsh7 in #9538
- docs(java): correct the deprecated setIndexCacheSize Javadoc by @jackylee-ch in #9559
Performance Improvements 馃殌
- perf: stop reading past the first manifest when resolving the latest version by @app/copilot-swe-agent in #8729
- perf(linalg): let AVX-512 hosts fall back to the AVX2 dist_table kernel by @LuciferYang in #8866
- perf(fts): decode exact-phrase positions from the rarest term first by @mikewhb in #9170
- perf(fts): complete MaxScore inner windows with SoA and top2-gap by @mikewhb in #9171
- perf: buffer BlobFile sequential reads in Rust by @geruh in #9251
- perf(linalg): restore u8 scalar L2 performance by @LeoReeYang in #9259
- perf(python): import the namespace client lazily by @dshepelev15 in #9273
- perf(io): decode large metadata messages straight from fetched chunks by @dshepelev15 in #9274
- perf(table): keep inline row ids as Bytes slices of the manifest buffer by @dshepelev15 in #9275
- perf(dataset): open a fragment's data files concurrently by @dshepelev15 in #9276
- perf(table): make U64Segment::len O(1) for RangeWithHoles by @dshepelev15 in #9277
- perf(encoding): build page schedulers only for requested pages by @BubbleCal in #9278
- perf(frag-reuse): hash the remap's fragment-id map with an integer hasher by @amunra in #9288
- perf(index): coalesce btree page loads into one read_ranges request by @westonpace in #9310
- perf(table): fold clustered deletions into range segments by @dshepelev15 in #9311
- perf: prune V2 LIMIT/OFFSET scans by live row counts by @LeoReeYang in #9344
- perf(fts): cut per-window and per-candidate overhead in bulk AND search by @BubbleCal in #9354
- perf(encoding): batch page reads of a structural column into one I/O request by @dshepelev15 in #9379
- perf(cleanup): concurrent deletion and robustness fixes by @jackye1995 in #9414
- perf: share one ScanScheduler per base within a scan by @LuQQiu in #9417
- perf(table): sparse mask_to_offset_ranges for RangeWithBitmap segments by @LeoReeYang in #9437
- perf: prune scalar index segments by fragment scope by @ddupg in #9449
- perf(fts): track posting validation with one byte per token by @LuQQiu in #9457
- perf(index): cache index file metadata across opens by @BubbleCal in #9461
- perf(index): assemble btree search results once instead of per page by @BubbleCal in #9462
- perf(index): stream btree pages sequentially for update, remap and prewarm by @BubbleCal in #9463
- perf: skip full-snapshot row-id prefilters and expose loader metrics by @Gabriel39 in #9471
- perf(encoding): batch a scheduling step's reads across columns into one I/O request by @dshepelev15 in #9473
- perf(vector): score duplicate pairs with SIMD group kernels by @BubbleCal in #9563
- perf(index): keep the HNSW graph when remapping IVF_HNSW partitions by @james-rms in #9590
- perf: skip redundant ANN segment row-id prefilters by @Gabriel39 in #9599
- perf(fts): cut per-partition fixed cost of match queries by @BubbleCal in #9612
- perf(datafusion): stop cloning the session state for every execution by @BubbleCal in #9614
Other Changes
- refactor(fts): make HybridCompoundQueryExec public by @hamersaw in #9166
- refactor: use DataFusion standard nested field accessor by @Xuanwo in #9207
- refactor(io): simplify single-range object reads by @Xuanwo in #9241
- perf(fts): reduce per-token index build memory by @BubbleCal in #9613
Full Changelog: release-root/13.0.0-beta.N...v13.0.0-rc.1