Skip to content

v13.0.0-rc.1

Pre-release
Pre-release

Choose a tag to compare

@lance-community lance-community released this 30 Sep 17:54
· 12 commits to main since this release

What's Changed

Breaking Changes 馃洜

  • perf(encoding)!: initialize only the page metadata a read will touch by @Ali2Arslan in #7465
  • feat(fts): add BM25F cross-field search by @sbrunk in #7905
  • feat(format): separate index keys from covering fields by @Ali2Arslan in #9159
  • feat!: expose file writer page options in dataset APIs by @ddupg in #9192
  • refactor: read JSON through one accessor in the JSON semantic module by @Xuanwo in #9479
  • fix!: derive merge_columns field ids from the dataset manifest by @zhangyue19921010 in #9547
  • perf(vector)!: score duplicate pairs in cache-blocked native tiles by @BubbleCal in #9551

New Features 馃帀

  • feat(java): add combined_fields (BM25F) full-text query by @sbrunk in #8549
  • feat(python): add combined_fields (BM25F) full-text query by @sbrunk in #8550
  • feat: expose exact V2 versions in bindings by @Xuanwo in #8585
  • feat(overlay): add a writer for data overlay files by @wjones127 in #8761
  • feat(encoding): implement plan_decoded_bytes for fixed-width structural decoder by @westonpace in #8792
  • feat(format): specify carried-column storage and allow carrying a keyed column by @vivek-bharathan in #8856
  • feat(io): apply byte-sized batch budget to blob materialization by @lichuang in #8927
  • feat(mem-wal): let a backpressure controller hold back memtable freezes by @hamersaw in #8994
  • feat(format): add a fragment metadata tree by @geruh in #9060
  • feat(index): implement stable partition mapping readers by @LuQQiu in #9064
  • feat(format): add the MinHash LSH scalar index specification (experimental) by @zhangyue19921010 in #9076
  • feat(core): define mapping readers with ordered compaction support by @LuQQiu in #9106
  • feat(index): apply shared FRI remapping to scalar and vector queries by @LuQQiu in #9107
  • feat(index): support MinHash LSH scalar index by @zhangyue19921010 in #9114
  • feat(format): feature flag for a fragment reuse index on stable row id tables by @brendanclement in #9119
  • feat: enable the fragment reuse index on stable row id datasets by @brendanclement in #9120
  • feat(index): accept a vector index a table is too small to train by @xuanyu-z in #9134
  • feat(format): define a unified tagged fragment reuse history by @LuQQiu in #9136
  • feat(mem_wal): resolve a schema change through one plan by @xuanyu-z in #9143
  • feat(mem_wal): evaluate compound FTS queries on the active memtable by @hamersaw in #9160
  • feat(mem_wal): cross-column FTS on the fresh tier by @hamersaw in #9172
  • feat: make commit retry timeout configurable by @mackrorysd in #9177
  • feat(python): expose registered dataset base paths by @ddupg in #9191
  • feat(io): expose Hugging Face resolve cache option by @Xuanwo in #9236
  • feat(format): define hidden row lineage columns by @BubbleCal in #9253
  • feat(knn): expose ANNIvfBatchExec's query, width and inputs by @hamersaw in #9262
  • feat(io): compile GooseFS metadata and page caches by @XuQianJin-Stars in #9263
  • feat(mem_wal): evaluate cross-column predicates on the fresh tier by @hamersaw in #9292
  • feat(knn): let a caller rebuild a batch KNN node with a new k by @hamersaw in #9334
  • feat(dataset): read and write spilled row lineage columns by @BubbleCal in #9336
  • feat(dataset): spill row lineage at compaction by @BubbleCal in #9337
  • feat(table): load spilled row lineage ahead of a commit by @BubbleCal in #9338
  • feat(dataset): spill row lineage when updating rows by @BubbleCal in #9339
  • feat(dataset): write compaction's spilled lineage into the fragment's data file by @BubbleCal in #9347
  • feat(io): add dataset label mode for object store metrics by @dshepelev15 in #9350
  • feat(mem_wal): serve single-column and nested multi-match on the fresh tier by @hamersaw in #9363
  • feat(java): expose add columns from stream api to java by @zhangyue19921010 in #9369
  • feat(io): accept block_size as a storage option by @dshepelev15 in #9380
  • feat(index): configure IVF shuffle offset preload budget by @jackye1995 in #9434
  • feat(fts): score unindexed fragments in combined_fields by @sbrunk in #9444
  • feat(blob): support mixed v1 and v2 files by @Xuanwo in #9459
  • feat(search): calibrate auto ivf probing for dot products by @BubbleCal in #9467
  • feat(java): add batch vector search binding by @sezruby in #9475
  • feat: check write inputs against the data file schema at a single boundary by @Xuanwo in #9478
  • feat(vector): add streaming duplicate pair enumeration by @BubbleCal in #9493
  • feat: prewrite MemWAL blob payloads by @jackye1995 in #9589
  • feat: expose ANN search stage timings by @Gabriel39 in #9602

Bug Fixes 馃悰

  • fix: reserve main branch name by @majin1102 in #7180
  • fix(namespace): reload rotated mTLS certificates in the REST client by @maswin in #8456
  • fix(fts): validate JSON index queries and targets by @app/lance-gatefixer in #8813
  • fix(label_list): bound index build memory with a sorted stream by @vivek-bharathan in #8840
  • fix(datafusion): coerce numeric literals to and from Float16 by @LuciferYang in #8847
  • fix(linalg): point length-contract panics at the distance function by @LuciferYang in #8864
  • fix(index): reject RabitQ indices whose dimension is not a multiple of 8 by @LuciferYang in #8869
  • fix(linalg): validate hamming batch layout by @app/lance-gatefixer in #8873
  • fix(linalg): preserve half-kernel fallback on AVX-512 hosts by @app/lance-gatefixer in #8877
  • fix(linalg): reject a null Int8 query instead of panicking by @LuciferYang in #8884
  • fix(index): use writer-side IVF shuffle offsets instead of re-decoding by @XuQianJin-Stars in #8941
  • fix: collect unreferenced data newer than retained manifests by @app/lance-gatefixer in #8944
  • fix(core): parse fixed-offset timezones in timestamp logical types by @jackylee-ch in #8950
  • fix(python): complete some straightforward API annotations by @jonasdedden in #8954
  • fix: honor nested projections in the mem-wal LSM scanner by @hamersaw in #8970
  • fix(index): keep the shuffler scratch directory alive for the whole build by @LuciferYang in #8993
  • fix(index): narrow f16/f64 training chunks to f32 in streaming IVF trainers by @LuciferYang in #8996
  • fix(index): reject a supplied PQ codebook that does not match the column by @LuciferYang in #8999
  • fix(python): accept numpy ivf_centroids without num_partitions by @LuciferYang in #9001
  • fix(index): report corrupt IVF metadata instead of aborting by @LuciferYang in #9006
  • fix(index): keep the requested index name after a failed uncommitted build by @LuciferYang in #9011
  • fix(fts): prevent overflow for unordered token positions by @app/lance-gatefixer in #9017
  • fix(encoding): emit NIL control words when rep/def levels are non-empty but zero-width by @lichuang in #9018
  • fix(table): surface a malformed manifest size from DynamoDB by @jackylee-ch in #9023
  • fix(index): reject RabitQ metadata without a code dimension by @LuciferYang in #9025
  • fix(merge_insert): surface a splitter panic as a stream error by @LuciferYang in #9085
  • fix(scanner): do not lower a negated InList into a take by @LuciferYang in #9097
  • fix(dataset): stop the fragment-reuse index blocking row id migration by @LuciferYang in #9099
  • fix: avoid index cache capacity fragmentation by @kamronis in #9115
  • fix(scanner): materialize list columns late as a unit by @zhangyue19921010 in #9148
  • fix(mem_wal): honor the query's distance type and ef in fresh-tier vector search by @hamersaw in #9163
  • fix(dataset): keep carried index bases through chained shallow clones by @LuQQiu in #9176
  • fix(datafusion): size the memory pool by the effective partition count by @LuQQiu in #9183
  • fix: preserve adaptive IVF probe bounds by @jackye1995 in #9184
  • fix: preserve deleted row zero in compaction remaps by @app/lance-gatefixer in #9203
  • fix: validate fully deleted row remaps by @app/lance-gatefixer in #9204
  • fix(deps): update rustls for RUSTSEC-2026-0285 by @Xuanwo in #9212
  • fix(index): stream prepared partitions through global top-k scoring by @BubbleCal in #9213
  • fix: preserve overlays across rebased deletes by @app/lance-gatefixer in #9218
  • fix: preserve projected files across rebased writes by @app/lance-gatefixer in #9219
  • fix(encoding): split legacy decode batches before i32 offset overflow by @app/lance-gatefixer in #9225
  • fix(dataset): handle evolved V2.0 struct projections by @app/lance-gatefixer in #9226
  • fix(dataset): fix open branch URI with tag to non-latest version by @dentiny in #9227
  • fix(dataset): reject a zero add_columns batch size instead of panicking by @zhangyue19921010 in #9233
  • fix: reject unsupported nulls in the add_columns stream path by @zhangyue19921010 in #9239
  • fix(python): support batched UInt8 Hamming queries by @ddupg in #9244
  • fix(rowids): reject duplicate values in a SortedArray segment by @jackylee-ch in #9245
  • fix: correct reserved metadata column error message typo by @ziwenzhang in #9248
  • fix(encoding): derive full-zip max_visible_def like the writer by @westonpace in #9254
  • fix(encoding): preserve special slots when recording validity by @Xuanwo in #9268
  • fix(linalg): gate L2 SimdSupport import on x86_64 by @Xuanwo in #9279
  • fix(namespace): avoid write starvation under heavy reads by @amunra in #9285
  • fix(dataset): reuse the prefetch window in BlobFile::read by @geruh in #9330
  • fix(index): prune stale label list entries during updates by @terapyon in #9352
  • fix(dataset): correct BlobFile seek semantics by @geruh in #9358
  • fix: a metadata update must not rebase over a restore by @wkalt in #9361
  • fix(io): allow goosefs:/// URLs to use site.properties by @XuQianJin-Stars in #9364
  • fix(mem_wal): build the transient FTS index with positions by @hamersaw in #9386
  • fix(python): normalize Arrow JSON in LanceFileWriter by @app/lance-gatefixer in #9394
  • fix: return actionable errors for binary/string offset overflow by @app/copilot-swe-agent in #9403
  • fix(datafusion): report the scan range in scan statistics by @vivek-bharathan in #9411
  • fix(format): move FLAG_UNKNOWN off the tagged FRI bit by @LuQQiu in #9420
  • fix(index): merge segments across a compaction without emptying the index by @xuanyu-z in #9421
  • fix: conflict an in-place column rewrite with a projection that dropped its field by @wkalt in #9436
  • fix(python): expose clone transactions by @everySympathy in #9443
  • fix(linalg): build on architectures without SIMD kernels by @sunyuechi in #9446
  • fix(commit): re-parent children when fix_schema renumbers a nested field by @shoemoney in #9447
  • fix(index): normalize dot k-means centroids by @BubbleCal in #9451
  • fix(index): make RaBitQ dot scoring query-scale invariant by @BubbleCal in #9452
  • fix(exec): preserve exact KNN ordering through late materialization by @BubbleCal in #9469
  • fix(core): reject dictionary types whose logical type cannot be parsed back by @LuciferYang in #9492
  • fix(core): release SharedStream state when one half is dropped by @LuciferYang in #9510
  • fix(core): make Dictionary equality reflexive for metadata-derived dictionaries by @LuciferYang in #9517
  • fix(merge_insert): build partial-schema fill as one projection by @timsaucer in #9523
  • fix: preserve carried external blobs in merge insert by @app/lance-gatefixer in #9532
  • fix: validate blob thresholds before schema commits by @app/lance-gatefixer in #9533
  • fix(index): preserve 4-bit PQ scores in bulk search by @Gabriel39 in #9537
  • fix(core): keep typed error signals across CloneableError::clone by @LuciferYang in #9542
  • fix: reject deep Clone commits that bypass deep_clone by @zhangyue19921010 in #9555
  • fix: tolerate temporary AWS STS errors in AwsCredentialAdapter by @cmccabe in #9561
  • fix: read schema-only nested columns added under a non-nullable parent by @LuQQiu in #9562
  • fix(python): preserve pyarrow expressions in mutation and fragment filters by @app/lance-gatefixer in #9573
  • fix: bound raw vector reads when splitting IVF partitions by @Xuanwo in #9574
  • fix(namespace): build the mTLS REST client with rustls by @Xuanwo in #9582
  • fix(label_list): split unnested batches to one sort chunk each by @vivek-bharathan in #9585
  • fix(mem_wal): dedup repeated row positions when training the flushed PK sidecar by @beinan in #9595
  • fix(encoding): skip full-width block bitpacking by @app/lance-gatefixer in #9598
  • fix(table): preserve spilled row lineage during updates by @majin1102 in #9623

Documentation 馃摎

  • docs: add experimental features to the governance docs by @westonpace in #8304
  • docs: correct nested vector index test description by @Xuanwo in #9210
  • docs: clarify sort spilling memory constraints by @Xuanwo in #9211
  • docs(knn): drop private intra-doc link that fails rustdoc on main by @hamersaw in #9410
  • docs: correct Blob v2 default storage thresholds by @ningsh7 in #9538
  • docs(java): correct the deprecated setIndexCacheSize Javadoc by @jackylee-ch in #9559

Performance Improvements 馃殌

  • perf: stop reading past the first manifest when resolving the latest version by @app/copilot-swe-agent in #8729
  • perf(linalg): let AVX-512 hosts fall back to the AVX2 dist_table kernel by @LuciferYang in #8866
  • perf(fts): decode exact-phrase positions from the rarest term first by @mikewhb in #9170
  • perf(fts): complete MaxScore inner windows with SoA and top2-gap by @mikewhb in #9171
  • perf: buffer BlobFile sequential reads in Rust by @geruh in #9251
  • perf(linalg): restore u8 scalar L2 performance by @LeoReeYang in #9259
  • perf(python): import the namespace client lazily by @dshepelev15 in #9273
  • perf(io): decode large metadata messages straight from fetched chunks by @dshepelev15 in #9274
  • perf(table): keep inline row ids as Bytes slices of the manifest buffer by @dshepelev15 in #9275
  • perf(dataset): open a fragment's data files concurrently by @dshepelev15 in #9276
  • perf(table): make U64Segment::len O(1) for RangeWithHoles by @dshepelev15 in #9277
  • perf(encoding): build page schedulers only for requested pages by @BubbleCal in #9278
  • perf(frag-reuse): hash the remap's fragment-id map with an integer hasher by @amunra in #9288
  • perf(index): coalesce btree page loads into one read_ranges request by @westonpace in #9310
  • perf(table): fold clustered deletions into range segments by @dshepelev15 in #9311
  • perf: prune V2 LIMIT/OFFSET scans by live row counts by @LeoReeYang in #9344
  • perf(fts): cut per-window and per-candidate overhead in bulk AND search by @BubbleCal in #9354
  • perf(encoding): batch page reads of a structural column into one I/O request by @dshepelev15 in #9379
  • perf(cleanup): concurrent deletion and robustness fixes by @jackye1995 in #9414
  • perf: share one ScanScheduler per base within a scan by @LuQQiu in #9417
  • perf(table): sparse mask_to_offset_ranges for RangeWithBitmap segments by @LeoReeYang in #9437
  • perf: prune scalar index segments by fragment scope by @ddupg in #9449
  • perf(fts): track posting validation with one byte per token by @LuQQiu in #9457
  • perf(index): cache index file metadata across opens by @BubbleCal in #9461
  • perf(index): assemble btree search results once instead of per page by @BubbleCal in #9462
  • perf(index): stream btree pages sequentially for update, remap and prewarm by @BubbleCal in #9463
  • perf: skip full-snapshot row-id prefilters and expose loader metrics by @Gabriel39 in #9471
  • perf(encoding): batch a scheduling step's reads across columns into one I/O request by @dshepelev15 in #9473
  • perf(vector): score duplicate pairs with SIMD group kernels by @BubbleCal in #9563
  • perf(index): keep the HNSW graph when remapping IVF_HNSW partitions by @james-rms in #9590
  • perf: skip redundant ANN segment row-id prefilters by @Gabriel39 in #9599
  • perf(fts): cut per-partition fixed cost of match queries by @BubbleCal in #9612
  • perf(datafusion): stop cloning the session state for every execution by @BubbleCal in #9614

Other Changes

  • refactor(fts): make HybridCompoundQueryExec public by @hamersaw in #9166
  • refactor: use DataFusion standard nested field accessor by @Xuanwo in #9207
  • refactor(io): simplify single-range object reads by @Xuanwo in #9241
  • perf(fts): reduce per-token index build memory by @BubbleCal in #9613

Full Changelog: release-root/13.0.0-beta.N...v13.0.0-rc.1