BlobV2Descriptor Structured Extension — Fine-Grained Random Access for Multimodal Blob #8739
wenxuanguan
started this conversation in
Lance File Format
Replies: 1 comment 1 reply
|
Precomputing the media index makes sense to me. Could we compare this with keeping the index in a regular column and projecting it alongside the blob descriptor? We already have If we do need storage-level support, could this fit into the Object Layer design? That seems like a natural home for representation-specific metadata. |
1 reply
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Background & Motivation
Blob v2 already provides efficient storage and access for large binary objects. However, the current BlobV2Descriptor only records the blob's overall position and size (kind / position / size / blob_id / blob_uri) and is unaware of the blob's internal structure.
This limits Lance's applicability in multimodal scenarios:
Building on prior discussions, #7174 points out that the Blob v2 descriptor only expresses physical storage location and cannot support richer logical objects such as multiple rows sharing the same object, delta encoding, or chunked blobs. It proposes an Object Layer abstraction that decouples physical location from logical representation through a unified resolve → plan → read → decode flow, enabling domain-specific blob layouts (e.g., GOP-level video storage) and cross-row object sharing (I/O coalescing). #4320 points out that the Blob API treats video as unstructured binary data and cannot perceive video metadata (e.g., frame rate); it proposes a video encoding extension that pluggably supports "encode-on-write" and "background progressive rewrite (at GOP level)" while retaining the native strengths of columnar storage.
I'd like to suggest a minimal, version-gated (file version >= 2.3) BlobV2Descriptor extension on top of these discussions, so that Lance can perceive a blob's internal structure and support fine-grained random access, while keeping the Lance core free of any codec dependency.
Core Design: A Minimal Extension to BlobV2Descriptor
On top of the existing 5 native fields, the only addition is a nullable sub-field blob_info, which carries the metadata and internal-structure index for various multimodal blobs. data_storage_version >= 2.3 allows the six-field layout. The existing 5 fields are unchanged. The full structure is defined as follows:
blob_info is a fixed Arrow envelope in the format specification (stream_info : binary, entries : binary), not a client-defined nested Struct. The binary contents of these two children are user-defined. Lance is unaware of their content or semantics; it only stores, validates, projects, and copies them, while their parsing is defined by a user-provided plugin. For video, this might include stream-level metadata (e.g., fps) and GOP/Frame offsets. For image, it could contain page-level metadata such as page count, and index entries for a multi-page document.
Architecture: How It Maps Onto the Object Layer
This proposal follows the implementation path of #7174 Object Layer. It adheres to resolve → plan → read → decode, using video frame extraction as an example.
Mapping of the blob read path:
blob_info is inlined in the descriptor, so the number of I/O reads stays the same.
Custom Plugins
Media processing capabilities are implemented by user-provided plugins. Lance passes the source blob and opaque configuration through a unified plugin interface, and receives the processed byte stream, stream_info, and entries produced by the plugin. Lance is unaware of the data type handled by the plugin and does not parse the contents of the two blob_info children.
On reads, the plugin parses stream_info and entries according to its own definitions, plans the required byte ranges, and decodes the bytes returned by Lance. lance-media-runtime is responsible for loading the plugin; Lance core only handles data storage and range reads, with no built-in codec dependencies such as FFmpeg.
On writes, do not extend write_dataset. Add write_dataset_with_blob_info(..., media_options={column_path: (library_path, config_binary)}). Configured Packed/Dedicated columns (version >= 2.3) go through MediaProvider::ingest: write the processed byte stream (e.g. GOP) and return stream_info / entries. Unconfigured columns and Inline keep the original path with blob_info = null. No video frame decoding or re-encoding is involved.
The boundary of changes is clear:
API Description
Using video as an example, this proposal adds the following two user-facing APIs:
The existing write_dataset, take_blobs, read_blobs, and BlobFile APIs gain no new parameters and retain their existing semantics.
Compatibility
Taking video frame-extraction scenarios as an example. Stored bytes, blob_info, and decoded frames are three different logical values: take_blobs / read_blobs return stored bytes only; extract_frame returns frames.
The proposal lets the plugin decompose the source container and stores the processed byte stream (e.g. GOP) in the Lance Dataset. Since original container metadata is stripped during ingest, legacy interfaces cannot correctly interpret these bytes or reproduce a valid MP4, so they error instead of returning them.
Performance Gains
Taking video frame extraction as an example, the blob_info extension reduces the number of I/O operations by an order of magnitude compared to the BlobFile-based random access approach, thereby significantly lowering the I/O cost for multimodal frame-extraction workloads.
End-to-End Latency Comparison
Test setup: MP4 video stored on MinIO (133.4 MB, H.264, fps 25, 40 min duration); Ten 500 ms time windows are randomly generated, and ten corresponding keyframe extraction operations are performed.The test implementation follows the official Lance lazy video decoding example: example-decode-video-frames-lazily. In that example, the
process_framefunction performs RGB conversion and materializes the frame as an ndarray.This proposal's extraction flow: extract_frame locates and retrieves the target GOP from the blob according to the time interval, and then decodes to all keyframes using the media plugin.
Baseline extraction flow: Using take_blobs to fetch the BlobFile, a container is built from it. The container is then seeked to the closest keyframe position based on the given start time, and all keyframes are subsequently extracted in sequential order.
Results:
Performance Breakdown
This proposal decodes from the GOP's keyframe sequentially to the target frame; the baseline computes a seek to the preceding keyframe from the frame number and fps (and related info), then decodes sequentially to the target frame. Decode times are close — the performance gap mainly comes from blob data reads.
Comparing the I/O traffic of reading blob data, the difference is significant:
Other approaches
Other approaches can achieve similar results, such as storing the internal-structure index in a separate user-table column; this adds one sequential projection and requires users to keep the blob and index column synchronized across append / compact / merge. Inlining blob_info in the descriptor keeps physical location and internal index as one piece of row-level metadata, reducing logical reads from 3 to 2.
Test setup: on remote MinIO, 30 real H.264 MP4s (about 7 MB each, 30 s duration, fps 25, 720P/1080P) with 30 real frame PTS values. The full flow is: project descriptor / blob_info → rebuild BlobEntry → plan GOP range → read GOP payload; timing stops when the payload is in memory, with no frame decode. The inlined layout obtains descriptor + blob_info in one projection; the separate-column layout projects the index column first, then the descriptor.
All reactions