Repository navigation
Proposal: ADBC file drivers with pluggable storage #4877
CurtHagenlocher
started this conversation in
Ideas
Replies: 2 comments
|
in conjunction with apache/spark#54603 this could be used as a language-neutral extension point for file parsing from inside Spark. |
0 replies
|
"Vibe-coded" prototypes in https://github.com/CurtHagenlocher/arrow-adbc/tree/file-api |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Proposal: ADBC file drivers with pluggable storage
Status: early draft for discussion. Feedback on the direction is wanted
before any detailed specification.
Summary
ADBC today connects to databases and query engines. This proposal extends it
to the large amount of tabular data that lives in files with no server or
engine in front of them. Examples are Excel workbooks, HDF5 and NetCDF,
dBASE, fixed-width text, and a ZIP of CSVs.
It introduces three pieces:
and exposes its contents as tables. It does not do its own I/O or query
processing.
driver. The driver gets all of its bytes through it, which makes storage
(local disk, S3, GCS, Azure, HTTP, archives) pluggable and shared.
dynamically loaded storage providers, much as the driver manager already
finds and loads drivers. The core ships with a local file-system provider
only.
Query processing is left to the consumer, such as DataFusion, DuckDB, Polars,
or pandas. The spec would aim to standardize only what is needed to make that efficient:
projection, filter and limit pushdown, statistics, and partitioned reads.
Motivation
Lots of tabular data is not behind a query engine
Spreadsheets, scientific array formats, legacy desktop-database files, and
archives of delimited text are everywhere. Reading them as Arrow today means
finding a format-specific library for each language and engine, if one exists.
ODBC solved this, but without any sharing
ODBC had many file-based drivers: CSV/TSV, fixed-width, Excel, dBASE, Paradox,
Access, XML, JSON, Parquet, ORC, Avro, and more. Each one had to bundle three
separate things:
Shipping two such drivers meant shipping two storage layers and two query
engines, or inventing a private way to share them. There was no standard way
for independent drivers to share a common component.
Today's engines repeat the same split, in-process
DataFusion, DuckDB, Arrow Datasets, GDAL, and others each separate "format"
from "file system" internally. That separation lives behind each project's own
in-process API, so every format has to be implemented again for every engine.
There is no standard, cross-language, dynamically loadable boundary for a
file-format adapter.
Databases have been deconstructed
Arrow standardized the data, and ADBC standardized the driver API. The
remaining pieces, format logic and byte access, can now be standardized too.
Then a format adapter written once, in any language, can be loaded by any
ADBC-capable engine and read from any storage the host can reach.
Goals and non-goals
Goals
produce Arrow.
or similar themselves.
pushdown, statistics, and parallel partitioned reads.
Non-goals
excellent implementations. The value is in the long tail of formats.
later if bulk ingest into files proves useful.
is shaped for format readers.
Proposal
1. File drivers
A file driver is an ordinary ADBC driver, loaded through the existing driver
manager. It identifies itself as a file driver (for example through
GetInfo)and can declare the file extensions, MIME types, and magic-byte signatures it
recognizes. With that, a driver manager can optionally pick a driver
automatically for a given file.
The driver's contents map onto existing ADBC concepts:
GetObjects/GetTableTypesGetTableSchemaGetStatisticsExecutePartitions/ReadPartitionObject names are format-specific and opaque to ADBC, for example
Sheet1,Sheet1!A1:C3,/group/dataset, or2024/part-0.csv. A driver may alsoaccept a small SQL-like convenience syntax, but this would be optional and
not part of the core contract.
2. The file-system interface
When a file-driver database is created, the host supplies a file-system
object: a versioned C function table, in the same style as the rest of
ADBC. The driver never opens files itself. It calls back through this table
to open, stat, list, and read.
Key properties:
Implementing a provider, or consuming one from a driver, should be easy.
from the end of a file are supported, so a footer can be fetched in one
request. The provider is free to merge, split, or parallelize the ranges.
This is most of the performance win on high-latency object storage.
complete through a callback. A blocking form is also available. A provider
implements whichever form is natural for it, and the driver manager supplies
the other. Drivers built on
Read + Seekstyle libraries use the blockingform, while drivers that know their ranges ahead of time overlap I/O
asynchronously.
until that range's callback fires. Every accepted read gets exactly one
callback, including on error or cancellation. Callbacks never run inline
during submission and must only signal completion, never do real work.
Closing a file cancels outstanding reads and waits for their callbacks.
These few rules map directly onto futures and tasks in Rust, C++, C#, and Go.
bytes or fails. Files expose an opaque version (an ETag or generation) so
drivers can cache parsed metadata safely.
views, streaming input, stronger consistency guarantees, writes, or
integration with an engine's own I/O scheduler.
3. File-system registry in the driver manager
The driver manager gains a registry of storage providers, keyed by URI
scheme, similar to Hadoop's
fs.<scheme>.impl:s3,gs,abfss,http(s),hdfs, ...) are separate,dynamically loaded libraries, found and installed the same way drivers are.
interface or the file driver.
resolves the scheme, constructs the provider, and hands it to the driver.
file-system object. An example is an engine exposing its configured object
store. That way, the engine's credentials and caching are reused.
Archives are file systems, not formats. A
zipprovider can sit on topof any other provider. A "ZIP of CSVs on S3" then becomes a composition of
a CSV driver, a zip provider, and an S3 provider, and none of them needs to
know about the others.
4. Pushdown and statistics
Engines keep doing query processing, so the spec covers only what lets them
read less:
Substrait. For each filter term the driver reports whether it applied it
exactly, inexactly (so the engine re-checks), or not at all.
GetStatistics.A likely addition is per-partition statistics, so engines can skip
chunks or row groups before reading them.
Example
A DataFusion query joins:
PricesinC:\data\budget.xlsx, through an Excel filedriver over the built-in local provider; and
orders/*.csvinsides3://bucket/archive.zip, through a CSV file driverover a zip provider over an S3 provider.
Every driver and provider is loaded dynamically through standard interfaces.
DataFusion pushes the column lists and filters down, reads the partitions in
parallel, and performs the join itself.
Changes required in ADBC
enough for prototypes; a dedicated, versioned entry point is likely the
long-term answer.
GetInfoidentification, table types,option names, and how scans and pushdown are expressed on a statement.
Existing drivers and applications are unaffected.
Planned prototypes
directories and archives, and HDF5 or NetCDF, which bridges a library that
has its own I/O layer.
table provider that shows pushdown and (if possible) parallel partitioned reads.
Some questions for consideration
that ADBC merely uses?
What should
SetSqlQuerymean for a file driver?file driver be expected to understand?
details of the file system API. Where is the right place to put this?
(Apologies in advance for a fair chunk of LLM-generated text.)
All reactions