A high-performance Rust library for local Apache Cassandra SSTable access
Status: v0.13.0 — the performance release. Core reading, CLI, output writers, Python & Node.js bindings, and write support are production-ready, now with read-path constant-factor speedups, byte-bounded result budgets, explicit
Database.refresh(), no-heuristics correctness fixes, and byte-for-byte compaction parity against Apache Cassandra, an Arrow Flight + Trino connector, canonical BTI (da) write/read, and CDC-style delta export. See CHANGELOG.md.
Upgrading to v0.13? See the v0.13 Migration Guide for the 3 breaking changes (Python duration/time types, unknown-table errors, CLI YAML removal).
CQLite provides SQLite-like local access to Apache Cassandra SSTables, enabling developers to read Cassandra 5.0+ data files without cluster dependencies. Built in Rust for performance and safety.
⭐ Find CQLite useful? Star the repo — it is the clearest signal that this work matters and directly drives how much time goes into it. 🐛 Hit a bug or need a feature? Open an issue. For questions and ideas, use Discussions. See Known Issues and the Roadmap before filing.
Full documentation is at https://pmcfadin.github.io/cqlite/:
| Section | URL |
|---|---|
| User Docs — install, quick start, CLI, Python, Node.js | /cqlite/user-docs/ |
| SSTable Format Guide — binary format deep-dive | /cqlite/sstable-format/ |
| For Agents: Using CQLite — LLM/agent integration | /cqlite/agents-using/ |
| For Agents: Developing CQLite — contributor doctrine, gate contract | /cqlite/agents-developing/ |
CQLite aims to become the standard tool for Cassandra SSTable manipulation outside of the main Apache Cassandra project, enabling new workflows for data analytics, migration, testing, and edge computing.
CQLite is designed by Patrick McFadin, Apache Cassandra PMC member with over a decade of Cassandra experience. The project embodies Apache Cassandra community values and will be donated to the Apache Cassandra project upon maturity.
The quickest path on macOS (Apple Silicon or Intel) and Linux (x86_64 or arm64). The formula verifies the release checksum before installing:
brew install pmcfadin/cqlite/cqlite
cqlite --helpcargo install cqlite-cli # installs the `cqlite` binary
cqlite --helpEach GitHub release attaches a
prebuilt cqlite CLI binary for the common platforms, each with a .sha256
checksum sidecar:
| Platform | Asset |
|---|---|
| macOS (Apple Silicon) | cqlite-aarch64-apple-darwin.tar.gz |
| macOS (Intel) | cqlite-x86_64-apple-darwin.tar.gz |
| Linux x86_64 (glibc) | cqlite-x86_64-unknown-linux-gnu.tar.gz |
| Linux x86_64 (static musl) | cqlite-x86_64-unknown-linux-musl.tar.gz |
| Linux arm64 (glibc) | cqlite-aarch64-unknown-linux-gnu.tar.gz |
| Windows x86_64 | cqlite-x86_64-pc-windows-gnu.zip |
# Example: macOS Apple Silicon
TARGET=aarch64-apple-darwin
curl -fsSLO https://github.com/pmcfadin/cqlite/releases/latest/download/cqlite-$TARGET.tar.gz
curl -fsSLO https://github.com/pmcfadin/cqlite/releases/latest/download/cqlite-$TARGET.tar.gz.sha256
shasum -a 256 -c cqlite-$TARGET.tar.gz.sha256 # verify (use sha256sum -c on Linux)
tar xzf cqlite-$TARGET.tar.gz
./cqlite --helpcargo add cqlite-core # use cqlite-core as a dependencySee Using cqlite-core as a dependency and the API docs.
pip install cqlite-py # Python
npm install @cqlite/node # Node.jsQuery a Cassandra node's SSTables over Arrow Flight (gRPC) with the
cqlite-flight server, published as a multi-arch image on every release tag.
Mount the data dir read-only and point --data-dir at it:
docker run --rm -p 8815:8815 \
-v /var/lib/cassandra:/var/lib/cassandra:ro \
ghcr.io/pmcfadin/cqlite-flight:latest \
--data-dir /var/lib/cassandra/data --listen 0.0.0.0:8815See cqlite-flight/README.md for image tags, the
ticket/predicate API, and the trino-connector that builds
on it.
# Clone the repository
git clone https://github.com/pmcfadin/cqlite.git
cd cqlite
# Build the project
cargo build --release
# Run the CLI tool
cargo run --package cqlite-cli -- \
--schema test-data/schemas/basic-types.cql \
--data-dir test-data/datasets/sstables \
--query "SELECT * FROM test_basic.simple_table LIMIT 5" \
--out jsonpip install cqlite-pyimport cqlite
with cqlite.open('path/to/sstables', schema='schema.cql') as db:
for row in db.execute('SELECT * FROM keyspace.table LIMIT 5'):
print(row.to_dict())npm install @cqlite/nodeimport { Database } from '@cqlite/node';
const db = await Database.open('path/to/sstables', { schema: 'schema.cql' });
const result = await db.execute('SELECT * FROM keyspace.table LIMIT 5');
for (const row of result.rows) {
console.log(row.name);
}
await db.close();CQLite v0.9.0 (M5) ships write support across all interfaces: Rust core, Python,
Node.js, and CLI. Written data flushes to portable Cassandra 5.0 SSTables that
Cassandra can read directly via nodetool refresh.
The schema file below is included in the repository at
test-data/schemas/write-test.cql.
import cqlite
# Open in writable mode — write_dir stores the WAL and flushed SSTables
with cqlite.open(
'test-data/datasets/sstables',
schema='test-data/schemas/write-test.cql',
writable=True,
write_dir='/tmp/my-writes',
) as db:
db.execute(
"INSERT INTO test_basic.simple_table (id, name, age) "
"VALUES (11111111-1111-1111-1111-111111111111, 'Alice', 30)"
)
path = db.flush_run()
print(f'Flushed SSTable: {path}')const { Database } = require('@cqlite/node');
const db = await Database.open('test-data/datasets/sstables', {
schema: 'test-data/schemas/write-test.cql',
writable: true,
writeDir: '/tmp/my-writes',
});
await db.execute(
"INSERT INTO test_basic.simple_table (id, name, age) " +
"VALUES (22222222-2222-2222-2222-222222222222, 'Bob', 25)"
);
const path = await db.flushRun();
console.log('Flushed SSTable:', path);
await db.close();# Build with write support
cargo build --package cqlite-cli --features write-support
# Write via CQL INSERT
cargo run --package cqlite-cli --features write-support -- \
--writable --write-dir /tmp/my-writes \
--schema test-data/schemas/write-test.cql \
--execute "INSERT INTO test_basic.simple_table (id, name, age) \
VALUES (33333333-3333-3333-3333-333333333333, 'Carol', 28)"
# Flush memtable to SSTable
cargo run --package cqlite-cli --features write-support -- \
--writable --write-dir /tmp/my-writes \
--schema test-data/schemas/write-test.cql \
--flushSee docs/write-support.md for the full write guide,
including the Cassandra export workflow and known limitations. To embed
cqlite-core in your own Rust project (dependency line, feature flags, and a
compiling write example), see
docs/using-cqlite-core-as-a-dependency.md.
cqlite-core gates optional functionality behind Cargo features. The table below
maps the public API you're likely to reach for to the feature that enables it.
| Want… | Enable feature | In defaults? |
|---|---|---|
Read / query path (Database::open, execute, scan, get) |
state_machine |
✅ yes |
| Compression (LZ4 / Snappy / Deflate / Zstd) | all-compression |
✅ yes |
Write path (WriteEngine, Mutation, WriteEngine::write/flush) |
write-support |
✅ yes |
Database::flush / Database::compact (high-level convenience) |
experimental |
❌ opt-in |
CLI ingestion / REPL helpers (cqlite-cli) |
cli-helpers |
❌ opt-in |
| Performance metrics collection | metrics |
❌ opt-in |
Default features are ["all-compression", "state_machine", "write-support"]
(see cqlite-core/Cargo.toml). write-support was folded into the defaults in
#558 — it gates only first-party
code and adds no extra dependencies, so read-only consumers pay nothing for it.
flush/compact on the high-level Database type remain behind experimental;
the equivalent engine-level WriteEngine::flush is part of write-support.
# Default build (read + write + compression)
cargo build
# Read-only consumer: drop the write path (still zero-cost to keep it, but explicit)
cargo build -p cqlite-core --no-default-features --features all-compression,state_machine
# Opt into high-level Database::flush / compact
cargo build -p cqlite-core --features experimental
# Minimal build (no compression, no query engine)
cargo build -p cqlite-core --no-default-features- Cassandra 5+ SSTable format parsing (100% of test tables)
- All CQL types including collections and UDTs
- All compression codecs (LZ4, Snappy, Deflate, Zstd)
- CLI tool with REPL and one-shot query modes
- SELECT with WHERE clause (partition/clustering key equality)
- Output formats: Table, JSON, CSV
- Parquet output format with Snappy compression
- Export command (
cqlite export) - Streaming export for large datasets
- Output formats: CSV, JSON, Parquet, CQL
- Python bindings with full CQL type support
- Node.js bindings with TypeScript definitions
- Streaming API for memory-efficient queries
- pip/npm installable packages (5 platform builds each)
- Type stubs for IDE support (Python mypy, TypeScript)
- Write support: WAL + memtable + flush to Cassandra SSTables
- STCS compaction via
maintenance_step() - Write API in Python, Node.js, and CLI
- Full type coverage: Inet, Varint, Duration, Tuple, Frozen
- E2E readback gate: write → flush → Cassandra
nodetool refresh→ verify
- Embeddable Parquet writer in
cqlite-core(behind aparquetfeature) +export_parquetin Python/Node - Version-gated reads for the Cassandra 5.0
oaformat; graceful handling ofda(BTI) - Real BTI trie node-type dispatch and schema-typed query result columns
- Published documentation site at pmcfadin.github.io/cqlite
- Read-path constant-factor speedups — query-engine hot-path cleanups (schema
Arc, single projection, cached sort keys, plan cache), read-path idiom bundle, and point-read I/O via a positional-read (ReadAt) trait - Node.js bindings throughput — batch-fetch streaming rows, move (not clone) row values in
executeNative, cachedSet/Mapconstructors - Byte-bounded result budget —
Error::ResultTooLarge+QueryConfig.max_result_bytes(default 64 MiB) - Per-surface SSTable freshness contract + explicit
Database.refresh() - No-heuristics correctness — removed blob-decode byte-pattern guessing; unknown-table reads fail honestly instead of fabricating a default schema
- See CHANGELOG.md for the full per-release detail
- Byte-for-byte compaction parity vs Apache Cassandra —
cqlite compact+ a differential harness in CI, full reconciliation rule set (complex deletions, tombstone tie-breaks,gc_gracepurging, range tombstones, per-cell/dropped-column purging, non-frozen UDT multi-cell) - Arrow Flight server + Trino connector — query SSTables as a federated source with predicate, token-range, and aggregation pushdown
- Canonical BTI (
da) write + end-to-end read — emit Cassandra-format trie-indexed SSTables - CDC-style delta-scan /
delta-export— project SSTable generations to Parquet envelopes with full tombstone fidelity -
WRITETIME()/TTL()inSELECTand query-engine completeness (PER PARTITION LIMIT, static columns, clustering order/bounds, partition-targeted lookups) - crates.io OIDC trusted publishing + Homebrew tap
- See CHANGELOG.md for the full per-release detail
See the Roadmap section below for in-flight epics and milestones.
CQLite is at v0.13.0 and production-ready for the use cases above. The path to v1.0 is tracked in the open. Full detail, with milestones, lives at pmcfadin.github.io/cqlite → Roadmap.
| Workstream | Epic |
|---|---|
| Wire storage-layer capabilities (bloom/index/BTI seeks) into the CQL query path + regression guards | #951 |
| Read-path performance & I/O backend (parallel single-reader scans, io_uring spike) | #906 |
| CLI & bindings polish (DX & cleanup) | #907 |
| Compaction byte-parity follow-ups (range tombstones e2e + edge cases) | #938 |
| M6 — WASM bindings · M7 — performance validation + v1.0 | planned |
The roadmap follows real-world use. Want something prioritized? Open or 👍 an issue — and ⭐ star the repo.
CQLite is honest about its sharp edges. The current release (v0.13.0) has a few
known gaps — none of which block the core read/export workflows. Full, dated list:
pmcfadin.github.io/cqlite → Known Issues.
| Issue | Impact | Tracking |
|---|---|---|
SET<FROZEN<UDT>> fails to deserialize in the Python bindings |
Python only; CLI/Rust unaffected | #804 |
Concurrent queries on one Database can race (Column not found) |
Use one handle per thread | #805 |
Wide partitions written by CQLite scan linearly (promoted_index_length = 0) |
Perf on 10k+ rows/partition | #751, #752 |
Pre-5.0 formats (md/mc/la/ma) unsupported |
By design — Cassandra 5.0 only | Limitations |
For what CQLite does not do by design (older formats, network access, query features), see Limitations.
Found something not listed? Open an issue — a good report (Cassandra version, schema, command, output) is the most valuable contribution you can make.
Design Philosophy:
- No cluster dependency - Read and write SSTables directly, with no running Cassandra node
- CQL parser - Native CQL support using an Antlr4 grammar
- Cassandra 5+ focus - Modern 'oa' format with BTI support
- Memory efficient - <128MB usage target for large files
- Self-contained engine - Pure-Rust parsing and writing, including STCS compaction
CQLite is developed in the open as an Apache-licensed project. We welcome contributions from the Cassandra community!
CQLite uses a spec-driven, agent-orchestrated, gate-enforced workflow built on Claude Code. In short:
- Specs are the source of truth. Requirements live in a durable OpenSpec spec under
openspec/specs/; GitHub issues (epics + sub-issues) are the execution ledger, not the contract. (spec layer rolling out in the v0.13 cycle.) - A Product-Manager orchestrator (
/prioritize,/pm-status,/start-epic) plans, prioritizes, and coordinates implementer agents — one stream per issue in an isolated git worktree. - Every task passes a deterministic gate (
scripts/agent-gate.sh:cargo fmt,clippy -D warnings, tests, smoke) before it's "done" — enforced by aTaskCompletedhook, not the honor system. - The author is never the reviewer. Work is reviewed in a fresh context by roborev (a second model family) +
rust-reviewer, and audited against the spec (spec-auditor) and for meaningful coverage (coverage-reviewer). - Humans decide product, agents decide implementation. Ambiguous scope and tradeoffs are escalated on a NEEDS YOU list, never guessed.
Definition of done: gate passes · spec-auditor confirms acceptance criteria · coverage-reviewer confirms tests are meaningful · roborev is clean.
📖 Full workflow, lifecycle, and how to run it yourself: docs/development/METHODOLOGY.md
# Prerequisites
# - Rust 1.85+
# Clone and build
git clone https://github.com/pmcfadin/cqlite.git
cd cqlite
cargo build
# Fetch test data (JSONL reference files are in git, SSTable binaries fetched separately)
bash test-data/scripts/fetch-datasets.sh
# Run tests
env CQLITE_DATASETS_ROOT=$PWD/test-data/datasets cargo test --package cqlite-core- Check Issues: Look for
good-first-issuelabels - Discuss: Join our community discussions
- Code: Follow Rust best practices and include tests
- Test: Ensure compatibility with real Cassandra data
- Document: Update docs for user-facing changes
- All SSTable components parsed (Data.db, Index.db, Summary.db, Statistics.db, TOC)
- 33/33 test tables passing (100% validation)
- All 21 CQL primitive types + collections + UDTs + frozen types
- All compression algorithms working
- Tiered test coverage targets (see PRD Section 5.1)
- CLI with one-shot and REPL modes
- SELECT queries with WHERE clause support
- Multiple output formats (Table, JSON, CSV)
- Parquet output format with Snappy compression
- Export command with CSV, JSON, Parquet, CQL formats
- Streaming export for memory-efficient large dataset handling
- Progress bar and statistics for exports
- Python bindings via PyO3 with sync-first API
- Node.js bindings via napi-rs with Promise-based API
- Full CQL type system (20+ types including collections, UDTs)
- Thread-safe database handles
- 500+ tests with 98%+ pass rate across both bindings
- Write support: WAL-backed memtable + flush to portable Cassandra 5.0 SSTables
- STCS compaction (
maintenance_step()) - Write API exposed in Python (
flush_run,maintenance_step,write_stats), Node.js (flushRun,maintenanceStep,writeStats), and CLI (--writable,--write-dir,--flush,maintenance,write-stats,export-sstable) - Type roundtrips verified for all major types including Inet, Varint, Duration, Tuple, Frozen
- E2E validation against live Cassandra 5.0 (write → flush →
nodetool refresh→cqlsh)
See docs/development/PRD.md for milestone details.
- Cassandra 5.0+: 'oa' format with BTI support
- File Types: Data.db, Index.db, Summary.db, Statistics.db
- Compression: LZ4, Snappy, Deflate, Zstd
- Parse Speed: 1GB files in <10 seconds
- Memory Usage: <128MB for large SSTables
- Query Latency: Sub-millisecond partition lookups
- Python: Production-ready sync API (see Python README)
- Node.js: Production-ready Promise API (see Node.js README)
- WASM: Planned (M6+)
- Documentation site: https://pmcfadin.github.io/cqlite/ — user docs, SSTable format guide, agent integration docs
- API docs (rustdoc): latest tag · published per release tag at
https://pmcfadin.github.io/cqlite/api/<tag>/ - Changelog: CHANGELOG.md — what each tagged release contains
- Performance: Methodology, local repro, and CI gate policy
- CQL Grammar: Patrick's Antlr4 CQL Grammar
- Issues: GitHub Issues
- Discussions: GitHub Discussions
- ⭐ Star the project: github.com/pmcfadin/cqlite — the single best way to support it and shape where the time goes
- 🐛 Bugs & feature requests: GitHub Issues
- 💬 Questions & ideas: GitHub Discussions
- 🛠 Contributing: see CONTRIBUTING.md, the Roadmap, and our Code of Conduct — look for
good-first-issuelabels
CQLite is an independent open-source project, not an Apache Software Foundation project. It is built in the spirit of the Apache Cassandra community, with the goal of contributing it upstream as it matures.
Licensed under the Apache License, Version 2.0. See LICENSE for details.
Special thanks to the Apache Cassandra community and the many contributors who make projects like this possible. CQLite builds on decades of database engineering innovation from the Cassandra project.
Note: M1 through M5 milestones are complete and the project is at v0.13.0. Core SSTable reading, CLI, output writers (including Parquet), Python and Node.js bindings, and write support with STCS compaction and byte-for-byte compaction parity vs Apache Cassandra are production-ready, alongside read-path performance wins, byte-bounded result budgets, an Arrow Flight + Trino connector, canonical BTI (da) write/read, and CDC-style delta export. Next: M6 (WASM bindings) and M7 (performance validation + v1.0).
