Releases: BenchBox-dev/BenchBox
Releases · BenchBox-dev/BenchBox
Release list
BenchBox 0.4.0
Post-publication accounting correction: This section was reconciled on
developafter v0.4.0 was published. It records shipped changes omitted from
the immutable tag's changelog. The tag and PyPI artifacts were not modified.
Removed
- BREAKING: bare
clickhouseplatform alias removed - The temporary
compatibility alias that let--platform clickhouse(andget_adapter( "clickhouse")) resolve to a first-class platform was added in v0.2.1 as a
deprecation shim and is removed in this release (window v0.2.1 → v0.4.0).
Useclickhouse-local(embedded chDB),clickhouse-server(self-hosted), or
clickhouse-cloud(managed); the explicitclickhouse:local/
clickhouse:server/clickhouse:cloudselectors also remain available.
Passing bareclickhousenow raises aValueErrornaming these replacements
instead of silently defaulting to a deployment mode. ThechCLI shorthand
now resolves toclickhouse-local. - BREAKING:
databricks-connectinstall extra removed - Replace
benchbox[databricks-connect]withbenchbox[cloud-spark-databricks]. This
renames only the BenchBox install extra; it still installs the upstream
databricks-connectpackage. Other DataFrame install extras are unchanged.
New
- Results Explorer preview - Browse and compare published benchmark results at
benchbox.dev/results/. Explore cross-platform
leaderboards, per-query timings, hardware details, comparison tools, and SQL
queries over the public data. The preview contains a curated set of results,
not a complete or certified ranking. - DuckLake platform - Run benchmarks against DuckLake (DuckDB lakehouse
format: Parquet table data + SQL-database catalog metadata) via
--platform ducklake. The catalog backend (--platform-option catalog=duckdb|sqlite|postgres) and the Parquetdata_path(local or
s3://) can be selected independently. Four documented deployment modes have
passed TPC-H scale-factor-1 correctness validation. Existing catalogs are
reused across runs and--forcerebuilds them; requires DuckDB >= 1.3,
installable asbenchbox[ducklake]. Beta.
Added
- Result provenance and funding labels - Results Explorer now shows whether
each published run came from BenchBox maintainers, a community contributor,
or a platform vendor, together with any disclosed funding source. These labels
appear in rankings, comparisons, and result details, with an in-page
explanation of what they mean. Community submissions remain visible but are
excluded from ranked tables. - Local MCP connections over HTTP - Local MCP clients can now connect with
benchbox-mcp --transport streamable-http; existing stdio integrations
continue to work unchanged. Authenticated non-local deployments also provide
persistent benchmark jobs, but shared deployment remains deferred and is not
supported for production use in this release.
Changed
- GitHub repository moved to the
BenchBox-devorganization - BenchBox is
now hosted at github.com/BenchBox-dev/BenchBox.
Public project links, issue and release tooling, CI, and package metadata now
use the organization-owned repository. Existing Git remotes using
github.com/joeharris76/BenchBox.gitcontinue to redirect; update local
remotes to the organization URL when convenient.
Fixed
- Correct data volumes for Write and Transaction Primitives - Reusing
benchmark data created for a different scale factor could produce successful
results against the wrong amount of staged data. BenchBox now detects and
rebuilds that data automatically on the next run; no action is needed. - TPC generators recover from incompatible bundled tools - The bundled
TPC-H generators now support macOS 15. BenchBox also checks bundled TPC-H and
TPC-DS generators before use and automatically builds a compatible version
when needed. - Correct TPC throughput metrics in newly exported results - TPC-H and
TPC-DS throughput drivers now include every executed query in
Throughput@Size, correcting the 22x and 99x understatements produced by
earlier result versions. Historical result bundles remain unchanged records. - Removed dead
BENCHBOX_TUNING_ENABLEDenv var - This variable set the
tuning.enabledconfig key, which nothing at runtime ever read (only a unit
test did); docs incorrectly claimed it "activates tuned runs in CI". Use
--tuning tuned/--tuning autoonbenchbox runto actually enable
tuning;BENCHBOX_TUNING_CONFIGstill works to point at a default tuning
file. If you were settingBENCHBOX_TUNING_ENABLED, it had no effect and can
simply be removed. - Accurate
--tuning autoguidance for SQL platforms - SQL runs now explain
thatautouses a basic constraints-only configuration: primary-key,
foreign-key, unique, and check constraints are enabled, with no other tuning.
DataFrame platforms continue to use smart defaults. - DataFusion accepts Python string table paths - Programmatic DataFusion
loads now accept paths supplied as either Python strings orpathlib.Path
objects.
Security
- Expanded secret redaction in result exports and common errors - Result
exports now remove connection credentials, account identifiers, and
service-account keys. MotherDuck, DuckLake, and MCP benchmark-execution errors
also redact known secrets, including values carried by chained exceptions.
Some other MCP errors can still include backend-provided exception text, so
avoid credentials in values that a backend might repeat in an error.
BenchBox 0.3.1
What's Changed
- Release v0.3.0 pandas runtime dependency repair by @joeharris76 in #473
- Release v0.3.0 prompts landing CSS hotfix by @joeharris76 in #484
- ci: relocate dependency audit by @joeharris76 in #482
- test: lock dependency bounds parser behavior by @joeharris76 in #480
- test: enforce import dependency contract by @joeharris76 in #481
- ci: add hermetic wheel smoke gates by @joeharris76 in #479
- Keep platform imports lazy by @joeharris76 in #485
- [codex] Harden release pipeline follow-ups by @joeharris76 in #486
- Release v0.3.1 by @joeharris76 in #1072
Full Changelog: v0.3.0...v0.3.1
BenchBox 0.3.0
What's Changed
- Release v0.3.0 by @joeharris76 in #466
- Release v0.3.0 recovery by @joeharris76 in #472
Full Changelog: v0.2.1...v0.3.0
BenchBox 0.2.1
Full Changelog: v0.2.0...v0.2.1
JoinOrder canonical IMDb 2013 dataset (v1)
JoinOrder canonical IMDb 2013 dataset (v1)
- Source: Harvard Dataverse DOI 10.7910/DVN/2QYZBT, file
imdb_pg11. - Snapshot: May 2013 IMDb list-file data parsed into the JOB 21-table schema.
- Attribution: Information courtesy of IMDb (https://www.imdb.com). Used with permission.
- Scope: research and database optimizer benchmarking.
- BenchBox redistribution disclaimer and takedown commitment are included in DATA-LICENSE.md.
- Archive:
joinorder-imdb-2013-v1.tar.zstsha256106ae78c5a1ab09f8e5ccc9ce0751ff8b5eae1f7b9e03a262773c26c528e9526.
Tables
aka_name: 901,343 rows, sha2562abc121754c0ee62aac325dcc5373fcf054ba0eb9740d65001f31daceb07169eaka_title: 361,472 rows, sha256d0f6a559ceeaac996191656d97a6ce9477d515a382aec96948e8e207210adde7cast_info: 36,244,344 rows, sha256c6bebb9d00b50a7f5d6a046763c8d5e725dc55237b9ad94f92288aa2e9e3cbaechar_name: 3,140,339 rows, sha2568427d9f1c7376df4a3f4175fed83b016277d25b835985fe9178a088c66379b25comp_cast_type: 4 rows, sha256f78755ff054ef0d5e3e00d245c7f3d9335c2d031eccbef0428c414c91058e6afcompany_name: 234,997 rows, sha2567210730eb45239af2f6c45b3bc229f0e21f890b11e8101a32fe7ad8bec478ffacompany_type: 4 rows, sha256bba80a9acac32fe57344fbea5904a0d6a23a977c00b1916a83a5be94594c1fe8complete_cast: 135,086 rows, sha2566c660c5665348174d9226ad6dfd118728c2352ba68a1c4d45c93de945f7bccc5info_type: 113 rows, sha256e83fa2db221564ffb3487c16614200899a042a9e66792326a74179523cefc907keyword: 134,170 rows, sha256fd61a10ba98c9ded14778649bf64e1c98defbe9b78d481b6d1fe51dc444ce1bfkind_type: 7 rows, sha256dd9d886414b4ee83b0b015e06a440f868a05bb0d590d64cc7bfaadec1021bde9link_type: 18 rows, sha256205d37c045011d4fc20cf566d787e7b9655ecdad7bb0d0e19befb9d32f0036b8movie_companies: 2,609,129 rows, sha2566d7939c6ebbb88bd6963b155485a48060df4b073099f3e86835a68f282b2c94dmovie_info: 14,835,720 rows, sha256c63b13e8f90965cdfb45263d2fb8c6b09dd7926a6c2e9d957508c927133be54fmovie_info_idx: 1,380,035 rows, sha25652458eba771af35067459e79988ad91d639b5dcb535e116854de214aa06ee159movie_keyword: 4,523,930 rows, sha25617e628c6a2214d265ed0ea67c2462034a198a0b91c42356ce94a04e44f7b7a64movie_link: 29,997 rows, sha25671208862a69db18e79bf414a9f2787dbf640a4a3558db165ca536c57eb746991name: 4,167,491 rows, sha2569c755fb800bbac00e6fa4c09d7bd344a759ff0d65ce8396c117fd012221f49dcperson_info: 2,963,664 rows, sha256a516f55bade10bf660eb40f159aa93b6c579349646d1c927110f72cfe78beac3role_type: 12 rows, sha2567f6a2148a3e34787dfe1c6f71f68e9d1cd0e3fb165e497bdf9a7c64e21d771cetitle: 2,528,312 rows, sha2566c5749b8c1569342c1a96ffaefe1146051a10f77db3b693df84a0d2d44f1462b
BenchBox 0.2.0
Full Changelog: v0.1.5...v0.2.0
BenchBox 0.1.5
Added
- Textcharts standalone library - Extracted all 15 ASCII chart types and base rendering
primitives into an independenttextchartspackage underpackages/textcharts/. The library
has its ownpyproject.toml, README with chart gallery and API reference, and zero BenchBox
dependencies. Clean standalone names (BarChart,Histogram,Heatmap, etc.) are exported
alongside BenchBox-compatible aliases. BenchBox now depends on textcharts as a path dependency
with compatibility shims preserving existing import paths. - Open table format loading - Added runtime loading support for external table formats
(Delta Lake, Iceberg, Hudi) through adapter-levelload_tableimplementations for Spark
mixin platforms, cloud SQL platforms, and Snowflake/ClickHouse adapters. Format support is
gated on adapter configuration so it is only available on platforms that implement it. - Expanded format capability registry - Registered format capabilities for Hudi,
Presto/Trino, Snowflake, ClickHouse, Redshift, BigQuery, and Spark-based platforms including
cloud lakehouse variants (EMR, Dataproc, Glue, Fabric Spark, Synapse Spark, Dataproc
Serverless). Removed registrations for platforms without actual loading code (including
LakeSail for delta/iceberg/hudi). - Mutation testing - Added
mutmutmutation testing targeting 5 critical modules
(duckdb.py,adapter.py,runner.py,chart_generator.py,run.py) with a
make mutation-testMakefile target for manual quality reviews.
Fixed
- Format capability registry accuracy - Normalized platform display names to match registry
keys and removed platforms from format registrations where adapter code has no actual loading
implementation. - CLI recursive import - Fixed a circular lazy-import in the benchmarks module that caused a
RecursionErroron CLI startup. - CoffeeShop SA2 query - Corrected
group_bycolumn name from'name'to
'product_name'. - Textcharts API migration - Migrated to textcharts v0.1.2 API after breaking changes,
renamedASCII*-prefixed classes across 10 source and test files, removed 3 unused deprecated
factory imports, and regenerated golden snapshots for neutralized defaults. - pytest-xdist worker title patch - Tightened the xdist worker title monkeypatch to prevent
test pollution across parallel workers. - Comprehensive Windows CI compatibility - Fixed 80+ Windows test failures spanning path
separators (.as_posix()for forward-slash comparison), file encoding (encoding="utf-8"
forwrite_text()), numpy int32 overflow on 64-bit multiplication, Rich Console width on
headless CI, Python ABI tag format differences (.pydvs.so),Path.touch()vs
time.time()mocking, NTFS directoryst_sizereturning 0, Windows CWD locks preventing
temp directory cleanup, andshutil.copytreereplacing symlinks for TPC-DS template setup. - TPC-DS dsqgen Windows option prefix - Fixed dsqgen invocation on Windows where the binary
expects/option prefixes instead of-(OPTION_STARTinr_params.c), and switched to
relative paths to stay under dsqgen's 80-charPARAM_MAX_LENbuffer. - Missing
tpcds.idxdistribution file - Added the required TPC-DS distribution index file
to Windows binary packages (both x86_64 and ARM64). - Throughput test timer resolution - Used
time.perf_counter()for throughput duration
calculation to avoid zero-duration results from low-resolutiontime.time()on Windows.
Changed
- Test suite quality overhaul - Deleted 13 hollow coverage-theater test files and replaced
them with behavior-verifying tests for DuckDB, SQLite, and DataFusion adapters. Replaced mock
credential tests with real file-based tests. Strengthened 150is-not-Noneassertions across
5 test files, replaced hollowisinstanceassertions with behavioral checks, and swapped
MagicMockforSimpleNamespaceon attribute-only objects. Removed per-file coverage
enforcement in favor of a suite-wide 60% threshold. - ~316 rendering tests migrated to textcharts - Pure chart-rendering tests moved from
BenchBox's test suite to the standalone textcharts library, with shim import smoke tests
retained in BenchBox to verify re-export paths. - Pytest lane restructure - Converted test lanes from implicit timing heuristics to explicit
source markers with measured-timing-based rebucketing. Restored a lightweight fast lane,
serialized stress tests, re-laned cloud adapter tests toslow+cloud_import, and documented
pytest-xdist safety requirements. - Chart subtitle simplified - Migrated chart subtitle storage from a metadata dict to a
plain string, removing an unnecessary layer of indirection. - Verbose logging extracted - Moved verbose logging configuration from
run.pyinto a
dedicatedcli/verbose_logging.pymodule. - Visualization constants - Extracted magic numbers into named constants across
visualization modules.
Full Changelog: https://github.com/joeharris76/BenchBox/blob/main/CHANGELOG.md#015---2026-03-10
BenchBox 0.1.4
Added
power_barchart type - Added a horizontal bar chart for TPC Power@Size comparisons.
Higher values are treated as better (opposite ofperformance_bar), powered by
summary.tpc_metrics.power_at_sizeand exposed inNormalizedResult.power_bartemplate coverage - Added toflagship,head_to_head,trends,
regression_triage, andexecutive_summary. The chart renders only when TPC metric data is
present and is skipped for non-TPC runs.- Driver-version-aware chart labeling - Multi-platform chart series labels and run summaries
now include driver version context so version comparisons stay explicit in rendered output. - Runtime ABI validation for isolated drivers - Added ABI compatibility checks to isolated
runtime discovery so driver auto-install paths fail fast with actionable validation errors
instead of late runtime crashes. - Presorted data-generation modes for table formats - Added
parquet-sortedoutput mode,
plusdelta-sortedandiceberg-sortedorganization paths with clustering primitives
(z-order, Hilbert, partition-aware sorting) andcluster-bytuning integration.
Fixed
- Query plan capture correctness and persistence - Fixed multiple plan-capture defects:
forwardingcapture_plansthroughRunConfig, DuckDB JSON plan parsing edge cases,
preservation ofquery_planthrough normalization, andshow-plan/compare-plans
loading through the standard result-file path. - SSB dot-notation query IDs -
--queriesnow accepts IDs likeQ2.1, and plan-oriented
CLI flows preserve dotted IDs instead of normalizing them away. - Result timing pipeline accuracy - Fixed datagen/load timing propagation end-to-end,
including per-table load timings intable_statistics, corrected load-phase duration keying,
datagen phase duration and manifest stats in metadata, and explicit total duration override
propagation in result builders. Data-only runs now correctly execute generation, and
force_regenerateis forwarded through CLI and runner paths. - ASCII visualization readability under skewed data - Fixed outlier handling across chart
types (bar, histogram, stacked, scatter, line, CDF, percentile ladder, heatmap), addressed
zero-heavy fallback truncation edge cases, improved natural query sorting and color cycling,
and raised effective render width cap from 120 to 400 characters. --quietoutput contract for automation - Quiet mode now emits only the bare result
filepath to stdout, removing decorative output that broke script parsing.- Runtime environment stability - Fixed interpreter targeting for driver auto-install,
correctedauto_install_usedstate propagation, and resolved SIGSEGV-class failures when
driver_auto_install=truereused an already-matching version. - Additional correctness fixes - Restored
ai_primitivesregistry resolution fallback,
corrected SQLiteforce_recreateoption handling, fixed SSB customer row-count expectation in
SSBRowCountStrategy, and resolved visualize command crashes / multi-series rendering issues.
Changed
- Plan-capture default now uses actual execution timing -
--capture-plansnow defaults to
EXPLAIN (ANALYZE, FORMAT JSON)behavior viaanalyze_plans=True, recording measured timing
in captured plans. Users can opt out withanalyze_plans: falsefor estimate-only capture. - Benchmark runtime/result internals harmonized - Refactored enhanced result construction to
use a shared factory path and aligned canonical runtime behavior for benchmarks like NYC Taxi
and TSBS DevOps. make test-allresource policy and parallelism - Resource-heavy tests are now serialized
to prevent machine stalls, while slow/performance suites are moved to a dedicated stress lane.
The test suite also replaces fixed sleeps with bounded polling, reduces fixture/harness
duplication, and shifts selected CLI/e2e coverage to in-process runners for faster execution.- CI quality gates tightened - Added required table-format integration coverage and promoted
doc checks (linkcheck, example validation, docstring coverage) plus security audit policy
controls to blocking CI behavior.
Full Changelog: https://github.com/joeharris76/BenchBox/blob/main/CHANGELOG.md#014---2026-03-03
BenchBox 0.1.3
Added
- Driver version pinning —
--platform-option driver_version=X.Y.Zpins any platform's Python driver; pair withdriver_auto_install=trueto have BenchBox install it automatically viauv. Active driver version is shown in the run announcement line. - Bulk multi-shard table loading — new
load_table_bulk()interface onFileFormatHandleringests multi-shard tables in a single native call. DuckDB (CSV, Parquet) and ClickHouse Native are the first implementations; TPC-DS sharded runs are measurably faster. - Greyscale / no-color ASCII chart fallbacks — all seven ASCII chart types now use fill-pattern and glyph differentiation when color is unavailable (CI logs,
NO_COLOR, piped output). - Five new ASCII chart types — percentile ladder, stacked bar, sparkline table, CDF, rank table, and normalized speedup (log₂-scaled). All registered in the chart registry and accessible via CLI and MCP.
- Post-run summary charts — charts are automatically generated and displayed in the terminal after every benchmark run and included in MCP
run_benchmarkresponses. - Three new chart template bundles —
latency_deep_dive,regression_triage, andexecutive_summary. fabric-dwas a preferred CLI alias for thefabric_dwplatform.
Fixed
- Driver auto-install version switching — stale
sys.modulesand metadata caching could return the wrong driver version afterdriver_auto_installswapped versions; module cache is now invalidated on switch. - DataFrame cache path mismatch — DataFrame and SQL modes now share a flat directory layout, eliminating redundant data generation when switching modes on the same scale factor.
- ClickHouse zstd double-decompression —
ClickHouseNativeHandlerwas applying manual decompression on top of the driver's built-in decompression, corrupting data for compressed bulk loads. - Platform display names — corrected for Amazon Athena, Google Cloud Dataproc, Microsoft Azure platforms, and Databricks SQL.
- CLI warning when a platform option's default value is not in the declared choices list.
- Ranking normalization crash when all metric values are negative finite numbers.
- PySpark SIGINT handler hanging
pytest-xdistworkers. --validation-modeCLI prompt crash whenspec.defaultis not a string.
Changed
- Four platform drivers moved to optional extras — DuckDB (
benchbox[duckdb]), Polars (benchbox[polars]), ClickHouse Connect (benchbox[clickhouse-connect]), and psycopg2 (benchbox[postgresql]) are no longer hard dependencies. Usepip install benchbox[all]to restore the previous behaviour. - All user-facing terminal output in the run pipeline flows through
emit(), making--quietsuppression and output capture in tests consistent. BaseQueryCatalogMixinandTranslatableQueryMixinextracted from duplicate query-catalog implementations.
Full Changelog: https://github.com/joeharris76/BenchBox/blob/main/CHANGELOG.md#0130---2026-02-23
BenchBox 0.1.2
Added
- DataFrame mode for all benchmarks — complete DataFrame query implementations across all 18 benchmarks including TPC-DS (102 queries), TPC-H (22 queries), SSB, ClickBench, NYC Taxi, TSBS DevOps, H2ODB, AMPLab, CoffeeShop, TPC-H Skew, and Data Vault. DataFrame platforms: Polars, DuckDB, DataFusion, PySpark, Pandas, Modin, Dask, and cuDF (GPU).
- ASCII chart visualizations — terminal-native ASCII rendering replacing Plotly HTML charts. Seven chart types: performance bar, distribution box, query heatmap, comparison bar, diverging bar, summary box, and query latency histogram. ANSI colors, Unicode box-drawing, and best/worst highlighting.
- 14 new SQL platform adapters — PostgreSQL, Trino, PrestoDB, Apache Spark, AWS Athena, Azure Synapse, Microsoft Fabric, Firebolt, MotherDuck, InfluxDB 3.x, TimescaleDB, ClickHouse Cloud, Onehouse Quanton, and managed Spark variants (EMR, Dataproc, Glue, Fabric Spark, Synapse Spark, Dataproc Serverless).
- Open table format support — Delta Lake, Apache Iceberg, Apache Hudi, DuckLake, and Vortex columnar format with format conversion orchestration and manifest v2 for multi-format tracking.
- Physical tuning DDL generation — platform-specific DDL generators for DuckDB, Snowflake, Redshift, BigQuery, ClickHouse, Firebolt, PostgreSQL, TimescaleDB, Trino/Presto/Athena, and Spark family with sort keys, partitioning, clustering, and compression.
- Query plan capture and comparison — plan parsers for DuckDB, PostgreSQL, Redshift, DataFusion, and SQLite. Comparison engine with regression detection, fingerprinting, historical tracking with flapping detection, and CLI visualization.
- Interactive CLI wizard — guided benchmark configuration with platform selection, tuning wizard, scale factor validation, phase/query selection, onboarding, and persistent preferences.
- TPC-DI benchmark — complete implementation across 4 phases: core schema, query suite, ETL pipeline, and validation/testing.
- Cross-platform comparison engine —
benchbox comparecommand with multi-platform analysis, SQL vs DataFrame comparison, and unified visualization. - Unified tuning configuration — YAML-based tuning system with per-platform DDL generation, write-time physical layout configuration, and dry-run preview.
- Cloud storage and deployment modes for S3/GCS/ADLS/DBFS with credential setup wizard and cost estimation.
- TPC compliance improvements: stream-aware validation, query permutations, warmup/measurement iterations, maintenance operations (RF1/RF2), and
--seedfor reproducibility. - New benchmarks: AI/ML Primitives, Metadata Primitives, Write Primitives, Transaction Primitives.
- MCP server:
suggest_chartsandgenerate_charttools;--queriesand--validation-modeCLI flags; tiered--help.
Fixed
- TPC-DS data generation reliability — segfaults with fractional scale factors, parallel generation errors, streaming compression, and chunked file handling.
- Cloud platform stability — credential refresh errors, schema creation ordering, UC Volume uploads, S3 key handling, and BigQuery/Snowflake/Redshift/Databricks adapter issues.
- Type safety — 150+ type errors resolved across production code with proper annotations and TYPE_CHECKING imports.
- SQL dialect translation — SQLGlot compatibility for DuckDB, ClickHouse, DataFusion, and Netezza; reserved keyword quoting and identifier case sensitivity.
- Security hardening: SQL injection prevention, parameterized queries, path traversal protection.
- CLI hanging in non-interactive mode, progress display precision,
--quietmode propagation. - TPC compliance: correct stream permutations, maintenance phase SQL execution, Power@Size calculation parity between SQL and DataFrame modes.
Changed
- Dropped Plotly HTML charts in favor of ASCII-only rendering.
- Lazy-load cloud platform adapters to speed up CLI startup and test suite.
- Optimized TPC-DS smoke tests with selective table generation.
Full Changelog: https://github.com/joeharris76/BenchBox/blob/main/CHANGELOG.md#012---2026-02-09