Repository navigation
Releases: AI45Lab/wt-data-platform-sdk
Release list
v0.6.4 — Adopt dldb v1.1.7 write-path optimization
Highlights
- Upgrade the SDK dependency from
dldb-v1.1.6todldb-v1.1.7. - Adopt dldb's HASH write-path optimization for
add()andupsert(). - Avoid unnecessary bucket-splitting and data copies when a batch maps to a single HASH bucket.
- Preserve the existing SDK and dldb API behavior for existing callers.
- Keep multi-bucket batch writes supported with the existing partitioning behavior.
Compatibility
- No application code changes are required for existing
add(),upsert(),update(), or query calls. - The optimization is implemented in dldb; the SDK only updates the pinned dldb dependency.
- Existing landing partial-upsert support remains available through
insert_missing=False. - Serving ETL continues to use the default complete-row upsert behavior.
Dependency
This release installs dldb from:
https://github.com/DeepLink-org/Persisting.git@dldb-v1.1.7#subdirectory=persisting/dldb
Notes
The reduced-copy path applies when all records in a batch belong to one HASH bucket. Batches spanning multiple buckets continue to be split and written per bucket as before.
v0.6.3 – Support Partial Upsert For landing Writing Path
Highlights
- Upgrade the dldb dependency from
dldb-v1.1.2todldb-v1.1.6. - Add optional dldb heartbeat configuration through
GatewayConfig. - Refine trainability detection:
- Keep Claude Code sessions in one append-only chain across
currentDaterollovers. - Normalize harness-managed current-date reminders only during fingerprint matching.
- Bump
update_is_trainablestage version from v7 to v8. - Bump
landing_enrichment_pipelineversion from v4 to v5.
- Keep Claude Code sessions in one append-only chain across
- Add ETL task discovery, enqueueing, worker execution, status tracking, warnings, and richer execution reports.
- Add support for partial landing upserts with
insert_missing=False, avoiding unnecessary reads and rewrites of wide payload columns. - Harden deletion for HASH-partitioned tables by pruning non-materialized buckets before issuing delete operations.
- Improve Arrow/Pandas columnar extraction and reduce unnecessary data conversion overhead.
- Add bucket health checks and scheduled table-maintenance utilities.
- Refine landing archive and cleanup workflows.
- Improve ETL timing, report, and operational logging.
Compatibility
- Python:
>=3.10,<3.13 - dldb:
v1.1.6 - Existing full-row landing and serving upsert behavior remains unchanged by default.
- Partial landing upsert requires dldb
v1.1.5or later.
Installation
Install from the release tag:
pip install --upgrade \
"wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.6.3"For Safactory integration, pin the dependency to:
wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.6.3
v0.6.2 – Support batch upsert landing function
Highlights
-
Added complete-row landing upsert APIs:
WTGatewayClient.upsert_landing()WTGatewayClient.upsert_landing_batch()
-
Added configurable
match_columnssupport for landing and serving upserts. -
The default upsert match key is now:
["job_id", "id"]
-
match_columnsis passed through to dldb as its upsert merge-key columns. -
HASH bucket routing remains fully managed by dldb from each record’s
job_id; callers do not need to provide a partition number. -
Added validation for:
- non-empty
job_id; - valid schema columns;
- duplicate match columns;
- duplicate match keys within one batch.
- non-empty
-
Updated English and Chinese documentation with the new upsert behavior and API signatures.
Dependency
This release uses dldb from the dldb-v1.1.2 tag.
Validation
- Unit tests:
164 passed - Real upsert smoke test passed against
v2_landing_test, including:- initial insert;
- repeated upsert with the same
job_idandid; - verification that the existing row was updated rather than duplicated;
- cleanup verification.
v0.6.1 — Upgrade dldb dependency to v1.1.2
Changes
- Upgrade the dldb dependency from
dldb-v1.1.0todldb-v1.1.2. - Include the dldb fix for Persisting Issue #101, which reuses a single catalog snapshot during cold multi-partition HASH/VALUE filters and removes the previous N+1
list_tables()behavior. - Update installation documentation and runtime compatibility notes.
Compatibility
- No wt-sdk public API changes.
- Existing dldb/LanceDB access continues to go through dldb.
- Requires the dependency versions declared by dldb-v1.1.2:
- Python 3.10–3.12
- LanceDB 0.34.x
- pylance 9.x
- pandas 2.3.x
- PyArrow 21–23
Validation
- wt-sdk dldb compatibility tests:
3 passed. - dldb catalog-listing regression tests:
2 passed. - Cold HASH filter catalog discovery was verified to use a single
list_tables()call. - A real
v2_landing_testcold query completed with one catalog listing. Full legacylanding_testvalidation remains subject to existing S3 fragment availability.
v0.6.0 – Replace Test Landing Table with v2_landing_test
Highlights
- Switched the test-profile landing table from
landing_testtov2_landing_test. - Updated active SDK scripts, ETL tools, integration tests, unit tests, and documentation to use the new test table.
- Added and verified the new
v2_landing_testwith the current schema andHASH(job_id)partitioning. - Preserved the original
landing_testtable for legacy diagnostics and dldb investigation. - Removed obsolete one-time landing migration and production smoke-test scripts.
- Production table defaults and behavior remain unchanged.
Breaking Changes
WT_SDK_PROFILE=testnow reads from and writes tov2_landing_test.- The
v2_landing_testtable must exist before running test-profile integrations or operational commands. - The old
landing_testtable is no longer the default test table, but can still be selected explicitly when needed.
v0.5.0 – Environment Isolation, dldb 1.1 Compatibility, and Operational Tooling
Highlights
- Added profile-aware environment-config isolation:
testusesenv_config_test, whileproductionandproduseevaluation_env_config; explicit table overrides remain supported. - Upgraded the dldb dependency to
dldb-v1.1.0, enabling correct unpartitioned SimpleTable reopening and exact logical-table resolution without SDK-side wrapper pinning. - Updated ETL checkpoint handling to use current dldb session APIs and removed dependencies on deleted table-opening APIs.
- Added reusable environment-config operations for table initialization, querying, cleanup, index maintenance, fragment inspection, and index-coverage reporting.
- Improved operational query performance with partition pruning, optimized counting and cleanup paths, distinct-value inspection, and HASH partition health diagnostics.
- Added controlled landing archive workflows for full-table rotation and online cold-data migration.
- Expanded serving delivery tooling with verified batch export and dataset-level counting.
- Enhanced ETL with trainability processing, FreeCoT enrichment, search-text generation, improved incremental execution, and broader integration coverage.
Upgrade Notes
- Downstream applications should update their SDK dependency to
v0.5.0. WT_SDK_PROFILEnow also selects the default environment-config table. Omitting the profile continues to default safely totest.- Existing applications that explicitly pass an environment-config table name retain their current behavior.
- Existing environments should force-reinstall dldb because the
dldb-v1.1.0source tag may still report distribution version1.0.0, which can cause pip to reuse an older installation.
v0.4.1 - Fresh Reads SNAPSHOT and More Ops Tools
Highlights
- Extended DevOps toolkit
scripts/inspect/query_data.pyto supportevaluation_env_configby table name, automatically resolvingWT_SDK_ENV_CONFIG_DB_URI. - Extended
scripts/ops/cleanup_data.pyto support filtered cleanup ofevaluation_env_config, including dry-run previews and latest-snapshot reads. - Fixed inspection/cleanup scripts to avoid partition pinning warnings for unpartitioned environment-config tables.
- Added unit coverage and a real DLDB/S3 integration test for environment-config latest-snapshot visibility.
v0.4.0 – SDK-Managed Timestamps and Serving Upsert
Highlights
- Added
source_updated_atandserving_updated_atas Unix epoch millisecond fields in the unified landing/serving schema. - Automatically initializes source timestamps and refreshes them on landing updates without changing existing caller code.
- Added native
upsert_serving()andupsert_serving_batch()APIs usingidas the business key. - Serving writes now preserve
source_updated_at, refreshserving_updated_at, and never mutate caller-owned models. - Added BTREE index definitions for the new timestamp fields.
- Expanded unit and real S3 integration coverage for timestamps, JSON payloads, landing updates, serving ingestion, and upsert behavior.
- Synchronized package metadata and runtime version reporting to
0.4.0.
Compatibility
Existing ingest_landing(), ingest_landing_batch(), query_data(), pull_data(), and update_landing() calls remain source-compatible.
update_landing() now refreshes source_updated_at by default. Pure operational updates may opt out with:
touch_source_updated_at=FalseServing upserts require a non-empty and immutable job_id. Callers must guarantee globally unique IDs because dldb HASH tables do not enforce uniqueness across buckets.
Migration
This release changes the physical LanceDB schema. Existing landing and serving tables must be rebuilt or migrated before upgrading writers to v0.4.0.
Downstream applications should update their dependency pin to:
wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.4.0
Validation
- 67 hermetic unit tests passed.
- 6 real S3 integration tests passed.
- Active landing and serving tables were rebuilt and verified with the new schema.
- Integration test data was cleaned successfully.
- Legacy archive tables were left unchanged.
v0.3.0 – Flexible JSON Payloads and Unified Index Maintenance
Highlights
- Changed
messages,response,chosen_trace, andrejected_traceto Arrow JSON columns, allowing provider-specific payloads without SDK-level schema validation. - Added
deserialize_jsonsupport across the main read APIs, returning either native JSON strings or Pythondict/listvalues. - Preserved best-effort
blob_manifestextraction while ensuring extraction failures never block data ingestion. - Unified landing and serving index maintenance through
maintain_table_indexes()andscripts/ops/maintain_table_indexes.py. - Added role-specific index maintenance for the four supported production and test tables, including missing-index creation and partition optimization.
- Removed ineffective in-memory dirty-bucket tracking from the landing ingestion path.
- Expanded unit and integration coverage for JSON payloads, deserialization behavior, production-style
job_idvalues, and landing/serving read workflows. - Updated the English and Chinese documentation for the new schema, APIs, and operational workflow.
Breaking Changes
messages,response,chosen_trace, andrejected_tracemust now be supplied as JSON strings when writing records.- These fields are returned as JSON strings by default; pass
deserialize_json=Trueto receive Pythondict/listvalues. maintain_landing_indexes()andscripts/ops/maintain_landing_indexes.pyhave been replaced by the table-awaremaintain_table_indexes()interface.- Downstream applications must update their SDK dependency to
v0.3.0before writing to tables using the new schema.
v0.2.0 – Unified Landing/Serving Schema and Read APIs
Highlights
- Unified the landing and serving schemas with
tags,search_text,chosen_trace, andrejected_trace. - Added Arrow JSON support for
meta_jsonwhile preserving JSON-string compatibility in the Python SDK. - Standardized both tables on 128-bucket
HASH(job_id)partitioning with workload-specific indexes. - Replaced
query_landing()with the table-awarequery_data(), which always returnsList[dict]and excludes null fields by default. - Replaced
fetch_data()withiter_data_batches()and added table selection to the primary read APIs. - Added
export_data_batches()for reliable offline exports using a fixed ID manifest and per-batch validation. - Preserved the existing landing defaults and DataFrame behavior of
pull_data()for SAfactory compatibility. - Updated operational scripts, documentation, and unit/integration coverage for the new schemas and APIs.
Breaking Changes
chosen_responseandrejected_responsehave been replaced bychosen_traceandrejected_trace.query_landing()has been replaced byquery_data().fetch_data()has been replaced byiter_data_batches().- Existing landing and serving tables must use the new schema and
HASH(job_id)partition layout.