Skip to content

Releases: AI45Lab/wt-data-platform-sdk

v0.6.4 — Adopt dldb v1.1.7 write-path optimization

Choose a tag to compare

@hsballoon hsballoon released this 20 Sep 12:00

Highlights

  • Upgrade the SDK dependency from dldb-v1.1.6 to dldb-v1.1.7.
  • Adopt dldb's HASH write-path optimization for add() and upsert().
  • Avoid unnecessary bucket-splitting and data copies when a batch maps to a single HASH bucket.
  • Preserve the existing SDK and dldb API behavior for existing callers.
  • Keep multi-bucket batch writes supported with the existing partitioning behavior.

Compatibility

  • No application code changes are required for existing add(), upsert(), update(), or query calls.
  • The optimization is implemented in dldb; the SDK only updates the pinned dldb dependency.
  • Existing landing partial-upsert support remains available through insert_missing=False.
  • Serving ETL continues to use the default complete-row upsert behavior.

Dependency

This release installs dldb from:

https://github.com/DeepLink-org/Persisting.git@dldb-v1.1.7#subdirectory=persisting/dldb

Notes

The reduced-copy path applies when all records in a batch belong to one HASH bucket. Batches spanning multiple buckets continue to be split and written per bucket as before.

v0.6.3 – Support Partial Upsert For landing Writing Path

Choose a tag to compare

@hsballoon hsballoon released this 20 Sep 10:13

Highlights

  • Upgrade the dldb dependency from dldb-v1.1.2 to dldb-v1.1.6.
  • Add optional dldb heartbeat configuration through GatewayConfig.
  • Refine trainability detection:
    • Keep Claude Code sessions in one append-only chain across currentDate rollovers.
    • Normalize harness-managed current-date reminders only during fingerprint matching.
    • Bump update_is_trainable stage version from v7 to v8.
    • Bump landing_enrichment_pipeline version from v4 to v5.
  • Add ETL task discovery, enqueueing, worker execution, status tracking, warnings, and richer execution reports.
  • Add support for partial landing upserts with insert_missing=False, avoiding unnecessary reads and rewrites of wide payload columns.
  • Harden deletion for HASH-partitioned tables by pruning non-materialized buckets before issuing delete operations.
  • Improve Arrow/Pandas columnar extraction and reduce unnecessary data conversion overhead.
  • Add bucket health checks and scheduled table-maintenance utilities.
  • Refine landing archive and cleanup workflows.
  • Improve ETL timing, report, and operational logging.

Compatibility

  • Python: >=3.10,<3.13
  • dldb: v1.1.6
  • Existing full-row landing and serving upsert behavior remains unchanged by default.
  • Partial landing upsert requires dldb v1.1.5 or later.

Installation

Install from the release tag:

pip install --upgrade \
  "wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.6.3"

For Safactory integration, pin the dependency to:

wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.6.3

v0.6.2 – Support batch upsert landing function

Choose a tag to compare

@hsballoon hsballoon released this 11 Sep 09:05

Highlights

  • Added complete-row landing upsert APIs:

    • WTGatewayClient.upsert_landing()
    • WTGatewayClient.upsert_landing_batch()
  • Added configurable match_columns support for landing and serving upserts.

  • The default upsert match key is now:

    ["job_id", "id"]
  • match_columns is passed through to dldb as its upsert merge-key columns.

  • HASH bucket routing remains fully managed by dldb from each record’s job_id; callers do not need to provide a partition number.

  • Added validation for:

    • non-empty job_id;
    • valid schema columns;
    • duplicate match columns;
    • duplicate match keys within one batch.
  • Updated English and Chinese documentation with the new upsert behavior and API signatures.

Dependency

This release uses dldb from the dldb-v1.1.2 tag.

Validation

  • Unit tests: 164 passed
  • Real upsert smoke test passed against v2_landing_test, including:
    • initial insert;
    • repeated upsert with the same job_id and id;
    • verification that the existing row was updated rather than duplicated;
    • cleanup verification.

v0.6.1 — Upgrade dldb dependency to v1.1.2

Choose a tag to compare

@hsballoon hsballoon released this 04 Sep 08:27

Changes

  • Upgrade the dldb dependency from dldb-v1.1.0 to dldb-v1.1.2.
  • Include the dldb fix for Persisting Issue #101, which reuses a single catalog snapshot during cold multi-partition HASH/VALUE filters and removes the previous N+1 list_tables() behavior.
  • Update installation documentation and runtime compatibility notes.

Compatibility

  • No wt-sdk public API changes.
  • Existing dldb/LanceDB access continues to go through dldb.
  • Requires the dependency versions declared by dldb-v1.1.2:
    • Python 3.10–3.12
    • LanceDB 0.34.x
    • pylance 9.x
    • pandas 2.3.x
    • PyArrow 21–23

Validation

  • wt-sdk dldb compatibility tests: 3 passed.
  • dldb catalog-listing regression tests: 2 passed.
  • Cold HASH filter catalog discovery was verified to use a single list_tables() call.
  • A real v2_landing_test cold query completed with one catalog listing. Full legacy landing_test validation remains subject to existing S3 fragment availability.

v0.6.0 – Replace Test Landing Table with v2_landing_test

Choose a tag to compare

@hsballoon hsballoon released this 03 Sep 09:27

Highlights

  • Switched the test-profile landing table from landing_test to v2_landing_test.
  • Updated active SDK scripts, ETL tools, integration tests, unit tests, and documentation to use the new test table.
  • Added and verified the new v2_landing_test with the current schema and HASH(job_id) partitioning.
  • Preserved the original landing_test table for legacy diagnostics and dldb investigation.
  • Removed obsolete one-time landing migration and production smoke-test scripts.
  • Production table defaults and behavior remain unchanged.

Breaking Changes

  • WT_SDK_PROFILE=test now reads from and writes to v2_landing_test.
  • The v2_landing_test table must exist before running test-profile integrations or operational commands.
  • The old landing_test table is no longer the default test table, but can still be selected explicitly when needed.

v0.5.0 – Environment Isolation, dldb 1.1 Compatibility, and Operational Tooling

Choose a tag to compare

@hsballoon hsballoon released this 01 Sep 05:52

Highlights

  • Added profile-aware environment-config isolation: test uses env_config_test, while production and prod use evaluation_env_config; explicit table overrides remain supported.
  • Upgraded the dldb dependency to dldb-v1.1.0, enabling correct unpartitioned SimpleTable reopening and exact logical-table resolution without SDK-side wrapper pinning.
  • Updated ETL checkpoint handling to use current dldb session APIs and removed dependencies on deleted table-opening APIs.
  • Added reusable environment-config operations for table initialization, querying, cleanup, index maintenance, fragment inspection, and index-coverage reporting.
  • Improved operational query performance with partition pruning, optimized counting and cleanup paths, distinct-value inspection, and HASH partition health diagnostics.
  • Added controlled landing archive workflows for full-table rotation and online cold-data migration.
  • Expanded serving delivery tooling with verified batch export and dataset-level counting.
  • Enhanced ETL with trainability processing, FreeCoT enrichment, search-text generation, improved incremental execution, and broader integration coverage.

Upgrade Notes

  • Downstream applications should update their SDK dependency to v0.5.0.
  • WT_SDK_PROFILE now also selects the default environment-config table. Omitting the profile continues to default safely to test.
  • Existing applications that explicitly pass an environment-config table name retain their current behavior.
  • Existing environments should force-reinstall dldb because the dldb-v1.1.0 source tag may still report distribution version 1.0.0, which can cause pip to reuse an older installation.

v0.4.1 - Fresh Reads SNAPSHOT and More Ops Tools

Choose a tag to compare

@hsballoon hsballoon released this 07 Aug 04:37

Highlights

  • Extended DevOps toolkit scripts/inspect/query_data.py to support evaluation_env_config by table name, automatically resolving WT_SDK_ENV_CONFIG_DB_URI.
  • Extended scripts/ops/cleanup_data.py to support filtered cleanup of evaluation_env_config, including dry-run previews and latest-snapshot reads.
  • Fixed inspection/cleanup scripts to avoid partition pinning warnings for unpartitioned environment-config tables.
  • Added unit coverage and a real DLDB/S3 integration test for environment-config latest-snapshot visibility.

v0.4.0 – SDK-Managed Timestamps and Serving Upsert

Choose a tag to compare

@hsballoon hsballoon released this 05 Aug 06:39

Highlights

  • Added source_updated_at and serving_updated_at as Unix epoch millisecond fields in the unified landing/serving schema.
  • Automatically initializes source timestamps and refreshes them on landing updates without changing existing caller code.
  • Added native upsert_serving() and upsert_serving_batch() APIs using id as the business key.
  • Serving writes now preserve source_updated_at, refresh serving_updated_at, and never mutate caller-owned models.
  • Added BTREE index definitions for the new timestamp fields.
  • Expanded unit and real S3 integration coverage for timestamps, JSON payloads, landing updates, serving ingestion, and upsert behavior.
  • Synchronized package metadata and runtime version reporting to 0.4.0.

Compatibility

Existing ingest_landing(), ingest_landing_batch(), query_data(), pull_data(), and update_landing() calls remain source-compatible.

update_landing() now refreshes source_updated_at by default. Pure operational updates may opt out with:

touch_source_updated_at=False

Serving upserts require a non-empty and immutable job_id. Callers must guarantee globally unique IDs because dldb HASH tables do not enforce uniqueness across buckets.

Migration

This release changes the physical LanceDB schema. Existing landing and serving tables must be rebuilt or migrated before upgrading writers to v0.4.0.

Downstream applications should update their dependency pin to:

wt-data-platform-sdk @ git+https://github.com/AI45Lab/wt-data-platform-sdk.git@v0.4.0

Validation

  • 67 hermetic unit tests passed.
  • 6 real S3 integration tests passed.
  • Active landing and serving tables were rebuilt and verified with the new schema.
  • Integration test data was cleaned successfully.
  • Legacy archive tables were left unchanged.

v0.3.0 – Flexible JSON Payloads and Unified Index Maintenance

Choose a tag to compare

@hsballoon hsballoon released this 04 Aug 08:45

Highlights

  • Changed messages, response, chosen_trace, and rejected_trace to Arrow JSON columns, allowing provider-specific payloads without SDK-level schema validation.
  • Added deserialize_json support across the main read APIs, returning either native JSON strings or Python dict/list values.
  • Preserved best-effort blob_manifest extraction while ensuring extraction failures never block data ingestion.
  • Unified landing and serving index maintenance through maintain_table_indexes() and scripts/ops/maintain_table_indexes.py.
  • Added role-specific index maintenance for the four supported production and test tables, including missing-index creation and partition optimization.
  • Removed ineffective in-memory dirty-bucket tracking from the landing ingestion path.
  • Expanded unit and integration coverage for JSON payloads, deserialization behavior, production-style job_id values, and landing/serving read workflows.
  • Updated the English and Chinese documentation for the new schema, APIs, and operational workflow.

Breaking Changes

  • messages, response, chosen_trace, and rejected_trace must now be supplied as JSON strings when writing records.
  • These fields are returned as JSON strings by default; pass deserialize_json=True to receive Python dict/list values.
  • maintain_landing_indexes() and scripts/ops/maintain_landing_indexes.py have been replaced by the table-aware maintain_table_indexes() interface.
  • Downstream applications must update their SDK dependency to v0.3.0 before writing to tables using the new schema.

v0.2.0 – Unified Landing/Serving Schema and Read APIs

Choose a tag to compare

@hsballoon hsballoon released this 31 Jul 15:08

Highlights

  • Unified the landing and serving schemas with tags, search_text, chosen_trace, and rejected_trace.
  • Added Arrow JSON support for meta_json while preserving JSON-string compatibility in the Python SDK.
  • Standardized both tables on 128-bucket HASH(job_id) partitioning with workload-specific indexes.
  • Replaced query_landing() with the table-aware query_data(), which always returns List[dict] and excludes null fields by default.
  • Replaced fetch_data() with iter_data_batches() and added table selection to the primary read APIs.
  • Added export_data_batches() for reliable offline exports using a fixed ID manifest and per-batch validation.
  • Preserved the existing landing defaults and DataFrame behavior of pull_data() for SAfactory compatibility.
  • Updated operational scripts, documentation, and unit/integration coverage for the new schemas and APIs.

Breaking Changes

  • chosen_response and rejected_response have been replaced by chosen_trace and rejected_trace.
  • query_landing() has been replaced by query_data().
  • fetch_data() has been replaced by iter_data_batches().
  • Existing landing and serving tables must use the new schema and HASH(job_id) partition layout.