Skip to content

v0.3.0 – Flexible JSON Payloads and Unified Index Maintenance

Choose a tag to compare

@hsballoon hsballoon released this 04 Aug 08:45
· 95 commits to master since this release

Highlights

  • Changed messages, response, chosen_trace, and rejected_trace to Arrow JSON columns, allowing provider-specific payloads without SDK-level schema validation.
  • Added deserialize_json support across the main read APIs, returning either native JSON strings or Python dict/list values.
  • Preserved best-effort blob_manifest extraction while ensuring extraction failures never block data ingestion.
  • Unified landing and serving index maintenance through maintain_table_indexes() and scripts/ops/maintain_table_indexes.py.
  • Added role-specific index maintenance for the four supported production and test tables, including missing-index creation and partition optimization.
  • Removed ineffective in-memory dirty-bucket tracking from the landing ingestion path.
  • Expanded unit and integration coverage for JSON payloads, deserialization behavior, production-style job_id values, and landing/serving read workflows.
  • Updated the English and Chinese documentation for the new schema, APIs, and operational workflow.

Breaking Changes

  • messages, response, chosen_trace, and rejected_trace must now be supplied as JSON strings when writing records.
  • These fields are returned as JSON strings by default; pass deserialize_json=True to receive Python dict/list values.
  • maintain_landing_indexes() and scripts/ops/maintain_landing_indexes.py have been replaced by the table-aware maintain_table_indexes() interface.
  • Downstream applications must update their SDK dependency to v0.3.0 before writing to tables using the new schema.