Skip to content

Releases: Mullassery/StatGuardian

v2.5.0

Choose a tag to compare

@Mullassery Mullassery released this 23 Aug 08:54

Fixes

  • Streaming was fake. StreamingBatcher called a full-file parse inside next_batch(), so pulling N batches re-read and re-parsed the whole file N times with no bound on per-batch memory. Now opens the source once via Polars' mmap-backed batched readers and advances a cursor. New tests/test_streaming.rs proves output parity with whole-file execution, bounded batch size, and cost that doesn't scale with batch count.
  • Seasonal anomaly detection had a timestamp bug. AdaptiveThresholdDetector.detect_seasonal_anomaly derived each history entry's day-of-week from its list index relative to now, silently wrong unless observations arrived exactly one per day, in order, ending today. Now uses each observation's actual recorded timestamp.
  • execute_sql()/execute_spark()/execute_cloud() import fix (carried from 2.4.0) — they referenced a nonexistent statguardian._statguard module instead of statguardian._statguardian.
  • polars/pyarrow compatibility. The polars pin is now >=1.6,<1.32.3 (old ==0.19.12 predated pl.String; 1.32.3+ breaks the compiled extension's pyo3-polars binding). pyarrow>=16.0 is now a base dependency — every execute() call needs it, not just the pandas/spark extras, and the old pyarrow==14.0.1 pin can't decode modern polars' string layout. Verified with clean pip install statguardian and pip install "statguardian[pandas]".

New

  • statguardian.merge_violations(report, extra_violations) -> MergedReport — combine a ValidationReport with run_custom_validators() output into one pass/fail result. Previously referenced in a docstring but didn't exist.
  • AdaptiveThresholdDetector.detect_hourly_anomaly — business-hours-vs-off-hours comparison, previously promised by the class docstring but not implemented.

Removed

  • cli_workflow.py, server_workflow.py — dead code, never imported anywhere.

v2.4.0

Choose a tag to compare

@Mullassery Mullassery released this 23 Aug 08:06

dbt integration

  • New installable dbt package: integrations/dbt-statguardian/ — captures dbt run/test results via on-run-start/on-run-end hooks, exposes a generic statguardian.contract_passed test and a statguardian.contract_recently_validated freshness test, plus reporting models.
  • New statguardian dbt validate CLI subcommand — runs .sg contracts against your dbt models' real warehouse tables (Postgres, Redshift, Snowflake, BigQuery, DuckDB) and writes results back for the dbt test to pick up.
  • Docs: docs/DBT_INTEGRATION.md, runnable end-to-end example in integrations/dbt-statguardian/integration_tests/ (DuckDB).
  • Install with pip install "statguardian[dbt]".

Fixes

  • execute_sql(), execute_spark(), and execute_cloud() were broken for everyone — they imported from a nonexistent statguardian._statguard module (typo) instead of the real compiled statguardian._statguardian. Fixed.

v2.3.2

Choose a tag to compare

@Mullassery Mullassery released this 17 Aug 16:47

Security remediation + build fix

  • Wired the previously-unused DSL validator into the CLI (statguardian check/validate) to reject oversized/malformed contracts before they reach the parser
  • Fixed silent exception swallowing in SQL connector fallback chain and OKF contract history parsing (now logged instead of discarded)
  • Added gitleaks secrets scanning and cargo audit to CI
  • Closed out the remaining items in docs/SECURITY_AUDIT.md — see that file for the full remediation status
  • Fixed a pre-existing Rust compile break in statguardian-io/src/sql.rs (stray variable rename, invalid lifetime, missing sqlx trait bounds) that was silently failing cargo build --all-features

v0.1.0 - Initial Release

Choose a tag to compare

@Mullassery Mullassery released this 28 Jul 02:04

Initial stable release of Data quality framework

Statguardian v1.0.0: Production-Ready Data Quality Engine

Choose a tag to compare

@Mullassery Mullassery released this 07 Jul 16:56

Statguardian v1.0.0 🚀 Production Beta-Ready

Production Score: 7.5/10 | 4-week hardening complete

What's New

  • ✅ Parser input limits (10MB, ReDoS prevention)
  • ✅ 50% test coverage (65+ tests, up from 2%)
  • ✅ Streaming window support (Tumbling/Sliding/Session)
  • ✅ Statistical drift detection

Features

  • DSL parsing for data contracts (.sg files)
  • Quality rule validation and schema enforcement
  • Statistical anomaly detection
  • 8 data format readers (Parquet, CSV, JSON, Avro, Delta, Iceberg)

Performance

  • Lazy evaluation planner (no eager materialization)
  • Columnar operator pipeline
  • Time-travel support for Delta/Iceberg snapshots

Upgrade: pip install --upgrade statguardian==1.0.0