Skip to content

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 10 Mar 11:04
· 288 commits to main since this release
bce1d41
  • Added uniqueness check(#200). A uniqueness check has been added, which reports an issue for each row containing a duplicate value in a specified column. This resolves issue 154.
  • Added sql expression support for limits in not less and not greater than checks, and updated docs (#200). This commit introduces several changes to simplify and enhance data quality checking in PySpark workloads for both streaming and batch data. The naming conventions of rule functions have been unified, and the is_not_less_than and is_not_greater_than functions now accept column names or expressions as limits. The input parameters for range checks have been unified, and the logic of is_not_in_range has been updated to be inclusive of the boundaries. The project's documentation has been improved, with the addition of comprehensive examples, and the contribution guidelines have been clarified. This change includes a breaking change for some of the checks. Users are advised to review and test the changes before implementation to ensure compatibility and avoid any disruptions. Resolves issues: 131, 197, 175, 205
  • Include predefined check functions by default when applying custom checks by metadata (#203). The data quality engine has been updated to include predefined check functions by default when applying custom checks using metadata in the form of YAML or JSON. This change simplifies the process of defining custom checks, as users no longer need to specify globals(). The default behavior now is to import all predefined checks. The validate_checks method has been updated to accept a dictionary of custom check functions instead of global variables. However, globals() can still be specified for backward compatibility. This improvement resolves the issue #48.

Contributors: @mwojtyczka