Skip to content

v1.12.0

Choose a tag to compare

@ChrisWRWX ChrisWRWX released this 16 Oct 21:31
· 1156 commits to main since this release
4aaff49

Highlights

If conditions

Add a condition to a read, wrangle, write or run. This will determine whether the action runs or not as a whole, in contrast to where which filters the rows that it will apply to.

# Send a failure email only if not testing
run:
  on_failure:
     - notification.email:
         if: ${environment} != 'test'
         ...

Matrix

Matrix can be used to run multiple actions defined by variables. Matrix was previously available for write. It is now also available for read, run and wrangles.

# Read all the files in a folder into a single dataframe
read:
  - matrix:
      variables:
        filename: dir(my_folder)
      read:
        - file:
            name: ${filename}

Concurrent

By default, all functions executed by a recipe happen sequentially. The concurrent connector allows functions to be executed in parallel.

# Write to a file and database simultaneously.
write:
  - concurrent:
      write:
        - file:
            name: file1.csv
        - postgres:
             host: postgres.domain
             ...

Try

Added a try wrangle. This allows trying a series of wrangles, catching any error that occurs and continuing anyway. Optionally, an except can be provided with wrangles to run in the event that the try fails, or a dictionary of keys and values to populate the dataframe with.
Try will fail as a whole, not per row.

# Try a wrangle that might fail and run another if it does
wrangles:
  - try:
      wrangles:
        - risky_wrangle:
            input: column
      except:
        - backup_wrangle:
             input: column

Not Columns

Added the ability to use -name syntax for column names where wildcards are supported. This will exclude that column. Columns are evaluated from first to last.

# Apply to all columns like column1, column2, ... except column 2
wrangles:
  - convert.case:
      input:
        - column*
        - -column2
      case: upper

compare.text

Added compare.text wrangle. This can be used to compare different strings.

wrangles:
  - compare.text:
       input:
         - col1
         - col2
       output: output
       method: intersection

Available methods include:

  • intersection: Return only the text in common between the inputs
  • difference: Return the text that differs between the inputs
  • overlap: Show the text in common between two inputs

Other Changes

  • log:
    • If using alternative logging methods such as write, default not to write to the console unless specifically set.
    • Fixed a bug where log used excessive memory for large dataframes.
  • lookup: Added schema definition for training read/write.
  • merge.concatentate: When concatenating a single column of strings, return that value rather than treating strings as lists.
  • merge.coalesce: Added support for coalescing lists within a single column.
  • rename: Fixed a bug where rename dropped the column if attempting to rename to itself.
  • select.group_by: Allow using custom functions as aggregation functions.
  • convert.to_yaml and convert.to_json: Accept multiple inputs for a single output. Will treat as a dictionary of the inputs.
  • extract.ai:
    • Set default model to gpt-4o-mini.
    • Ensure examples are passed as an array even if set as a scalar value.
  • python: Added an except parameter to provide a value to return in the case of an error.
  • extract.brackets:
    • Added find parameter to to specify which type of brackets to return values for.
    • Added include_brackets boolean parameter to set whether to include the brackets in the returned strings. Default false.
  • select.highest_confidence: Refinements to parameter and value handling.
  • split.tokenize: Added method parameter. New methods are available to use 'boundary' and 'boundary_ignore_space' to split on regex word boundaries. Also regex:<pattern> to split on a custom regex pattern and custom.<function> to use a custom function.
  • S3: Fixed a bug where reading or writing gzipped files directly failed.
  • input: New connector. This can be used in a read to reference the dataframe that was passed as part of the recipe.run function, e.g. to union or join with another source.
  • recipe wrangle:
    • Added support for using input, output and where to only apply the recipe to a subset of the data.
    • Default to pass through all variables if variables parameter is not used instead of none.
  • Improved the behaviour of where and handling of more edge cases.
  • Support multiple reads. If a specific aggregation isn't used, the default will be to union together into a single dataframe.
  • Custom functions for variables can now reference other variable by name in the arguments. e.g. def func(other_variable):
  • Full support for nested function calls for custom functions. e.g. custom.module.class.function
  • Fixed a bug where regex special characters had unintended side effects when using wildcards (*) in column names.
  • Allow read, wrangles, write and run to be defined as strings and not require an empty dict for parameters if no parameters are needed.
  • Fixed a bug where variables were not passed correctly to the recipe when they contained falsy values.
  • Fixed a bug where certain special characters were not escaped correctly in user credentials.
  • Fixed a bug where matrix write would not wait for all threads to complete fully.
  • Updates to tests due to backend changes.
  • General improvements to error messages.