Repository navigation
v1.12.0
Highlights
If conditions
Add a condition to a read, wrangle, write or run. This will determine whether the action runs or not as a whole, in contrast to where which filters the rows that it will apply to.
# Send a failure email only if not testing
run:
on_failure:
- notification.email:
if: ${environment} != 'test'
...Matrix
Matrix can be used to run multiple actions defined by variables. Matrix was previously available for write. It is now also available for read, run and wrangles.
# Read all the files in a folder into a single dataframe
read:
- matrix:
variables:
filename: dir(my_folder)
read:
- file:
name: ${filename}Concurrent
By default, all functions executed by a recipe happen sequentially. The concurrent connector allows functions to be executed in parallel.
# Write to a file and database simultaneously.
write:
- concurrent:
write:
- file:
name: file1.csv
- postgres:
host: postgres.domain
...Try
Added a try wrangle. This allows trying a series of wrangles, catching any error that occurs and continuing anyway. Optionally, an except can be provided with wrangles to run in the event that the try fails, or a dictionary of keys and values to populate the dataframe with.
Try will fail as a whole, not per row.
# Try a wrangle that might fail and run another if it does
wrangles:
- try:
wrangles:
- risky_wrangle:
input: column
except:
- backup_wrangle:
input: columnNot Columns
Added the ability to use -name syntax for column names where wildcards are supported. This will exclude that column. Columns are evaluated from first to last.
# Apply to all columns like column1, column2, ... except column 2
wrangles:
- convert.case:
input:
- column*
- -column2
case: uppercompare.text
Added compare.text wrangle. This can be used to compare different strings.
wrangles:
- compare.text:
input:
- col1
- col2
output: output
method: intersectionAvailable methods include:
- intersection: Return only the text in common between the inputs
- difference: Return the text that differs between the inputs
- overlap: Show the text in common between two inputs
Other Changes
- log:
- If using alternative logging methods such as write, default not to write to the console unless specifically set.
- Fixed a bug where log used excessive memory for large dataframes.
- lookup: Added schema definition for training read/write.
- merge.concatentate: When concatenating a single column of strings, return that value rather than treating strings as lists.
- merge.coalesce: Added support for coalescing lists within a single column.
- rename: Fixed a bug where rename dropped the column if attempting to rename to itself.
- select.group_by: Allow using custom functions as aggregation functions.
- convert.to_yaml and convert.to_json: Accept multiple inputs for a single output. Will treat as a dictionary of the inputs.
- extract.ai:
- Set default model to gpt-4o-mini.
- Ensure examples are passed as an array even if set as a scalar value.
- python: Added an except parameter to provide a value to return in the case of an error.
- extract.brackets:
- Added find parameter to to specify which type of brackets to return values for.
- Added include_brackets boolean parameter to set whether to include the brackets in the returned strings. Default false.
- select.highest_confidence: Refinements to parameter and value handling.
- split.tokenize: Added method parameter. New methods are available to use 'boundary' and 'boundary_ignore_space' to split on regex word boundaries. Also
regex:<pattern>to split on a custom regex pattern andcustom.<function>to use a custom function. - S3: Fixed a bug where reading or writing gzipped files directly failed.
- input: New connector. This can be used in a read to reference the dataframe that was passed as part of the recipe.run function, e.g. to union or join with another source.
- recipe wrangle:
- Added support for using input, output and where to only apply the recipe to a subset of the data.
- Default to pass through all variables if variables parameter is not used instead of none.
- Improved the behaviour of where and handling of more edge cases.
- Support multiple reads. If a specific aggregation isn't used, the default will be to union together into a single dataframe.
- Custom functions for variables can now reference other variable by name in the arguments. e.g.
def func(other_variable): - Full support for nested function calls for custom functions. e.g.
custom.module.class.function - Fixed a bug where regex special characters had unintended side effects when using wildcards (*) in column names.
- Allow read, wrangles, write and run to be defined as strings and not require an empty dict for parameters if no parameters are needed.
- Fixed a bug where variables were not passed correctly to the recipe when they contained falsy values.
- Fixed a bug where certain special characters were not escaped correctly in user credentials.
- Fixed a bug where matrix write would not wait for all threads to complete fully.
- Updates to tests due to backend changes.
- General improvements to error messages.