Skip to content

v1.20.0

Choose a tag to compare

@thomasstvr thomasstvr released this 30 Aug 20:38
· 27 commits to main since this release
e78dfb6

WranglesPY 1.20.0 expands AI extraction and web search, adds text cleaning, similarity scoring, new data connectors and Excel formatting, and improves reliability across DataFrame and batch workflows.

New Additions

  • extract.ai now uses OpenAI's Responses API and gpt-5.4-mini by default, with configurable reasoning, concurrency, timeouts, retries, examples, structured outputs, and caching. Existing Chat Completions definitions remain supported. #1124
  • extract.ai also adds opt-in native web search through web_search: true, with a deduplicated web_search_sources column for each row. The public guidance option is now instructions, while existing messages calls remain supported. #1148
  • Added standardize.clean for local Unicode and text repair using ftfy, available through Python, recipes, and the DataFrame API. Existing model-backed standardization remains available as standardize.custom, and legacy standardize calls continue to work. #1133
  • compare.text adds bounded similarity scoring with token_sort, token_set, and damerau_levenshtein metrics. A shared normalization pipeline handles case, Unicode, punctuation, and whitespace while returning null for missing or blank values. #1144
  • create.embeddings now supports Jina through provider: jina; OpenAI remains the default provider. #961
  • Added DuckDB and Microsoft Access connectors with read, write, and run operations for recipes and Python. #1024
  • The Akeneo connector can now read paginated and filtered data in addition to writing it. #979
  • lookup now accepts n to return multiple ranked matches, either in one list or distributed across matching output columns. #1001
  • select.group_by adds a counts aggregation, while split.dictionary adds output_format: to_lists for separate key and value lists. #1021 #1129
  • Extract wrangles now offer consistent list, string, dictionary, and column output formats. Labeled extract.custom results can include empty labels or expand labels into columns. #880 #1032
  • extract.codes now exposes input-order sorting and extract_raw in recipe schemas, with clearer documentation for strategies, length limits, sorting, multi-part tokens, and disallowed patterns. #1022 #1023
  • Added a nested formatting.columns contract for excel.sheet, supporting alignment, number formats, bold text, and checkboxes in WranglesXL output. #1136
  • Added the direct Python API wrangles.train.delete(model_id) for deleting a model by ID. #997
  • Recipes can now use the automatic applied_permission_group variable to branch on the current user's effective permission group. Remote recipes prefer model metadata, local runs use the authenticated token, and explicitly supplied variables still take precedence. #1070
  • Remote recipes now accept an explicit :production selector and fall back to the latest version with a clear log message when no production version is set. #1128

Enhancements

  • extract.ai saved models now support expanded Excel schema fields, human-friendly list and object shorthand, and saved reasoning effort, with a new guide for defining schemas in Excel or YAML. Legacy model layouts remain supported. #1151
  • create.column no longer fails when the output already exists; recipes can preserve it, fill gaps, or replace it. #957
  • compute.case_when now uses pandas behavior and preserves existing output values for unmatched rows when no default is supplied. convert.case also uses faster vectorized conversions while leaving non-string values unchanged. #947 #949
  • Coalescing is now shared by the Python and recipe APIs, with more consistent handling of ragged data, whitespace, and falsey values. #964
  • Recipe loading and file reads/writes now accept pathlib.Path and other path-like objects. convert.from_json and convert.from_yaml also accept one default per input column. #987 #988
  • Logging now covers more operations while keeping normal INFO output focused on completed wrangles. Recipe errors also identify the failing step and source line, including nested and concurrent recipes. #954 #965 #1120 #1121
  • The python wrangle's except option now accepts falsey fallbacks such as {}, 0, false, and empty strings. Model access errors now distinguish expired or invalid authentication from insufficient permissions. #1020 #1107
  • Batch workers now receive copied DataFrame slices, reducing pandas view/copy ambiguity; yearly date ranges and regex handling are also compatible with newer pandas and Python versions. #1123

Bug Fixes

  • Models trained by name through classify, extract, lookup, or standardize are now created with complete payloads and work immediately after creation. #977
  • Fixed coalescing failures when all input columns contain floating-point values. #970
  • extract.ai now safely maps output names containing spaces, parentheses, and other special characters back to the requested columns. Malformed successful API responses now return a structured failure without reusing data from an earlier retry. #989 #1148
  • Blank Nullable, Required, and Additional Properties cells in saved extract.ai models now use their documented defaults, fixing nested object and array definitions imported from Excel. Explicit values and invalid nonblank entries are unchanged. #1153
  • Wrangles guarded by a false if condition are skipped before column validation, and a where clause that matches no rows no longer runs the wrangle on an empty DataFrame. #1000 #1007
  • Excel writes now remove invalid XML characters. Batched sheet writes retain every row, avoid partial outputs, and align changing columns by name instead of position. #1004 #1118
  • Excel files written without an explicit formatting block once again receive the default table style and top alignment. Explicit advanced formatting remains unchanged. #1140
  • Extraction wrangles now keep all matches when a single output name is written as a list. Multi-input extract.custom results remain lists, mismatched input/output counts fail clearly, and label capitalization is preserved. #1101 #1119 #1127
  • DataFrame accessor calls now return an independent result instead of mutating the original DataFrame. #1008
  • Sorting now handles columns containing a mix of numbers, numeric strings, and empty values. Rename steps can also skip a missing source when the normalized output already exists. #1066 #1068
  • The generated recipe schema now accepts where, where_params, and if on nested recipe wrangles. #1093

Breaking Changes

  • Named nested properties in extract.ai schemas now default to required and non-null, and named objects reject unknown properties by default. Set nullable: true, required, or additionalProperties explicitly where a different contract is needed. #1151
  • s3.upload_files now uses file for the local source path and save_as for the S3 destination key. Update recipes that use the previous save_as and file_key parameter names. #1003
  • compare.text difference, intersection, and overlap comparisons are now case-insensitive by default; set case_sensitive: true to retain exact-case matching. When overlap uses include_ratio: true, provide two output columns for the mask and numeric ratio instead of receiving both values in one cell. #1144
  • Excel file writes now reject constant_memory: true, which is incompatible with pandas column writes and table creation. Remove that option or use a different writing path. #1140

Developer and Release Improvements

  • Package metadata now reports version 1.20.0, so artifacts built from this release identify themselves correctly.
  • Optional connector dependencies are separated into requirements-full.txt, and release automation can build the full package variant when those integrations are needed. PyArrow is now loaded only when Parquet support is used; install the full dependency set when working with Parquet files. #939 #1035 #1146
  • CI and release workflows now build release candidates and main packages, publish the development recipe schema, and publish tagged distributions to CodeArtifact and PyPI through a consolidated pipeline. #999 #1030 #1090 #1115
  • Development and production deployments can be triggered automatically, record the initiating user and an optional reason, and wait for the downstream deployment result. The dispatch payload is now emitted in the format GitHub Actions expects. #1031 #1125 #1137
  • Everyday CI is faster and less expensive: macOS coverage is reserved for tagged releases, redundant test runs were removed, and workflow actions were upgraded for the Node 24 runtime. #1017 #1025 #1026
  • The repository's CI test image now runs Python 3.13 with pinned pandas 2.3.3 and NumPy 2.4.6, while the reusable package continues to support Python 3.11 through 3.13 and prohibits pandas 3.x. This does not change the separately owned production Lambda runtime. #1141
  • Container, concurrency, and regression tests are more reliable, generated test outputs have a dedicated ignored directory, and training tests clean up the models they create. #1009 #1027 #1028 #1029 #1078 #1122 #1143
  • Local development is standardized around a reproducible Python 3.13 setup, with improved Windows/WSL devcontainer support and repository-wide line-ending guidance. Package compatibility remains declared for Python 3.11 through 3.13. #1091 #1130
  • Added documented pull-request ownership, review, release, and issue-intake workflows to make handoffs and required actions clearer. #1079 #1094

Full Changelog: v1.19.6...v1.20.0