Repository navigation
v1.20.0
WranglesPY 1.20.0 expands AI extraction and web search, adds text cleaning, similarity scoring, new data connectors and Excel formatting, and improves reliability across DataFrame and batch workflows.
New Additions
extract.ainow uses OpenAI's Responses API andgpt-5.4-miniby default, with configurable reasoning, concurrency, timeouts, retries, examples, structured outputs, and caching. Existing Chat Completions definitions remain supported. #1124extract.aialso adds opt-in native web search throughweb_search: true, with a deduplicatedweb_search_sourcescolumn for each row. The public guidance option is nowinstructions, while existingmessagescalls remain supported. #1148- Added
standardize.cleanfor local Unicode and text repair usingftfy, available through Python, recipes, and the DataFrame API. Existing model-backed standardization remains available asstandardize.custom, and legacystandardizecalls continue to work. #1133 compare.textadds bounded similarity scoring withtoken_sort,token_set, anddamerau_levenshteinmetrics. A shared normalization pipeline handles case, Unicode, punctuation, and whitespace while returning null for missing or blank values. #1144create.embeddingsnow supports Jina throughprovider: jina; OpenAI remains the default provider. #961- Added DuckDB and Microsoft Access connectors with read, write, and run operations for recipes and Python. #1024
- The Akeneo connector can now read paginated and filtered data in addition to writing it. #979
lookupnow acceptsnto return multiple ranked matches, either in one list or distributed across matching output columns. #1001select.group_byadds acountsaggregation, whilesplit.dictionaryaddsoutput_format: to_listsfor separate key and value lists. #1021 #1129- Extract wrangles now offer consistent list, string, dictionary, and column output formats. Labeled
extract.customresults can include empty labels or expand labels into columns. #880 #1032 extract.codesnow exposes input-order sorting andextract_rawin recipe schemas, with clearer documentation for strategies, length limits, sorting, multi-part tokens, and disallowed patterns. #1022 #1023- Added a nested
formatting.columnscontract forexcel.sheet, supporting alignment, number formats, bold text, and checkboxes in WranglesXL output. #1136 - Added the direct Python API
wrangles.train.delete(model_id)for deleting a model by ID. #997 - Recipes can now use the automatic
applied_permission_groupvariable to branch on the current user's effective permission group. Remote recipes prefer model metadata, local runs use the authenticated token, and explicitly supplied variables still take precedence. #1070 - Remote recipes now accept an explicit
:productionselector and fall back to the latest version with a clear log message when no production version is set. #1128
Enhancements
extract.aisaved models now support expanded Excel schema fields, human-friendly list and object shorthand, and saved reasoning effort, with a new guide for defining schemas in Excel or YAML. Legacy model layouts remain supported. #1151create.columnno longer fails when the output already exists; recipes can preserve it, fill gaps, or replace it. #957compute.case_whennow uses pandas behavior and preserves existing output values for unmatched rows when no default is supplied.convert.casealso uses faster vectorized conversions while leaving non-string values unchanged. #947 #949- Coalescing is now shared by the Python and recipe APIs, with more consistent handling of ragged data, whitespace, and falsey values. #964
- Recipe loading and file reads/writes now accept
pathlib.Pathand other path-like objects.convert.from_jsonandconvert.from_yamlalso accept one default per input column. #987 #988 - Logging now covers more operations while keeping normal INFO output focused on completed wrangles. Recipe errors also identify the failing step and source line, including nested and concurrent recipes. #954 #965 #1120 #1121
- The
pythonwrangle'sexceptoption now accepts falsey fallbacks such as{},0,false, and empty strings. Model access errors now distinguish expired or invalid authentication from insufficient permissions. #1020 #1107 - Batch workers now receive copied DataFrame slices, reducing pandas view/copy ambiguity; yearly date ranges and regex handling are also compatible with newer pandas and Python versions. #1123
Bug Fixes
- Models trained by name through
classify,extract,lookup, orstandardizeare now created with complete payloads and work immediately after creation. #977 - Fixed coalescing failures when all input columns contain floating-point values. #970
extract.ainow safely maps output names containing spaces, parentheses, and other special characters back to the requested columns. Malformed successful API responses now return a structured failure without reusing data from an earlier retry. #989 #1148- Blank
Nullable,Required, andAdditional Propertiescells in savedextract.aimodels now use their documented defaults, fixing nested object and array definitions imported from Excel. Explicit values and invalid nonblank entries are unchanged. #1153 - Wrangles guarded by a false
ifcondition are skipped before column validation, and awhereclause that matches no rows no longer runs the wrangle on an empty DataFrame. #1000 #1007 - Excel writes now remove invalid XML characters. Batched sheet writes retain every row, avoid partial outputs, and align changing columns by name instead of position. #1004 #1118
- Excel files written without an explicit formatting block once again receive the default table style and top alignment. Explicit advanced formatting remains unchanged. #1140
- Extraction wrangles now keep all matches when a single output name is written as a list. Multi-input
extract.customresults remain lists, mismatched input/output counts fail clearly, and label capitalization is preserved. #1101 #1119 #1127 - DataFrame accessor calls now return an independent result instead of mutating the original DataFrame. #1008
- Sorting now handles columns containing a mix of numbers, numeric strings, and empty values. Rename steps can also skip a missing source when the normalized output already exists. #1066 #1068
- The generated recipe schema now accepts
where,where_params, andifon nested recipe wrangles. #1093
Breaking Changes
- Named nested properties in
extract.aischemas now default to required and non-null, and named objects reject unknown properties by default. Setnullable: true,required, oradditionalPropertiesexplicitly where a different contract is needed. #1151 s3.upload_filesnow usesfilefor the local source path andsave_asfor the S3 destination key. Update recipes that use the previoussave_asandfile_keyparameter names. #1003compare.textdifference, intersection, and overlap comparisons are now case-insensitive by default; setcase_sensitive: trueto retain exact-case matching. Whenoverlapusesinclude_ratio: true, provide two output columns for the mask and numeric ratio instead of receiving both values in one cell. #1144- Excel file writes now reject
constant_memory: true, which is incompatible with pandas column writes and table creation. Remove that option or use a different writing path. #1140
Developer and Release Improvements
- Package metadata now reports version
1.20.0, so artifacts built from this release identify themselves correctly. - Optional connector dependencies are separated into
requirements-full.txt, and release automation can build the full package variant when those integrations are needed. PyArrow is now loaded only when Parquet support is used; install the full dependency set when working with Parquet files. #939 #1035 #1146 - CI and release workflows now build release candidates and main packages, publish the development recipe schema, and publish tagged distributions to CodeArtifact and PyPI through a consolidated pipeline. #999 #1030 #1090 #1115
- Development and production deployments can be triggered automatically, record the initiating user and an optional reason, and wait for the downstream deployment result. The dispatch payload is now emitted in the format GitHub Actions expects. #1031 #1125 #1137
- Everyday CI is faster and less expensive: macOS coverage is reserved for tagged releases, redundant test runs were removed, and workflow actions were upgraded for the Node 24 runtime. #1017 #1025 #1026
- The repository's CI test image now runs Python 3.13 with pinned pandas 2.3.3 and NumPy 2.4.6, while the reusable package continues to support Python 3.11 through 3.13 and prohibits pandas 3.x. This does not change the separately owned production Lambda runtime. #1141
- Container, concurrency, and regression tests are more reliable, generated test outputs have a dedicated ignored directory, and training tests clean up the models they create. #1009 #1027 #1028 #1029 #1078 #1122 #1143
- Local development is standardized around a reproducible Python 3.13 setup, with improved Windows/WSL devcontainer support and repository-wide line-ending guidance. Package compatibility remains declared for Python 3.11 through 3.13. #1091 #1130
- Added documented pull-request ownership, review, release, and issue-intake workflows to make handoffs and required actions clearer. #1079 #1094
Full Changelog: v1.19.6...v1.20.0