Skip to content

Release v0.17.5

Choose a tag to compare

@Edwardvaneechoud Edwardvaneechoud released this 14 Sep 20:53
· 74 commits to main since this release

Changes since v0.17.4.

A new List Files node turns a folder into a table, every Excel read moved onto one shared reader with a rebuilt settings panel, and custom nodes get a File Picker control.

List Files

A new input node (#733): one row per entry in a folder, with a fixed schema — file_name, file_path, directory, relative_path, file_type, size_bytes, last_modified, created_date, is_directory.

  • Filter by extension, include folders or hidden entries, recurse to a Max depth, cap the row count.
  • The schema is fixed, so the folder is only walked at run time, and a long walk over a big or slow tree is cancellable.
  • Python: ff.list_files(directory, file_types=None, recursive=False, ...). Both code exports emit the same call.
  • Not in Flowfile Lite; share links demote it to a placeholder.

Inventory a drop folder, filter on size or modification date, feed file_path into whatever reads the files. To read a whole folder as one table, the Read node's Directory scan mode is still the right tool.

Docs: List Files.

Excel reading

Core and the worker now read Excel through one module instead of two copies that had drifted apart (#734). Engine choice is explicit — calamine (fast, the default), openpyxl (permissive, what Type inference selects), xlsx2csv for headerless reads — logged to the flow log, and a sheet the fast engine cannot handle degrades to openpyxl with a warning instead of failing the run.

Fixed along the way:

  • Toggling Type inference changed which cells were read: xlsx2csv treated end_row as a row count after the skip, reading start_row rows too many. All three engines now honour the same 0-based, end-inclusive bounds.
  • A headerless table starting at A1 lost its first row to calamine's header. Headerless reads go through xlsx2csv now.
  • Metadata above the table or a stray value beside it raised ShapeError: N column names provided for a DataFrame of width M.
  • A blank sheet name now means "read the first sheet" rather than failing.
  • A column mixing ints and floats (141000, 87500, 162500.25) crashed the read; a purely numeric column is widened to float instead.

The settings panel was rebuilt: Sheet name is optional and offers the workbook's actual sheets, with a warning when the configured name isn't one of them, and the row/column range moved into a collapsible optional group.

Node designer: File Picker

nd.FilePicker — a path field with a Browse button that opens the same file browser the Read and Write nodes use (#731). Configure mode ("open" / "create"), file_types and allow_directory. The browser runs on the machine hosting flowfile_core, so in Docker and the web UI it lists the server's filesystem, not the viewer's.

Docs: File Picker.

Designer

Flow tabs scroll (#732). Chevrons appear at whichever edge is clipped, a vertical wheel scrolls the strip horizontally, and the active tab is scrolled into view.

Previews: ready for complex types

Groundwork for Polars 2, where extension and nested dtypes become ordinary (#730).

  • One bad cell no longer blanks the whole preview. Unserializable values fall back to a bounded string; lists and structs stay JSON.
  • Binary cells and stats bounds render as hex — 0x48656C6C6F… (24 bytes) — in previews, column stats, catalog and Lite.
  • Geometry groundwork. Columns carry a semantic_type, declared from the dtype and never sniffed, that every surface labels from. Nothing emits a geoarrow.* dtype yet, so it's inert until a producer lands.