Skip to content

Release v0.21.0

Choose a tag to compare

@Edwardvaneechoud Edwardvaneechoud released this 26 Sep 16:22
· 34 commits to main since this release
ec3274e

Changes since v0.20.0.

This release brings the Python API up to the canvas. Everything you can place in the designer you can now place from Python, and the result opens in the designer as ordinary editable nodes: gates, subflows, custom nodes, Python Script notebooks, SQL nodes and typed flow parameters. Building a flow in Python no longer runs subflows, kernels or writers; they run when the flow runs. No migration.

New features

Native node classes in the Python API (#773)

Nodes that have no fluent method get a class. Each call places one canvas node and returns its outputs as frames.

  • Gates. ff.Gate(frame, "[amount] > 1000") or ff.Gate(frame, parameter=MODE, value="full"). Chain from .then and .otherwise; only the live side runs, the other is marked skipped, exactly as on the canvas. A formula can also be an expression: ff.Gate(frame, ff.col("amount") > 1000).
  • Subflows. ff.RunFlow(ff.flow_ref("Sales", "Clean orders"), inputs={"orders": orders}, params={"min_amount": 25}) runs a registered flow and exposes its outputs as run["orders_clean"]. Declare a child's ports with ff.FlowInput("orders", sample=...) and frame.to_flow_output("orders_clean"); register it with ff.register_flow(child, name=...) or schema.register_flow(...).
  • Custom nodes. ff.custom_nodes.<key>(frame, source_column="x", threshold=10) places an installed node with flat keyword settings, ff.custom_nodes.list() shows what is installed, ff.custom_nodes.install(MyNode) writes a class to your node directory. ff.CustomNode(cls_or_key, *inputs, settings=..., kernel=...) is the long form.
  • Python Script notebooks. Decorate a function with @ff.python_script(kernel="lite", outputs=["forecast"]): parameters are the input frames, # %% splits the body into cells, the return is the published output. Call it like a function to place the node. ff.PythonScript(cells=[...], kernel=...) places one from cell text.
  • SQL. ff.sql("SELECT ... FROM orders o JOIN regions r ...", orders=clean, regions=regions) and frame.sql("SELECT * FROM self WHERE ...").
  • Typed flow parameters. ff.Parameter("min_amount", default=0, type="integer") with ff.add_flow_parameter(flow, p) and ff.set_flow_parameter(flow, p, 25). Use a parameter directly in a filter or a custom node setting; it is stored as ${min_amount} and resolved when the flow runs.
  • Any other node. ff.Node("record_count", frame) places a built-in node from its settings. ff.NodeType, ff.GateOperator and ff.ParamType enumerate the valid names.
  • Catalog. ff.get_catalog("Demo").get_schema("sales") gives you read_table, register_flow and get_flow; ff.read_catalog_sql is now exported at the top level.

Deferred frames (#773)

A frame whose node only produces data at run time (a subflow, a kernel script, an external source, or a writer below one of those or below a gate) is built with a typed empty placeholder instead of being executed. You keep chaining on it and the downstream nodes type-check; collect() runs that frame's lineage and returns the real rows, or an empty frame for the dead side of a gate. Writers on other branches are left alone.

PyCharm console (#773)

@ff.python_script and map_elements lambdas work for code executed as a selection in PyCharm's Python console, which normally keeps no source. Other embedded consoles can opt in with FLOWFILE_CONSOLE_SOURCE=1.

Code export

  • SQL Query nodes export to Python, as pl.SQLContext(...) in the Polars export and ff.sql(...) in the FlowFrame export. (#773)

Improvements

  • Dashboards. Faceted Graphic Walker charts fit inside their tiles. (#772)
  • Custom nodes. ${param} references in settings reach process() and predict_output_schema resolved; numeric and toggle settings turn substituted text into numbers and booleans, so a toggle set to "false" is now off. Kernel nodes without a schema hook no longer try to predict their columns by running; the drawer says to run the node. (#773)
  • Run Flow. A child flow without outputs shows its run-summary columns before it runs. The stored reference now carries the namespace and name and resolves by uuid first, so a parent opened on another install fails with a clear message instead of running whichever flow holds that id.
  • Merging graphs keeps second output handles and keyed inputs, so a filter split's fail branch or a gate's else exit survives a merge. (#773)
  • The kernel-integration CI job now also runs the Python API's Docker tests.

Fixes

  • A parameter value containing quotes, newlines or backslashes no longer breaks a Polars-code node when it is used inside a string literal.
  • A Python Script "run cell" in the designer no longer swaps inputs once a flow has ten or more nodes. An upstream node referenced as main is refused with a message, since main is the positional alias of all inputs.
  • fuzzy_join and concat in the Python API read the output handle they were given; a filter split's fail branch used to fall through to the pass branch.
  • A lambda defined in the PyCharm console is matched on its default values too, so a redefined lambda no longer answers for an older handle.

Upgrade notes

  • Python API: ff.col("${x}") and alias("${x}") now raise; use ff.Parameter for a parameter and a plain name for a column.
  • Python API: errors from wiring nodes are NativeNodeError (a ValueError) instead of an HTTP exception object; ff.flow_ref, ff.register_flow and ff.RunFlow raise NativeNodeError with the fix in the message, with the catalog error as its cause.
  • New environment variable FLOWFILE_CONSOLE_SOURCE (truthy 1/true/yes/on) forces the console-source hook on; it otherwise installs only in an interactive console.
  • No database migration.