Skip to content

v0.7.9

Latest

Choose a tag to compare

@milysun milysun released this 29 Sep 14:37
· 20 commits to main since this release

Wordflow 0.7.9 lets you add files of any size, and makes Concordance easier to sort and read.

Files

  • No file size limit: files over 64 MiB can now be previewed and added as a Data Block, for example a 500 MB Parquet file. Adding a very large file needs a few GB of memory for a moment.

Analysis tools

  • Concordance: with Highlight L1/R1 for sorting on (the default), clicking the left context header sorts by L1 and the right context header by R1; off, each context sorts by its own text. L1 and R1 now show as coloured words instead of a highlighter tint. The table no longer shows the match offsets (CONC_start_idx, CONC_end_idx); Data Blocks added to the Project keep them.
  • Topic Modelling: each Data Block's sample uses the Seed setting, so the same Seed and percentage pick the same documents whichever order the Data Blocks are in.
  • Frequency and Topic Modelling: picking a stop words list switches the stop words filter on.

Everywhere

  • Tasks panel: every task starts with its tool's letter, for example "C - yeah" for a Concordance opened from a word. Imports read "L - Sample data import"; a finished import says what it imported, and its arrow opens that folder in the Data Loader.
  • Windows: Topic Modelling and other analyses no longer fail with a "file not found" error under the default data folder (also in the updated 0.7.8 Windows installer).

Sample data

  • Two new collections from RAPID-CDL's database of the Australian Federal Parliament's proceedings (Hansard): paragraphs about the GST, and paragraphs about housing, 1998 to 2026, each with the paragraphs either side. Import them from the Data Loader's sample data.

Upgrading

Re-running a Topic Modelling tab whose second Data Block is sampled below 100% picks a different sample for that Data Block than 0.7.8 did.