Wordflow 0.7.9 lets you add files of any size, and makes Concordance easier to sort and read.
Files
- No file size limit: files over 64 MiB can now be previewed and added as a Data Block, for example a 500 MB Parquet file. Adding a very large file needs a few GB of memory for a moment.
Analysis tools
- Concordance: with Highlight L1/R1 for sorting on (the default), clicking the left context header sorts by L1 and the right context header by R1; off, each context sorts by its own text. L1 and R1 now show as coloured words instead of a highlighter tint. The table no longer shows the match offsets (
CONC_start_idx,CONC_end_idx); Data Blocks added to the Project keep them. - Topic Modelling: each Data Block's sample uses the Seed setting, so the same Seed and percentage pick the same documents whichever order the Data Blocks are in.
- Frequency and Topic Modelling: picking a stop words list switches the stop words filter on.
Everywhere
- Tasks panel: every task starts with its tool's letter, for example "C - yeah" for a Concordance opened from a word. Imports read "L - Sample data import"; a finished import says what it imported, and its arrow opens that folder in the Data Loader.
- Windows: Topic Modelling and other analyses no longer fail with a "file not found" error under the default data folder (also in the updated 0.7.8 Windows installer).
Sample data
- Two new collections from RAPID-CDL's database of the Australian Federal Parliament's proceedings (Hansard): paragraphs about the GST, and paragraphs about housing, 1998 to 2026, each with the paragraphs either side. Import them from the Data Loader's sample data.
Upgrading
Re-running a Topic Modelling tab whose second Data Block is sampled below 100% picks a different sample for that Data Block than 0.7.8 did.