Skip to content

v1.56 Hot town, summer in the city

Choose a tag to compare

@stijn-uva stijn-uva released this 09 Jul 13:26
· 53 commits to master since this release
Immutable release. Only release title and notes can be modified.

This 4CAT update comprises expanded support for LLMs; LLMs from multiple servers and providers can be made available to LLM-using processors, and 4CAT can be run via Docker with an embedded Ollama server through which LLMs can be loaded and enabled (requires a re-build; see README).

Additionally, there are various smaller improvements for bug fixes for processors, data sources, and the web interface.

⚠️ Docker users are recommended to rebuild their containers to benefit from some of the speed boosts implemented in previous updates as well as to use the latest available versions of its dependencies (without which some processors, particularly those processing videos, will fail more readily). This will become mandatory in a future release. For now rebuilding is optional and 4CAT will otherwise function normally, but sometimes slower and less effectively than it could be.

⚠️ Please also follow these instructions for upgrading if you have trouble upgrading to the latest version of 4CAT.

Otherwise, you can upgrade 4CAT via the 'Restart or upgrade' button in the Control Panel. This release of 4CAT incorporates the following fixes and improvements:

New processor features and other processor updates

  • Update the 'Word Tree' processor to account for emoji, do wildcard matching more accurately, use TweetTokenizer to tokenize text by default, and be more efficient (9f3e1fe)
  • Update the ‘Random sample’ processor to allow the creation of random samples of datasets consisting of a zip file (i.e. containing a random sample of files) (82d627b)
  • Update the ‘Image wall with captions’ processor to allow it to be run on the output of LLM-based processors (bddfa37)

GenAI-related features and fixes

  • An alternative Docker Compose file is available to create a 4CAT instance with an embedded Ollama server, for easy use of LLM-based processors locally. See README.md in the ‘Docker’ folder for further instructions. (#576)
  • It is now possible to add multiple LLM servers as a source for LLMs 4CAT uses for e.g. the ‘LLM Prompting’ processor. Currently, support is available for Ollama, LiteLLM, OpenAI-compatible servers (e.g. LM Studio or vLLM), and a selection of third-party commercial APIs (e.g. Gemini, GPT, Mistral, Deepseek). Ollama models can be added and deleted via the 4CAT web interface. (#576)

Processor and data source bug fixes

  • Fix an issue where BlueSky datasets could not be processed properly if they contained a date in an unexpected format (a3a3c7f)
  • Fix an issue where Instagram datasets could not be processed properly if they were missing certain user data (98f281c)
  • Fix an issue in the ‘Download videos’ processor where videos would not be downloaded with yt-dlp even if yt-dlp was enabled and available (#612)
  • Fix an issue in the ‘Thread metrics’ processor where the ‘op_length’ column was erroneously labeled ‘op_replies’ (4b85bd0)
  • Fix an issue where ‘imported from X/Twitter’ datasets could not be filtered for and would not be listed on the data sources overview page (c1035bb)

New Web UI features and other general 4CAT updates

  • If a processor/job crashes while executing, this is now made clearer in the UI and will not prevent other procesors of the same type from running (#602)
  • Hide ‘indirect’ settings (that cannot be changed directly) from the Settings page in the Control Panel (7c88e0b)

Deprecations

  • Processors that were specific to DMI-TCAT have been removed from 4CAT (8dbf76a)

Full Changelog: v1.55...v1.56