Skip to content

v2.0.0-alpha — a new app, a real corpus

Pre-release
Pre-release

Choose a tag to compare

@maehr maehr released this 21 Sep 07:04
· 13 commits to main since this release
v2.0.0-alpha
c49677c

The first public build of version 2. This is an alpha. Expect rough edges.

Version 2 shares no code with version 1. The app is now a marimo notebook
that runs on Pyodide, so a static web server is enough to host it.

Try it: https://maehr.github.io/simple-topic-modeling/

Your documents stay in your browser. The app sends no text to an analysis server, and it calls no
model API.

What is new since version 1

  • NMF and LDA. Version 1 offered LDA only. NMF is now the default, and it is faster on short
    and medium documents.
  • A real demo corpus. The app opens with 295 articles from the Journal de Genève and the
    Gazette de Lausanne of 1914, with a real date column and the newspaper's own section headings.
    A first run gives six topics that you can check against those sections.
  • Onboarding for a newcomer. An intro, a four-step walkthrough, a jargon-free explanation of a
    topic model, an orientation line over the result tabs, and a glossary.
  • Five result tabs. Overview, Topics, Documents, Metadata, and Diagnostics, with a topic map, a
    similarity heatmap, word clouds, and a document explorer.
  • Progressive disclosure. Step 2 shows the model and the number of topics. The tuning controls
    wait in a closed Advanced settings panel.
  • Repeatable runs. A fixed random seed, a config.json export that records every setting and
    the app version, and a config.json upload that restores them.
  • Exports. documents_topics.csv, topics.csv, topic_terms.csv, topic_similarity.csv,
    config.json, and a project.zip that bundles all of them.
  • Five stop-word languages. English, German, French, Italian, and Spanish, each editable.

Known limits of this alpha

  • The demo corpus is French. The interface is English.
  • The demo text comes from a 1914 scan, so some words carry OCR errors. A real archive looks like
    this, and the minimum document frequency clears most of the noise.
  • PDF, DOCX, images, and ZIP input are out of scope.

The demo corpus

The Digital Humanities Laboratory of the EPFL digitised the historical archive of
Le Temps and published the year 1914 under CC BY 4.0, for the
2015 Swiss Open Cultural Data Hackathon. The articles
are anonymous newspaper text from 1914, so they left copyright in 1985.
NOTICE holds the full
statement, and scripts/build_demo_corpus.py rebuilds the corpus and records its provenance.

Feedback

Please report a problem or ask for a
feature
. The full list of changes is in
CHANGELOG.md.