v2.0.0-alpha — a new app, a real corpus
Pre-releaseThe first public build of version 2. This is an alpha. Expect rough edges.
Version 2 shares no code with version 1. The app is now a marimo notebook
that runs on Pyodide, so a static web server is enough to host it.
Try it: https://maehr.github.io/simple-topic-modeling/
Your documents stay in your browser. The app sends no text to an analysis server, and it calls no
model API.
What is new since version 1
- NMF and LDA. Version 1 offered LDA only. NMF is now the default, and it is faster on short
and medium documents. - A real demo corpus. The app opens with 295 articles from the Journal de Genève and the
Gazette de Lausanne of 1914, with a real date column and the newspaper's own section headings.
A first run gives six topics that you can check against those sections. - Onboarding for a newcomer. An intro, a four-step walkthrough, a jargon-free explanation of a
topic model, an orientation line over the result tabs, and a glossary. - Five result tabs. Overview, Topics, Documents, Metadata, and Diagnostics, with a topic map, a
similarity heatmap, word clouds, and a document explorer. - Progressive disclosure. Step 2 shows the model and the number of topics. The tuning controls
wait in a closed Advanced settings panel. - Repeatable runs. A fixed random seed, a
config.jsonexport that records every setting and
the app version, and aconfig.jsonupload that restores them. - Exports.
documents_topics.csv,topics.csv,topic_terms.csv,topic_similarity.csv,
config.json, and aproject.zipthat bundles all of them. - Five stop-word languages. English, German, French, Italian, and Spanish, each editable.
Known limits of this alpha
- The demo corpus is French. The interface is English.
- The demo text comes from a 1914 scan, so some words carry OCR errors. A real archive looks like
this, and the minimum document frequency clears most of the noise. - PDF, DOCX, images, and ZIP input are out of scope.
The demo corpus
The Digital Humanities Laboratory of the EPFL digitised the historical archive of
Le Temps and published the year 1914 under CC BY 4.0, for the
2015 Swiss Open Cultural Data Hackathon. The articles
are anonymous newspaper text from 1914, so they left copyright in 1985.
NOTICE holds the full
statement, and scripts/build_demo_corpus.py rebuilds the corpus and records its provenance.
Feedback
Please report a problem or ask for a
feature. The full list of changes is in
CHANGELOG.md.