Repository navigation
LexiCore-5000 · release 1.0.0 — 5,000 English headwords, 1,000 per CEFR band (A1–C1), counted over 13,632,611,970 words in five corpora. Licence: CC BY 4.0.
Files in the tag
data/LexiCore_5000.csv (canonical) · .json (adds per-corpus ppm) · .txt (words only) · oxford_comparison_sets.csv · MANIFEST.json (full provenance) · data/lists/* · data/derived/* (the whole 215,416-candidate pool with its feature vector)
Verify what you download
sha256sum data/LexiCore_5000.csv
# 49ae7a92ef5a9301286a6607f2f2e1e4b8a0f381784ad65996c36af187854de4
python3 scripts/verify_checksums.py --data data # ledger + manifest + derived
python3 -m lexicore validate # release contractAsset: the release checkpoint (743 MB of pickles, not committed to the tree)
ckpt_lx5000v2_finalize_20260914T175041Z.tar.gz — 374,861,085 bytes,
SHA-256 74aed1c2e249da6074dcc03ebbc00cfc4d1543efcbd8d080af0d8534acff26aa.
Contains deliverables/ (as released), out/*.pkl (per-source counters, feature table, labels,
lemma map, selection), 12 stage logs and the 11 done/*.flag markers. Extract and rebuild the
derived tables with:
tar -xzf ckpt_lx5000v2_finalize_20260914T175041Z.tar.gz
python3 scripts/build_derived.py --checkpoint <dir>/out --out data/derived # deterministicCompanion product
LexiDeck Multilingual is the Anki system built on this list: the 5,000 words plus ~700
high-utility additions (≈5,700 words, A1–C1), 5 rotating example sentences per word
(28,000+), 38+ hours of human-recorded sentence audio, headword audio, IPA, an illustration per
word and translations in 12 languages.
- Product page: https://whop.com/lexideck/lexideck-multilingual/
- Free 50-word sample (.apkg): https://assets-2-prod.whop.com/public/uploads/2026-09-20/305d9e95-7a9d-44b1-8471-76eb2b4f8748/application.apkg
- In-repo documentation:
ecosystem/lexideck-multilingual/(description, measured card structure, screenshots)
Read before citing
The recency axis (recent_share_2022plus, newest_share_2025_2026) is empty for all 5,000 rows,
pos is a dictionary guess, 437 levels come from a model with CV accuracy 0.38, and the pipeline
code is not part of the checkpoint. Full list: docs/KNOWN_ISSUES.md.
1.0.0 - 2026-09-14
Added
LexiCore_5000— 5,000 single-token English headwords, 1,000 each in A1/A2/B1/B2/C1,
ranked by a weighted score over frequency (0.34), dispersion (0.18), teaching-list utility
(0.18), source entropy (0.10), range (0.08), document frequency (0.06) and recency (0.06).- Counting evidence over five corpora totalling 13,632,611,970 words:
FineWebsample/10BT, FineWeb-Edu-score-2, Wikipedia20231101.en,
OPUS-OpenSubtitles v2024, Common CrawlCC-MAIN-2026-34. - CEFR levels from CEFR-J v1.5 (3,898) and Octanove C1/C2 (665), with a multinomial logistic
model for the remaining 437 (CV accuracy 0.3814, macro-F1 0.3829). - External comparison against Oxford 5000 (never used in construction): 3,526 shared,
Jaccard 0.5508, exact level agreement 0.466, quadratic-weighted κ 0.7124. - Release manifest with per-file URLs, byte sizes and SHA-256 for every input.
Known in this release (documented, not hidden)
recent_share_2022plusandnewest_share_2025_2026are empty for all 5,000 rows: the recency
axis never reached the feature table, so its 0.06 weight was inert
(KI-2).posis a dictionary lookup, not corpus tagging (KI-3).confidencefor profile-derived labels is a constant cap, not a posterior
(KI-5).- The upstream
REPORT.md"modern additions, ranked by 2025-26 evidence" section is alphabetical
(KI-1). - Pipeline source code and the anchor wordlist files are not in the release checkpoint
(KI-12).
[1.0.0+repo] - 2026-09-21
Added
- Public repository packaging of the release: data, docs, tooling.
data/SHA256SUMS.txtintegrity ledger over 53 files, plusmake checksums/make ledger.data/lists/per-level handouts (rank order and alphabetical).data/derived/:candidate_features.csv.gz(215,416 candidates × the full feature vector),
candidate_labels.csv.gz(levels, sources, confidences, class probabilities),
surface_to_lemma.csv.gz(283,502 forms) — deterministic rebuild viamake derived.lexicorePython package:Entrymodel, CSV/JSON/TXT loaders,validate_data(),
a re-scoring harness for the published formula, and a CLI (words,show,stats,
validate,reweight).- 16 pytest cases: release contract + library behaviour, including loss-free CSV round-trip.
docs/: methodology, provenance (with live re-verification of source URLs/hashes), data
dictionary, validation, known issues, reproduction guide, artefact ledger.- CI (GitHub Actions) running validate / test / lint on Python 3.10, 3.12, 3.13.
ecosystem/lexideck-multilingual/: the companion Anki deck (≈5,700 words = this list plus
~700 additions), its card model schema as measured from the free sample, and screenshots.
Changed
- Upstream artefacts kept unmodified under
reports/upstream/; only the files indata/are
renamed/documented by this repository.
Not changed (on purpose)
- No word was added, removed or re-levelled. 1.0.0's content is exactly the released content.