Skip to content

LexiCore-5000 · release 1.0.0

Latest

Choose a tag to compare

@X-Trivle X-Trivle released this 21 Sep 14:30
· 5 commits to main since this release

LexiCore-5000 · release 1.0.0 — 5,000 English headwords, 1,000 per CEFR band (A1–C1), counted over 13,632,611,970 words in five corpora. Licence: CC BY 4.0.

Files in the tag

data/LexiCore_5000.csv (canonical) · .json (adds per-corpus ppm) · .txt (words only) · oxford_comparison_sets.csv · MANIFEST.json (full provenance) · data/lists/* · data/derived/* (the whole 215,416-candidate pool with its feature vector)

Verify what you download

sha256sum data/LexiCore_5000.csv
# 49ae7a92ef5a9301286a6607f2f2e1e4b8a0f381784ad65996c36af187854de4
python3 scripts/verify_checksums.py --data data   # ledger + manifest + derived
python3 -m lexicore validate                       # release contract

Asset: the release checkpoint (743 MB of pickles, not committed to the tree)

ckpt_lx5000v2_finalize_20260914T175041Z.tar.gz — 374,861,085 bytes,
SHA-256 74aed1c2e249da6074dcc03ebbc00cfc4d1543efcbd8d080af0d8534acff26aa.
Contains deliverables/ (as released), out/*.pkl (per-source counters, feature table, labels,
lemma map, selection), 12 stage logs and the 11 done/*.flag markers. Extract and rebuild the
derived tables with:

tar -xzf ckpt_lx5000v2_finalize_20260914T175041Z.tar.gz
python3 scripts/build_derived.py --checkpoint <dir>/out --out data/derived   # deterministic

Companion product

LexiDeck Multilingual is the Anki system built on this list: the 5,000 words plus ~700
high-utility additions (≈5,700 words, A1–C1)
, 5 rotating example sentences per word
(28,000+), 38+ hours of human-recorded sentence audio, headword audio, IPA, an illustration per
word and translations in 12 languages.

Read before citing

The recency axis (recent_share_2022plus, newest_share_2025_2026) is empty for all 5,000 rows,
pos is a dictionary guess, 437 levels come from a model with CV accuracy 0.38, and the pipeline
code is not part of the checkpoint. Full list: docs/KNOWN_ISSUES.md.


1.0.0 - 2026-09-14

Added

  • LexiCore_5000 — 5,000 single-token English headwords, 1,000 each in A1/A2/B1/B2/C1,
    ranked by a weighted score over frequency (0.34), dispersion (0.18), teaching-list utility
    (0.18), source entropy (0.10), range (0.08), document frequency (0.06) and recency (0.06).
  • Counting evidence over five corpora totalling 13,632,611,970 words:
    FineWeb sample/10BT, FineWeb-Edu-score-2, Wikipedia 20231101.en,
    OPUS-OpenSubtitles v2024, Common Crawl CC-MAIN-2026-34.
  • CEFR levels from CEFR-J v1.5 (3,898) and Octanove C1/C2 (665), with a multinomial logistic
    model for the remaining 437 (CV accuracy 0.3814, macro-F1 0.3829).
  • External comparison against Oxford 5000 (never used in construction): 3,526 shared,
    Jaccard 0.5508, exact level agreement 0.466, quadratic-weighted κ 0.7124.
  • Release manifest with per-file URLs, byte sizes and SHA-256 for every input.

Known in this release (documented, not hidden)

  • recent_share_2022plus and newest_share_2025_2026 are empty for all 5,000 rows: the recency
    axis never reached the feature table, so its 0.06 weight was inert
    (KI-2).
  • pos is a dictionary lookup, not corpus tagging (KI-3).
  • confidence for profile-derived labels is a constant cap, not a posterior
    (KI-5).
  • The upstream REPORT.md "modern additions, ranked by 2025-26 evidence" section is alphabetical
    (KI-1).
  • Pipeline source code and the anchor wordlist files are not in the release checkpoint
    (KI-12).

[1.0.0+repo] - 2026-09-21

Added

  • Public repository packaging of the release: data, docs, tooling.
  • data/SHA256SUMS.txt integrity ledger over 53 files, plus make checksums / make ledger.
  • data/lists/ per-level handouts (rank order and alphabetical).
  • data/derived/: candidate_features.csv.gz (215,416 candidates × the full feature vector),
    candidate_labels.csv.gz (levels, sources, confidences, class probabilities),
    surface_to_lemma.csv.gz (283,502 forms) — deterministic rebuild via make derived.
  • lexicore Python package: Entry model, CSV/JSON/TXT loaders, validate_data(),
    a re-scoring harness for the published formula, and a CLI (words, show, stats,
    validate, reweight).
  • 16 pytest cases: release contract + library behaviour, including loss-free CSV round-trip.
  • docs/: methodology, provenance (with live re-verification of source URLs/hashes), data
    dictionary, validation, known issues, reproduction guide, artefact ledger.
  • CI (GitHub Actions) running validate / test / lint on Python 3.10, 3.12, 3.13.
  • ecosystem/lexideck-multilingual/: the companion Anki deck (≈5,700 words = this list plus
    ~700 additions), its card model schema as measured from the free sample, and screenshots.

Changed

  • Upstream artefacts kept unmodified under reports/upstream/; only the files in data/ are
    renamed/documented by this repository.

Not changed (on purpose)

  • No word was added, removed or re-levelled. 1.0.0's content is exactly the released content.