-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings persistence
A tokenizer is only correct if its vocabulary is the model's own. Lodestar.Embeddings.Persistence
reads the four file formats models actually ship, and bounds what it will accept from them.
Nothing here is assembled by hand. The settings that change tokenization — whether the model was trained lowercased, what marks a continuation piece, which pieces are control markers, how a merge list is ordered — are read from the file wherever the file carries them, because a caller guessing one produces embeddings that do not match the model and look fine.
| The file the model ships | Loader | Produces |
|---|---|---|
vocab.txt (BERT) |
VocabTxtLoader |
WordPieceVocabulary |
spiece.model (SentencePiece) |
SentencePieceModelLoader |
SentencePieceVocabulary |
vocab.json + merges.txt (GPT-2) |
BpeFilesLoader |
BpeVocabulary |
tokenizer.json (HuggingFace) |
TokenizerJsonLoader |
any of the three |
Every loader has the same three shapes: Load(Stream), Load(string path) and an async
counterpart. A stream you pass in is never disposed for you.
The split is not arbitrary — it is whatever the format records.
vocab.txt is one token per line and nothing else, so
VocabTxtLoader.Load takes unkToken,
continuationPrefix and lowercase as parameters: the file cannot tell you them, and getting
lowercase wrong silently changes every embedding.
spiece.model and tokenizer.json carry their settings, so the loaders read them instead of
asking. That is why SentencePieceModelLoader.Load
takes only bounds — the piece types, the scores and the normalizer map are all in the file.
A vocabulary is something you downloaded, and every count it declares would otherwise size a
buffer. ArtifactLoadOptions is the ceiling on all five of
them, applied while reading rather than after. Exceeding one raises InvalidDataException
naming the limit and the value — never an OutOfMemoryException, which is the failure this type
exists to prevent.
This is a different type from Lodestar.Text.Persistence.ArtifactLoadOptions, which bounds a
saved vectorizer. The two are declared separately rather than shared;
decision 0011 has why, and the practical consequence
is that the defaults differ because what they bound differs.
Each tokenizer here implements one fixed pipeline, and a tokenizer.json describing another is
refused by name rather than loaded into an approximation of itself. Stock BERT is refused by
LoadWordPiece — its route is
VocabTxtLoader — and Llama-2 and Mistral v0.1 are refused by
LoadBpe for declaring byte_fallback.
A refusal is the correct outcome: the alternative is embeddings that do not match the model and carry nothing to say so.
| Type | What it is |
|---|---|
ArtifactLoadOptions |
The five bounds every load here is held to. |
BpeFilesLoader |
The vocab.json + merges.txt pair GPT-2 ships. |
SentencePieceModelLoader |
The trained spiece.model. |
TokenizerJsonLoader |
A HuggingFace tokenizer.json, whichever model it declares. |
VocabTxtLoader |
A BERT-style vocab.txt. |
- Embeddings, end to end — "Loading vocabularies", with the models that are refused.
- Python → C# equivalence — the loader rows, with the quirks reproduced.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels