Skip to content

Embeddings persistence

github-actions[bot] edited this page Aug 22, 2026 · 23 revisions

Loading vocabularies — Lodestar.Embeddings.Persistence

A tokenizer is only correct if its vocabulary is the model's own. Lodestar.Embeddings.Persistence reads the four file formats models actually ship, and bounds what it will accept from them.

Nothing here is assembled by hand. The settings that change tokenization — whether the model was trained lowercased, what marks a continuation piece, which pieces are control markers, how a merge list is ordered — are read from the file wherever the file carries them, because a caller guessing one produces embeddings that do not match the model and look fine.

Which loader?

The file the model ships Loader Produces
vocab.txt (BERT) VocabTxtLoader WordPieceVocabulary
spiece.model (SentencePiece) SentencePieceModelLoader SentencePieceVocabulary
vocab.json + merges.txt (GPT-2) BpeFilesLoader BpeVocabulary
tokenizer.json (HuggingFace) TokenizerJsonLoader any of the three

Every loader has the same three shapes: Load(Stream), Load(string path) and an async counterpart. A stream you pass in is never disposed for you.

What the file carries, and what stays a parameter

The split is not arbitrary — it is whatever the format records.

vocab.txt is one token per line and nothing else, so VocabTxtLoader.Load takes unkToken, continuationPrefix and lowercase as parameters: the file cannot tell you them, and getting lowercase wrong silently changes every embedding.

spiece.model and tokenizer.json carry their settings, so the loaders read them instead of asking. That is why SentencePieceModelLoader.Load takes only bounds — the piece types, the scores and the normalizer map are all in the file.

A file is untrusted until it has been bounded

A vocabulary is something you downloaded, and every count it declares would otherwise size a buffer. ArtifactLoadOptions is the ceiling on all five of them, applied while reading rather than after. Exceeding one raises InvalidDataException naming the limit and the value — never an OutOfMemoryException, which is the failure this type exists to prevent.

This is a different type from Lodestar.Text.Persistence.ArtifactLoadOptions, which bounds a saved vectorizer. The two are declared separately rather than shared; decision 0011 has why, and the practical consequence is that the defaults differ because what they bound differs.

Refusing a model is a feature

Each tokenizer here implements one fixed pipeline, and a tokenizer.json describing another is refused by name rather than loaded into an approximation of itself. Stock BERT is refused by LoadWordPiece — its route is VocabTxtLoader — and Llama-2 and Mistral v0.1 are refused by LoadBpe for declaring byte_fallback.

A refusal is the correct outcome: the alternative is embeddings that do not match the model and carry nothing to say so.

Types

Type What it is
ArtifactLoadOptions The five bounds every load here is held to.
BpeFilesLoader The vocab.json + merges.txt pair GPT-2 ships.
SentencePieceModelLoader The trained spiece.model.
TokenizerJsonLoader A HuggingFace tokenizer.json, whichever model it declares.
VocabTxtLoader A BERT-style vocab.txt.

See also

Lodestar

Project

Clone this wiki locally