-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings 0.4.0 sentencepiecemodelloader
Lodestar.Embeddings 0.4.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
Reads a SentencePiece spiece.model — the trained unigram vocabulary, its scores, its piece types
and the model's special-token ids.
public static class SentencePieceModelLoaderExample — T5, ALBERT, camemBERT and XLM-R all ship this file.
using Lodestar.Embeddings.Persistence;
using Lodestar.Embeddings.Tokenization;
SentencePieceVocabulary vocab = SentencePieceModelLoader.Load("spiece.model");
var tokenizer = new SentencePieceTokenizer(vocab);Remarks — spiece.model is a protobuf, and it carries far more than a word list. It records
the type of every piece, so the tokenizer knows which entries are control markers instead of
guessing from their ids, and it carries the scores unigram Viterbi segmentation needs.
It also carries the normalizer, as a compiled character map. That map is read from the file and
never assumed to be identity — a stock model ships nmt_nfkc, and applying it is what makes
tokenization here match Python on the same text.
Because all of that is in the file, Load takes only bounds.
There is nothing left for a caller to get wrong.
Reference behaviour is sentencepiece.SentencePieceProcessor(model_file=…).
Applies to — net10.0, netstandard2.0.
See also — TokenizerJsonLoader,
ArtifactLoadOptions, the persistence index.
| Member | What it does |
|---|---|
SentencePieceModelLoader.Load |
Reads a spiece.model. |
SentencePieceModelLoader.LoadAsync |
The same, asynchronously. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels