-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings sentencepiecemodelloader
Development build. This page describes
main, not a released package. The latest published Lodestar.Embeddings is 0.4.0 — read its documentation.
Reads a SentencePiece spiece.model — the trained unigram vocabulary, its scores, its piece types
and the model's special-token ids.
public static class SentencePieceModelLoaderExample — T5, ALBERT, camemBERT and XLM-R all ship this file.
using Lodestar.Embeddings.Persistence;
using Lodestar.Embeddings.Tokenization;
SentencePieceVocabulary vocab = SentencePieceModelLoader.Load("spiece.model");
var tokenizer = new SentencePieceTokenizer(vocab);Remarks — spiece.model is a protobuf, and it carries far more than a word list. It records
the type of every piece, so the tokenizer knows which entries are control markers instead of
guessing from their ids, and it carries the scores unigram Viterbi segmentation needs.
It also carries the normalizer, as a compiled character map. That map is read from the file and
never assumed to be identity — a stock model ships nmt_nfkc, and applying it is what makes
tokenization here match Python on the same text.
Because all of that is in the file, Load takes only bounds.
There is nothing left for a caller to get wrong.
Reference behaviour is sentencepiece.SentencePieceProcessor(model_file=…).
Applies to — net10.0, netstandard2.0.
See also — TokenizerJsonLoader,
ArtifactLoadOptions, the persistence index.
| Member | What it does |
|---|---|
SentencePieceModelLoader.Load |
Reads a spiece.model. |
SentencePieceModelLoader.LoadAsync |
The same, asynchronously. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels