-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings sentencepiecetokenizer
Development build. This page describes
main, not a released package. The latest published Lodestar.Embeddings is 0.4.0 — read its documentation.
Unigram encoding over a SentencePiece vocabulary — T5, ALBERT, XLM-R, camemBERT.
public sealed class SentencePieceTokenizer : ISubwordTokenizerConstructor — takes a SentencePieceVocabulary.
Example — two words, each one piece, each carrying its space.
using Lodestar.Embeddings.Tokenization;
SentencePiece[] pieces =
[
new SentencePiece("<unk>", 0.0, 0),
new SentencePiece("<s>", 0.0, 1),
new SentencePiece("▁alpha", -1.5, 2),
new SentencePiece("▁beta", -2.5, 3),
];
SentencePieceType[] types =
[
SentencePieceType.Unknown,
SentencePieceType.Control,
SentencePieceType.Normal,
SentencePieceType.Normal,
];
var vocabulary = new SentencePieceVocabulary(pieces, types, UnkId: 0, BosId: 1, EosId: -1, PadId: -1);
var tokenizer = new SentencePieceTokenizer(vocabulary);
TokenizationResult encoded = tokenizer.Encode("alpha beta");
int count = encoded.Tokens.Count; // => 2
string first = encoded.Tokens[0]; // => ▁alphaRemarks — no pre-tokenizer runs. The text is a stream, the space is encoded as ▁ inside the
pieces, and the segmentation is whichever one maximises the sum of the scores. That is the whole
difference from WordPiece, which splits on whitespace first and then matches greedily inside each
word.
The practical consequence: leading spaces matter. "alpha" and " alpha" can tokenize
differently, because one begins a word and the other continues the stream — and a model trained on
sentences expects the ▁.
Where the vocabulary carries a PrecompiledNormalizer, it runs first.
Applies to — net10.0, netstandard2.0.
See also — SentencePieceVocabulary,
ISubwordTokenizer.
| Member | What it does |
|---|---|
SentencePieceTokenizer.Encode |
Tokens and ids for one string. |
SentencePieceTokenizer.TryGetId |
The id of a piece, if the vocabulary holds it. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels