-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings addedtoken
A token matched literally, before the model's vocabulary sees the text.
public sealed record AddedTokenProperties — Content and Id are constructor parameters: the exact string to match, and the id it maps to. Special
marks it as a control token rather than content. SingleWord requires the match to stand alone
rather than fall inside a word. Lstrip and Rstrip absorb whitespace to the left or right into
the match. Normalized says whether the normalizer runs over it first.
Example — a mask token, matched whole where the model would otherwise split it.
using Lodestar.Embeddings.Tokenization;
var token = new AddedToken("[MASK]", 103)
{
Special = true,
Lstrip = true,
};
string content = token.Content; // => [MASK]Remarks — added tokens exist because a vocabulary cannot express "this exact string is one
token, whatever my merge rules say". [MASK] would otherwise become [, MA, ##SK, ] and
mean nothing. They are matched before the sub-word algorithm runs, so they win over it.
Lstrip is the one that surprises: with it on, the space before [MASK] is absorbed into the
token string and disappears from the ids. That is BERT's own behaviour and it changes the token
text you see without changing the id count.
SingleWord is what keeps a token like <s> from matching inside a<s>b.
Applies to — net10.0, netstandard2.0.
See also — WordPieceVocabulary,
BpeVocabulary, the Python equivalence table.
| Member | What it does |
|---|---|
AddedToken.Equals |
Value equality over every flag. |
AddedToken.GetHashCode |
A hash consistent with it. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels