-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings vocabtxtloader load
Reads a vocab.txt into a WordPiece vocabulary.
public static WordPieceVocabulary Load(Stream source, ArtifactLoadOptions options = null, string unkToken = "[UNK]", string continuationPrefix = "##", bool lowercase = false)
public static WordPieceVocabulary Load(string path, ArtifactLoadOptions options = null, string unkToken = "[UNK]", string continuationPrefix = "##", bool lowercase = false)Parameters — source is a readable stream, never disposed here; path is the file to read. options bounds what will be accepted and defaults to
ArtifactLoadOptions's own defaults.
unkToken is the piece an unknown word maps to, continuationPrefix marks a word-internal
piece, and lowercase says whether the model was trained on lowercased text.
Returns — WordPieceVocabulary, ids assigned by line number: the first line is id 0.
Exceptions — ArgumentNullException for a null source or path. InvalidDataException
when the content is not the format expected, declares a model this loader does not read, or
exceeds a bound in options — the message names both the limit and the value.
Example — an uncased BERT checkpoint, which is the case lowercase exists for.
using Lodestar.Embeddings.Persistence;
using Lodestar.Embeddings.Tokenization;
WordPieceVocabulary vocab = VocabTxtLoader.Load("bert-base-uncased/vocab.txt", lowercase: true);Remarks — lowercase is the parameter to get right, and the file cannot tell you its value. Loading an
uncased checkpoint without it leaves every capitalised word mapping to the unknown piece, which
does not throw, does not look wrong, and produces embeddings that are quietly meaningless. The
name of the checkpoint is usually the only evidence — bert-base-uncased against
bert-base-cased.
The id is the line number, so the file's order is the vocabulary's order and a reordered
vocab.txt is a different vocabulary. Two quirks of the Python loop this matches are reproduced
deliberately rather than corrected; docs/equivalence.md's loader row names them.
The defaults [UNK] and ## are BERT's. A model using other markers has to say so here, because
vocab.txt records neither.
Applies to — net10.0, netstandard2.0.
See also — VocabTxtLoader, VocabTxtLoader.LoadAsync,
the persistence index.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels