-
Notifications
You must be signed in to change notification settings - Fork 0
Text artifactloadoptions
The bounds a load is held to, so a file cannot ask for more memory than you meant to give it.
public sealed record ArtifactLoadOptionsProperties — MaxVocabularySize (default 1_000_000) is how many vocabulary entries will be
accepted. MaxTokenLength (default 1024) is the longest single token, in characters.
MaxJsonDepth (default 32) is how deeply the document may nest. MaxTotalBytes (default 256
MiB) is how much will be read from the source in total. MaxArrayLength (default 1_000_000) is
the longest single JSON array.
Example — a stricter set than the defaults, for a file from somewhere untrusted.
using Lodestar.Text.Persistence;
using Lodestar.Text.Vectorization;
var strict = new ArtifactLoadOptions
{
MaxVocabularySize = 50_000,
MaxTotalBytes = 8L * 1024 * 1024,
};
var cv = new CountVectorizer();
cv.Fit(["the cat eats", "the dog eats"]);
using var buffer = new MemoryStream();
cv.Save(buffer);
buffer.Position = 0;
CountVectorizer restored = CountVectorizer.Load(buffer, strict);
int columns = restored.Transform(["the cat"]).ColumnCount; // => 4Remarks — the bounds are checked as the content is read, not after, so an oversized file is refused before it is allocated rather than afterwards. That ordering is the whole point: a check that runs once the array exists has already lost.
Every default is generous enough that a real model never meets one — a million vocabulary entries is far past any corpus this package is likely to see — so tightening them is a decision about the source, not about the model. Tighten when the file came from a user, a network, or a build you do not control; leave them when it came from your own training run.
Exceeding a bound raises InvalidDataException, and the artifact is refused rather than
truncated. A model that quietly loaded smaller than it was saved would score differently and give
no sign, which is worse than a failure.
This type is declared separately from Lodestar.Embeddings's namesake rather than shared, so that
neither package depends on the other for its loading contract;
decisions/0011 has the reasoning, along with
the comparison to pickle.load that motivates bounding at all.
Applies to — net10.0, netstandard2.0.
See also — CountVectorizer.Load,
TfidfVectorizer.Load,
HashingVectorizer.Load, the
vectorization guide.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels