-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings embeddingindex load
Development build. This page describes
main, not a released package. The latest published Lodestar.Embeddings is 0.3.1 — read its documentation.
Reads an index back, ready to search without embedding the corpus again.
public static EmbeddingIndex Load(Stream source, ArtifactLoadOptions options = null)
public static EmbeddingIndex Load(string path, ArtifactLoadOptions options = null)Parameters — source is a readable stream, left open for the caller to dispose; path is the
file to read. options bounds what will be accepted and defaults to
ArtifactLoadOptions's own defaults.
Returns — EmbeddingIndex, with the same Dimension, Count, ids and normalization setting
it was saved with.
Exceptions — InvalidDataException when the content is not an embedding index, is of an
unsupported version, is internally inconsistent, holds a non-finite value, or exceeds a bound in
options.
Example — a saved index reloaded and queried.
using Lodestar.Embeddings.Search;
var original = new EmbeddingIndex(dimension: 2);
original.Add(new float[] { 1f, 0f }, "east");
original.Add(new float[] { 0f, 1f }, "north");
using var buffer = new MemoryStream();
original.Save(buffer);
buffer.Position = 0;
EmbeddingIndex reloaded = EmbeddingIndex.Load(buffer);
int size = reloaded.Count; // => 2
string top = reloaded.GetId(reloaded.Search(new float[] { 1f, 0f }, k: 1)[0].Index)!; // => eastRemarks — vectors are restored exactly as stored and never replayed through
Add. Re-normalizing an already normalized vector would move its bits,
and a reloaded index would then score slightly differently from the one that was saved.
The normalization flag travels in the file and cannot be supplied here. That is why neither overload takes one. An index built with normalization on and reloaded with it off would rank a corpus wrongly while looking entirely healthy, which is the class of bug a file format should make impossible rather than document.
options is what stands between a file and an allocation. Counts are bounded before they size
anything, and the vector block is capped in bytes by MaxTotalBytes before parsing begins —
an element-count limit sized for a vocabulary is orders of magnitude away from what a corpus of
embeddings needs. A file that exceeds a bound is refused, never truncated: a quietly smaller index
is a wrong answer.
Internal consistency is checked too. A file whose count and dimension do not account for the
number of values in its vector block is refused, as is one whose id array is a different length
from its count.
Applies to — net10.0, netstandard2.0.
See also — EmbeddingIndex.Save,
EmbeddingIndex.LoadAsync, EmbeddingIndex.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- 0051-the-save-paths-cost-is-the-buffer-not-the-encoding
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels