-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings blocknormalization
Development build. This page describes
main, not a released package. The latest published Lodestar.Embeddings is 0.4.0 — read its documentation.
How a block handed to a bulk ingest relates to the index's normalization.
public enum BlockNormalization { Normalize, AlreadyNormalized, Off }Members — Normalize L2-normalizes every vector of the block and turns the index's
normalization on, which is what makes a dot product a cosine. AlreadyNormalized also turns it on
but stores the block bit for bit, on the caller's promise that its vectors are already unit length.
Off turns normalization off on insertion and on query, which is a raw dot product.
Example — the same block ingested two ways, and the scores that follow from it.
using Lodestar.Embeddings.Search;
float[] raw = { 3f, 4f };
var normalizing = EmbeddingIndex.FromBlock(raw, dimension: 2, BlockNormalization.Normalize);
var verbatim = EmbeddingIndex.FromBlock(raw, dimension: 2, BlockNormalization.Off);
float scored = normalizing.Search(new float[] { 3f, 4f }, 1)[0].Score; // => 1
float unscored = verbatim.Search(new float[] { 3f, 4f }, 1)[0].Score; // => 25Both were given (3, 4). The first normalized it to (0.6, 0.8) and normalizes the query the
same way, so a vector queried against itself scores 1 however long it was. The second compares
the two verbatim, and 3² + 4² is 25 — a number that grows with the length of whatever is being
compared, which is exactly what cosine exists to remove.
Remarks — one argument rather than a normalize flag beside an alreadyNormalized one. The
index's flag governs the query as well as the store, so a pair of booleans would make a fourth
combination representable that means nothing; three named answers make each one sayable and
nothing else.
Normalize is the zero value, so a default(BlockNormalization) reaching an ingest by accident
yields the correct-but-slower behaviour rather than a silently wrong score. That ordering is
deliberate and is the reason the enum is not alphabetical.
AlreadyNormalized is the one that can hurt. It is a promise, not a check: a block that is not
unit length taken this way is stored as it is, scored against a normalized query, and every
similarity comes back scaled by the vector's own length — a ranking that is wrong and looks
plausible. Reach for it when the vectors came out of a model that normalizes, or out of an index
that was saved normalized; otherwise Normalize costs one pass over the block and removes the
question.
A value outside the three — a cast from an int, most likely — is refused by both
EmbeddingIndex.FromBlock and
EmbeddingIndex.FromOwnedBlock rather than being read as
AlreadyNormalized, which is what the fall-through would otherwise have made it.
Applies to — net10.0, netstandard2.0.
See also — EmbeddingIndex.FromBlock,
EmbeddingIndex.FromOwnedBlock,
EmbeddingIndex, the search index.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- 0051-the-save-paths-cost-is-the-buffer-not-the-encoding
- 0052-pre-sizing-the-artifact-file-buys-nothing-on-a-delayed-allocation-filesystem
- 0053-the-payload-buffer-is-not-pooled-because-residency-outlives-the-load
- 0054-the-payload-buffer-is-pooled-after-all-because-the-collection-is-the-cost
- 0055-the-artifact-gets-a-binary-sidecar-once-a-block-can-be-ingested-whole
- 0056-a-block-may-be-adopted-and-the-invariant-is-the-callers-to-keep
- 0057-the-npy-read-serves-a-stream-and-a-buffer-differently
- 0058-the-npy-ingest-is-memcpy-bound-and-the-allocation-is-not-the-cost
- 0059-phase-0-verifications-two-confirmed-voids-do-not-survive-nuget
- 0060-tensorprimitives-beats-our-kernel-and-the-knn-is-still-not-redundant
- 0061-the-ingest-gap-was-a-collection-landing-wherever-the-collector-ran
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels