-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings search
Embedding a corpus is the expensive half and it happens once. What comes after is cheap and constant: given a query vector, which of the stored vectors point most nearly the same way.
Lodestar.Embeddings.Search answers that with an exhaustive index — every stored vector is
scored on every query — plus the two SIMD primitives it is built on.
flowchart TD
A["How many vectors?"] --> B{"Up to a few<br/>hundred thousand?"}
B -->|yes| C["EmbeddingIndex — exhaustive,<br/>exact, nothing to tune"]
B -->|"more"| D["An approximate index (HNSW).<br/>Not in this package."]
Scoring every vector is linear, and a SIMD dot product makes the constant small enough that the crossover with an approximate index sits far higher than most corpora ever reach. The trade the approximate structures make — recall for speed, plus parameters to tune and a graph to build — is not worth taking before the linear scan is actually the bottleneck.
The consequence is that EmbeddingIndex.Search is exact.
There is no recall parameter, because nothing is skipped.
An index normalizes on insertion by default, and normalizes the query too. Once both sides are
unit vectors, cosine similarity is the dot product — so the hot loop is
VectorMath.Dot and nothing else.
using Lodestar.Embeddings.Search;
var index = new EmbeddingIndex(dimension: 2);
index.Add(new float[] { 1f, 0f });
index.Add(new float[] { 0f, 1f });
IReadOnlyList<SearchResult> hits = index.Search(new float[] { 2f, 0f }, k: 1);
int best = hits[0].Index; // => 0
float score = hits[0].Score; // => 1The query was (2, 0) and the score is 1: length was normalized away on both sides, which is
the point of cosine and the reason a query need not be scaled by the caller.
Search returns
SearchResult — a position and a score, and deliberately not your
document. The id is fetched separately with GetId, so the
scored array stays a block of 8-byte structs the garbage collector never has to look inside.
Save writes the index and
Load reads it back, vectors restored bit for bit rather than
re-normalized. Two things about that file are worth knowing before relying on it:
- The normalization flag travels in the file and cannot be supplied on load. An index reloaded under the other setting would rank a corpus wrongly and never look wrong.
-
A vector holding
NaNor an infinity can be added but cannot be saved. The refusal is deliberate —Addhas the reasoning.
| Type | What it is |
|---|---|
EmbeddingIndex |
The exhaustive cosine index: add, search, save, load. |
SearchResult |
One hit — a position and a score. |
VectorMath |
The two SIMD primitives the index is built on. |
- Embeddings, end to end — "Index a corpus and query it".
- Python → C# equivalence — what this replaces on the numpy and faiss side.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels