-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings onnx
Development build. This page describes
main, not a released package. The latest published Lodestar.Embeddings is 0.4.0 — read its documentation.
One type, OnnxTextEmbedder: it runs a sentence-transformer model and
gives you a vector per text. It is the only place in Lodestar where a model file is required,
and the only place ONNX Runtime is referenced — that dependency is deliberately confined to this
namespace so the rest of the package has none.
Weights are never committed to this repository. A running example would need a model of tens of
megabytes, and decisions/0003 rules that
out; tools/fetch_*.py pulls vocabularies against a pinned SHA-256 when they are needed, and
weights are not among them.
So the fences on these pages compile against the packed package and are marked
docs-run: skip, which is what that marker is for. The same exclusion is declared in the
packaging sample, where OnnxTextEmbedder is one of its two documented exclusions.
This namespace produces vectors. Turning text into the token ids it wants is
Lodestar.Embeddings.Tokenization; reducing a sequence of vectors to one is
Lodestar.Embeddings.Pooling; searching a set of them is
Lodestar.Embeddings.Search. This page is the middle step of four, and the only one
that needs a file from outside. The other three are named without links here because their pages
do not exist yet — they arrive with #226, #233 and #231, and a link written ahead of its target is
a link that is broken until then.
| Type | What it is |
|---|---|
OnnxTextEmbedder |
Runs an ONNX sentence-transformer and returns vectors. |
- Semantic search with embeddings — the guide, end to end.
- Python → C# equivalence.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels