-
Notifications
You must be signed in to change notification settings - Fork 0
Embeddings pooler l2normalize
Scales a vector in place to unit L2 norm.
public static void L2Normalize(Span<float> vector)Parameters — vector is the vector to scale, modified in place. A float[] converts implicitly.
Returns — nothing — the argument is the result.
Exceptions — none. A zero vector is left alone rather than producing NaN.
Example — the 3-4-5 triangle, scaled to unit length.
using Lodestar.Embeddings.Pooling;
float[] v = { 3f, 4f };
Pooler.L2Normalize(v);
float x = v[0]; // => 0.6
float y = v[1]; // => 0.8Remarks — This is the one place in the package that refuses to vectorize on purpose. The sum of squares
is accumulated in double and computed by a scalar loop: a Vector<float> accumulator would lose
precision, and it would make the answer depend on the SIMD width of the machine that ran it. The
scaling pass afterwards is exact whichever way it is done, so that half is vectorized.
The consequence is worth stating plainly: a pooled vector is bit-identical across
net10.0, netstandard2.0, and machines with different vector widths. That is the opposite trade
from VectorMath.Dot, where the SIMD accumulation is the point and the last bits may differ.
A zero vector is a no-op, not a NaN. That matters because an all-padding sequence pools to
zero, and a NaN there would spread into every score it ever touched.
Matches torch.nn.functional.normalize(v, p=2, dim=1).
Applies to — net10.0, netstandard2.0.
See also — Pooler.MeanPoolAndNormalize, Pooler,
the pooling index.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels