-
Notifications
You must be signed in to change notification settings - Fork 0
Text 0.4.0 hashingvectorizeroptions
Lodestar.Text 0.4.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
How many columns the hash lands in, and what it does with signs.
public sealed record HashingVectorizerOptionsProperties — NumFeatures (default 1048576, which is 2^20) is how many columns the matrix
has, and therefore how often two different terms collide into one. AlternateSign (default
true) gives half the terms a negative weight, so that collisions tend to cancel rather than
accumulate. Norm (default SparseNorm.L2) is the norm each row is divided by, and is nullable: null leaves the rows unnormalized, which is scikit-learn's norm=None. Count carries
the tokenization settings — the token pattern, lowercasing, n-gram range and the rest — because
hashing changes only what happens after a document is cut into terms.
Example — a deliberately tiny feature space, so collisions are certain.
using Lodestar.Text.Vectorization;
var hv = new HashingVectorizer(new HashingVectorizerOptions { NumFeatures = 16 });
CsrMatrix hashed = hv.Transform(["the cat eats", "the dog eats", "the cat and the dog"]);
int columns = hashed.ColumnCount; // => 16
double rowLength = hashed.RowL2Norm(0); // => 1Remarks — NumFeatures is the whole trade. Large enough and collisions are rare and the matrix
is wide; small enough and the matrix is compact and two unrelated terms share a column. The default
of 2^20 is scikit-learn's n_features=2**20, chosen so that collisions are negligible for most
corpora while the matrix stays sparse — the columns cost nothing until something lands in them.
AlternateSign is the part that looks like a trick and is not. When two terms collide, one of them
carrying a negative sign means the collision partly cancels instead of doubling, so the inner
product between two documents stays close to what it would have been without the collision. It is
alternate_sign=True in scikit-learn and on by default for the same reason.
There is no MinDf or MaxDf here, and there cannot be: both need document frequencies, and
counting those would mean the pass over the corpus that this vectorizer exists to avoid.
Applies to — net10.0, netstandard2.0.
See also — HashingVectorizer,
CountVectorizerOptions, SparseNorm, the
Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels