-
Notifications
You must be signed in to change notification settings - Fork 0
Text tfidfvectorizer
CountVectorizer and TfidfTransformer in one pass
— the equivalent of sklearn.feature_extraction.text.TfidfVectorizer.
Counts the terms, then weights each count by how rare the term is across the corpus, so that words appearing everywhere stop dominating the vectors.
public sealed class TfidfVectorizerConstructor — TfidfVectorizer(TfidfVectorizerOptions? options = null), whose two halves
default to scikit-learn's defaults.
Properties — Idf is the inverse document frequency per column, available after fitting.
Example — the same corpus as CountVectorizer, weighted.
using Lodestar.Text.Vectorization;
string[] docs = ["the cat eats", "the dog eats", "the cat and the dog"];
CsrMatrix weighted = new TfidfVectorizer().FitTransform(docs);
// Rows come out L2-normalized, so every row's length is 1.
double rowLength = weighted.RowL2Norm(0); // => 1Remarks — this is the vectorizer to reach for by default. Raw counts make long documents look important and common words look meaningful; TF-IDF fixes both, and the L2 normalization that follows is what makes two documents of different lengths comparable.
It is exactly the two other types composed, and the composition is the only difference. Where
counts already exist, TfidfTransformer weights them without re-reading
text. Where the vocabulary must not be held in memory,
HashingVectorizer gives that up instead.
the appears in all three documents above and still has a non-zero weight, which surprises
readers: with SmoothIdf on, the IDF of a ubiquitous term is log(1) + 1, not 0. See
TfidfOptions.
Applies to — net10.0, netstandard2.0.
See also — TfidfVectorizerOptions,
TfidfTransformer, CountVectorizer, the
vectorization guide.
| Member | What it does |
|---|---|
TfidfVectorizer.Fit |
Learn the vocabulary and the document frequencies. |
TfidfVectorizer.FitTransform |
Learn them and weight the same corpus. |
TfidfVectorizer.GetFeatureNames |
The term each column stands for. |
TfidfVectorizer.Load |
Read a fitted vectorizer back. |
TfidfVectorizer.LoadAsync |
The same, without blocking. |
TfidfVectorizer.Save |
Write a fitted vectorizer out. |
TfidfVectorizer.SaveAsync |
The same, without blocking. |
TfidfVectorizer.Transform |
Weight a corpus against what was learned. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels