-
Notifications
You must be signed in to change notification settings - Fork 0
Text tfidfvectorizeroptions
Development build. This page describes
main, not a released package. The latest published Lodestar.Text is 0.4.0 — read its documentation.
The counting half and the weighting half, as one object.
public sealed record TfidfVectorizerOptionsProperties — Count is a CountVectorizerOptions, deciding what
counts as a term. Tfidf is a TfidfOptions, deciding how a count becomes a
weight. Both default to their own defaults, so new TfidfVectorizerOptions() is scikit-learn's
TfidfVectorizer().
Example — stop words removed before the weighting, which is the usual reason to reach for this.
using Lodestar.Text.Vectorization;
var options = new TfidfVectorizerOptions
{
Count = new CountVectorizerOptions { StopWords = StopWords.English },
Tfidf = new TfidfOptions { SublinearTf = true },
};
CsrMatrix weighted = new TfidfVectorizer(options).FitTransform(["the cat eats", "a dog eats"]);Remarks — the split into two records is the one place this surface reads differently from
scikit-learn, where TfidfVectorizer takes every keyword of CountVectorizer and every keyword of
TfidfTransformer in one flat constructor. Flattening them here would have meant one record with
thirteen properties and no way to hand the counting half to a
CountVectorizer or the weighting half to a
TfidfTransformer; keeping them apart means the same options object can be
used with any of the three.
The order the two halves apply in is fixed and worth stating: Count decides what a term is,
then Tfidf decides what it is worth. So a stop word removed by Count never reaches the
weighting, and a document frequency computed by the weighting is over the terms that survived.
Applies to — net10.0, netstandard2.0.
See also — TfidfVectorizer,
CountVectorizerOptions, TfidfOptions, the
Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels