-
Notifications
You must be signed in to change notification settings - Fork 0
Text tfidftransformer transform
Development build. This page describes
main, not a released package. The latest published Lodestar.Text is 0.4.0 — read its documentation.
Weight a count matrix using document frequencies already learned.
public CsrMatrix Transform(CsrMatrix counts)Parameters — counts is the matrix to weight. It must have the same number of columns as the
matrix that was fitted.
Returns — CsrMatrix, weighted and normalized, the same shape as counts.
Exceptions — InvalidOperationException when nothing has been fitted yet and
TfidfOptions.UseIdf is true. With UseIdf = false there is no document
frequency to be missing, so an unfitted transformer weights happily.
ArgumentNullException when counts is null. ArgumentException when the column count disagrees
with the fit.
Example — training frequencies applied to a later document.
using Lodestar.Text.Vectorization;
var cv = new CountVectorizer();
CsrMatrix training = cv.FitTransform(["the cat eats", "the dog eats"]);
var tfidf = new TfidfTransformer();
tfidf.Fit(training);
CsrMatrix weighted = tfidf.Transform(cv.Transform(["the cat"]));
int width = weighted.ColumnCount; // => 4The row length is 1 only up to floating point — weighting then normalizing lands a
hair under it, which is why the assertion above is on the shape rather than the norm.
Remarks — the column count is checked and a mismatch throws, which is the one shape error this type can catch. What it cannot catch is a matrix of the right width whose columns mean something else — two vectorizers fitted on different corpora can easily agree in width and disagree in every column. Reusing the vectorizer that produced the training counts is what avoids that, and there is no way for this method to verify it.
Applies to — net10.0, netstandard2.0.
See also — TfidfTransformer.Fit,
TfidfTransformer.FitTransform.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels