-
Notifications
You must be signed in to change notification settings - Fork 0
Text 0.4.0 countvectorizer
Lodestar.Text 0.4.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
Term counts over a vocabulary learned from the corpus — the equivalent of
sklearn.feature_extraction.text.CountVectorizer.
Fitting learns which terms exist and fixes their column order; transforming counts them. Column 7
means the same term in every row, and GetFeatureNames can
say which.
public sealed class CountVectorizerConstructor — CountVectorizer(CountVectorizerOptions? options = null). The default options
are scikit-learn's defaults, so new CountVectorizer() is CountVectorizer().
Example — three documents, and the vocabulary they produce.
using Lodestar.Text.Vectorization;
string[] docs = ["the cat eats", "the dog eats", "the cat and the dog"];
var cv = new CountVectorizer();
CsrMatrix counts = cv.FitTransform(docs);
int features = counts.ColumnCount; // => 5Remarks — the vocabulary is sorted, as scikit-learn's is, so the column order depends only on the terms and not on the order the documents arrived in. Two fits over the same corpus give the same matrix, and a corpus shuffled before fitting gives the same matrix too.
What decides the counts is CountVectorizerOptions rather than
anything here — the token pattern that drops single-letter words, the stop words, the n-gram
range. This class only applies them.
To weight rare terms above common ones, TfidfVectorizer is this followed
by TfidfTransformer. To skip the vocabulary entirely,
HashingVectorizer trades naming for not having to hold one.
Applies to — net10.0, netstandard2.0.
See also — CountVectorizerOptions, CsrMatrix,
TfidfVectorizer, the
vectorization guide.
| Member | What it does |
|---|---|
CountVectorizer.Fit |
Learn the vocabulary, and return the same instance. |
CountVectorizer.FitTransform |
Learn it and count the same corpus in one pass. |
CountVectorizer.GetFeatureNames |
The term each column stands for, in column order. |
CountVectorizer.Load |
Read a fitted vectorizer back. |
CountVectorizer.LoadAsync |
The same, without blocking. |
CountVectorizer.Save |
Write a fitted vectorizer out. |
CountVectorizer.SaveAsync |
The same, without blocking. |
CountVectorizer.Transform |
Count a corpus against the vocabulary already learned. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels