-
Notifications
You must be signed in to change notification settings - Fork 0
Text 0.4.0 stopwords
Lodestar.Text 0.4.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
The six built-in stop-word lists, one per language with a Snowball stemmer.
public static class StopWordsProperties — English, French, German, Italian, Portuguese and Spanish, each an
IReadOnlyCollection<string> of lowercase words — a collection rather than a set because
IReadOnlySet<T> does not exist on netstandard2.0, which this package still targets. They are frozen sets built once and shared, so reading
one allocates nothing and the same instance comes back every time.
Example — the English list, passed where a vectorizer expects one.
using Lodestar.Text.Vectorization;
int english = StopWords.English.Count; // => 318
var cv = new CountVectorizer(new CountVectorizerOptions { StopWords = StopWords.English });
CsrMatrix counts = cv.FitTransform(["the cat eats", "a dog eats"]);Remarks — a stop word is removed after tokenization and lowercasing, so a list of
lowercase words is all that is needed however the document was written. Removal happens before
n-grams are formed, which is why NgramRange = (2, 2) over a corpus with stop words removed
produces bigrams of words that were never adjacent in the source.
The English list is scikit-learn's 'english' exactly, which is worth knowing because that list
is documented by scikit-learn itself as problematic — it is aggressive,
it removes words that carry meaning in some domains, and it exists mostly for parity. The other
five have no scikit-learn counterpart at all: stop_words='english' is the only built-in list
there, so those five are additions rather than reproductions.
Nothing stops a caller passing its own words instead; the option takes any
IReadOnlyCollection<string>.
Applies to — net10.0, netstandard2.0.
See also — CountVectorizerOptions,
CountVectorizer, the
Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels