-
Notifications
You must be signed in to change notification settings - Fork 0
Text englishsnowballstemmer stem
Development build. This page describes
main, not a released package. The latest published Lodestar.Text is 0.4.0 — read its documentation.
The Porter2 stem of one English word.
public static string Stem(string word)Parameters — word is a single word, not a sentence. It is lowercased before anything else
runs, so the caller need not do it.
Returns — string, the stem, always lowercase. It is a key rather than a word: poni and
agre are both correct answers.
Exceptions — ArgumentNullException when word is null. An empty string is returned
unchanged, as is any word of two characters or fewer.
Example — a plural, a participle and a word the rules would over-trim.
using Lodestar.Text.Stemming;
string plural = EnglishSnowballStemmer.Stem("ponies"); // => poni
string participle = EnglishSnowballStemmer.Stem("agreed"); // => agre
string kept = EnglishSnowballStemmer.Stem("communism"); // => communismRemarks — communism is the one worth reading twice. The -ism suffix is real, and the
algorithm still declines to remove it, because doing so would cut into R1. That restraint is the
difference from PorterStemmer.Stem, which returns commun — a fragment short enough to be
reached from several unrelated words.
Words of two characters or fewer skip the algorithm entirely and come back lowercased. Nothing
sensible can be trimmed from an, and the rules would find a suffix anyway.
Call it once per token, after tokenizing — it has no notion of whitespace, and a whole sentence passed in is treated as one long word.
Applies to — net10.0, netstandard2.0.
See also — EnglishSnowballStemmer,
PorterStemmer.Stem, the stemming index.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels