-
Notifications
You must be signed in to change notification settings - Fork 0
Text germansnowballstemmer
Development build. This page describes
main, not a released package. The latest published Lodestar.Text is 0.4.0 — read its documentation.
German stemming by the Snowball algorithm.
public static class GermanSnowballStemmerExample — a noun's case endings, collapsed.
using Lodestar.Text.Stemming;
string nominative = GermanSnowballStemmer.Stem("kind"); // => kind
string plural = GermanSnowballStemmer.Stem("kinder"); // => kind
string dative = GermanSnowballStemmer.Stem("kindern"); // => kindRemarks — German differs from the Romance stemmers here in three ways worth knowing before reading a stem.
ß becomes ss first, so the two spellings of a word cannot produce two keys: straße and
strasse both stem to strass. Umlauts are removed at the end, once the suffix rules have
run, which is why größe comes back as gross rather than größ.
And there is no RV region — the algorithm uses R1 and R2 only, with R1 floored at three characters. That floor is what stops short words from being trimmed to nothing, the job RV does in the Romance languages.
Compound nouns are not split. Donaudampfschiff is one word to this algorithm and comes back as
one stem; splitting German compounds needs a dictionary, and none ships here.
Reference behaviour is nltk.stem.snowball.SnowballStemmer("german"), matched over 88 words.
Applies to — net10.0, netstandard2.0.
See also — the stemming index.
| Member | What it does |
|---|---|
GermanSnowballStemmer.Stem |
The Snowball stem of one German word. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels