-
Notifications
You must be signed in to change notification settings - Fork 0
Text frenchsnowballstemmer
French stemming by the Snowball algorithm.
public static class FrenchSnowballStemmerExample — a masculine, a feminine and an adverb, all one key.
using Lodestar.Text.Stemming;
string masculine = FrenchSnowballStemmer.Stem("heureux"); // => heureux
string feminine = FrenchSnowballStemmer.Stem("heureuse"); // => heureux
string adverb = FrenchSnowballStemmer.Stem("heureusement"); // => heureuxRemarks — French inflects heavily at the end of the word, which is exactly what a suffix
stemmer is good at: national, nationale and nationaux all reduce to national, and the
-issait/-issant forms of second-group verbs reduce with their infinitive.
The algorithm uses RV alongside R1 and R2 — a region defined from the start of the word rather than from a suffix, which is what lets it protect short stems that the R1/R2 pair alone would eat. Six steps run in order, and which of them run at all depends on whether the previous one changed the word.
Input is normalised to NFC before the rules see it, so é written as a combining accent behaves
like é written as one character.
Reference behaviour is nltk.stem.snowball.SnowballStemmer("french"), matched over 152 words.
Applies to — net10.0, netstandard2.0.
See also — the stemming index,
ItalianSnowballStemmer and
SpanishSnowballStemmer, which share the Romance region scheme.
| Member | What it does |
|---|---|
FrenchSnowballStemmer.Stem |
The Snowball stem of one French word. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels