-
Notifications
You must be signed in to change notification settings - Fork 0
Text spanishsnowballstemmer
Development build. This page describes
main, not a released package. The latest published Lodestar.Text is 0.4.0 — read its documentation.
Spanish stemming by the Snowball algorithm.
public static class SpanishSnowballStemmerExample — a noun's four forms on one key.
using Lodestar.Text.Stemming;
string masculine = SpanishSnowballStemmer.Stem("musico"); // => music
string feminine = SpanishSnowballStemmer.Stem("musica"); // => music
string plural = SpanishSnowballStemmer.Stem("musicos"); // => musicRemarks — Spanish attaches object pronouns to the end of a verb — dámelo, cantándome —
which would leave a suffix stemmer trimming the pronoun's ending instead of the verb's. The
algorithm therefore opens with a step 0 that removes those attached pronouns before any
ordinary suffix rule runs. It is the one structural difference from the French and Italian
stemmers.
Accents are stripped last, after the rules. dámelo and damelo both reach damel, so a
corpus that omits accents — as plenty of real Spanish text does — indexes to the same keys as one
that keeps them.
Verb conjugation is where the collapse is largest: the present, preterite, imperfect, conditional
and subjunctive forms of cantar all reduce to cant.
Reference behaviour is nltk.stem.snowball.SnowballStemmer("spanish"), matched over 127 words.
Applies to — net10.0, netstandard2.0.
See also — PortugueseSnowballStemmer, its nearest neighbour,
and the stemming index.
| Member | What it does |
|---|---|
SpanishSnowballStemmer.Stem |
The Snowball stem of one Spanish word. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels