-
Notifications
You must be signed in to change notification settings - Fork 0
Text 0.4.0 distances
Lodestar.Text 0.4.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
How different are two pieces of text? Every type on this page answers that, and they disagree on what "different" means: some count the edits that turn one string into the other, some check how many characters line up in roughly the same place, and one looks for the longest stretches the two have in common. Picking the wrong one is the usual cause of a similarity score that looks nothing like what a reader expects.
Two conventions run through the whole namespace, and knowing them saves reading every entry.
- A distance counts how far apart two inputs are:
0means identical, and a bigger number is a worse match. A similarity runs the other way,1meaning identical. What decides whether a number can be compared across pairs of different lengths is its type, not its name: aDistancereturningintcounts edits and has no upper bound, while every member returningdoubleon this page — theNormalized…ones and alsoJaro,JaroWinklerandRatcliffObershelp— is already scaled to[0, 1]and is comparable. - Every member that takes a
stringalso takes aTextElementsaying what counts as one character. (The generic overloads —Distance<T>,SubsequenceLength<T>,SubstringLength<T>— do not: they compare whatever elements you hand them.) The default,TextElement.Utf16Unit, is .NET's own unit and gives the same answer as Python for every character in the Basic Multilingual Plane. Outside it — emoji, rare ideographs — one character is two UTF-16 units, and the two disagree on purpose; passTextElement.CodePointfor Python's answer. The reasoning is in decision 0002.
Comparing two bags of words or characters, where position does not matter at all, is a
different question. It is answered by the Lodestar.Text.Similarity namespace —
Jaccard, SorensenDice, Overlap, Tversky and Cosine — not by anything here.
flowchart TD
A["What are you comparing?"] --> B["Two short strings:<br/>names, codes, typos"]
A --> C["Two longer texts"]
A --> D["Two bags of words<br/>or characters"]
B --> E{"Do the two line up<br/>position by position?"}
E -->|yes| F["Hamming"]
E -->|no| G{"Is agreement on the first<br/>few letters strong evidence?"}
G -->|yes| H["JaroWinkler"]
G -->|no| I{"Are swapped neighbours<br/>a common mistake?"}
I -->|yes| J["DamerauLevenshtein,<br/>or Osa when speed matters more"]
I -->|no| K{"Do you want a count of edits,<br/>or a forgiving score?"}
K -->|a count| L["Levenshtein"]
K -->|a score| M["Jaro"]
C --> N{"Do you want a score,<br/>or the shared text itself?"}
N -->|the text itself| O["Lcs"]
N -->|a score| P{"Does the shared material come in a<br/>few long passages, or scattered?"}
P -->|long passages| Q["RatcliffObershelp"]
P -->|scattered| R["Indel"]
D --> S["Not here — see<br/>Lodestar.Text.Similarity"]
| Type | What it measures |
|---|---|
DamerauLevenshtein |
Insertions, deletions, substitutions and swaps of neighbouring characters, with no limit on re-editing a stretch. |
Hamming |
How many positions hold a different character, plus the difference in length. |
Indel |
Insertions and deletions only, never substitutions — the basis of rapidfuzz's fuzz.ratio. |
Jaro |
How many characters the two share near the same position, and how many of those arrive out of order. |
JaroWinkler |
Jaro, raised for pairs that already agree on their first few characters. |
Lcs |
The length of the longest run the two have in common, contiguous or not. |
Levenshtein |
Insertions, deletions and substitutions. |
Osa |
The same as DamerauLevenshtein, except that no stretch of text may be edited twice. |
RatcliffObershelp |
How much of the two texts their matching blocks cover, taken longest first. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels