-
Notifications
You must be signed in to change notification settings - Fork 0
Text ratcliffobershelp similarity
Development build. This page describes
main, not a released package. The latest published Lodestar.Text is 0.4.0 — read its documentation.
Scores two strings as twice the total length of their recursively matched blocks, divided by the sum of their lengths.
public static double Similarity(ReadOnlySpan<char> a, ReadOnlySpan<char> b, TextElement element = TextElement.Utf16Unit)Parameters — a and b are the two strings to compare. element says what counts as one
character; difflib works on code points, so TextElement.CodePoint is what reproduces its
numbers
on supplementary-plane text.
Returns — double in [0, 1], larger meaning more alike. 1 for equal inputs, and 1 when
both are empty.
Example — the matched blocks are st and e: three characters, counted twice, over ten.
using Lodestar.Text.Distances;
double s = RatcliffObershelp.Similarity("state", "taste"); // => 0.6Remarks — this is the measure for longer text whose overlap comes in passages: it rewards
long unbroken runs and does not care how much unmatched material sits between them. It is exactly
difflib.SequenceMatcher(None, a, b).ratio(), so it is the port for anything written against
Python's standard library rather than against rapidfuzz.
The page's other recommendation for longer text is Indel, and the two are not interchangeable
even though they agree on plenty of pairs. The difference is contiguity: Indel credits every
character the two share in order however scattered, while this credits only characters inside a
shared unbroken run, and it commits greedily to the longest run before looking at what is left. On
("state", "taste") — the example above — that is 0.6 here against 0.8 from
Indel.NormalizedSimilarity, and on ("conversation", "voicesranton") it is 0.25 against
0.5833…. Reach for this when a long verbatim passage should count for more than the same number
of
characters sprinkled about, and for Indel when it should not.
Two things to know, and the first is the one that catches people. This measure is not
symmetric:
swapping the arguments can change the answer, sometimes by a lot. Similarity("bbcabba", "bacaa")
is 0.6666… and Similarity("bacaa", "bbcabba") is 0.3333…, because the recursion anchors on
the
longest matching block and difflib's tie-break — earliest start in a, then earliest in b — is
reproduced here, so a tie broken one way for (a, b) breaks the other way for (b, a). Fix an
argument order and keep it, or you will get two different scores for the same pair of records.
And on inputs longer than 200 elements it deliberately diverges from difflib's default. difflib
applies an autojunk heuristic there, ignoring any element that appears in more than 1% of
positions; this implementation does not, matching difflib(autojunk=False) at every length. The
reasoning is in decision 0006.
Applies to — net10.0, netstandard2.0.
See also — RatcliffObershelp.Distance, Lcs.SubstringLength, Indel.NormalizedSimilarity,
decision 0006,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels