Skip to content

Fuzzy migrating from rapidfuzz

github-actions[bot] edited this page Aug 26, 2026 · 28 revisions

Development build. This page describes main, not a released package. The latest published Lodestar.Fuzzy is 0.4.0 — read its documentation.

Migrating from rapidfuzz — fuzzy matching

Lodestar.Fuzzy reproduces rapidfuzz.fuzz and rapidfuzz.process, plus a blocking deduplication.

dotnet add package Lodestar.Fuzzy

The fuzz.* ratios

All return a score in [0, 100]. Like rapidfuzz, no preprocessing by default (case-sensitive).

rapidfuzz Lodestar.Fuzzy
fuzz.ratio(a, b) Fuzz.Ratio(a, b)
fuzz.partial_ratio(a, b) Fuzz.PartialRatio(a, b)
fuzz.token_sort_ratio(a, b) Fuzz.TokenSortRatio(a, b)
fuzz.token_set_ratio(a, b) Fuzz.TokenSetRatio(a, b)
fuzz.WRatio(a, b) Fuzz.WRatio(a, b)
using Lodestar.Fuzzy;

Fuzz.Ratio("new york mets", "new york yankees");             // 65.0
Fuzz.TokenSortRatio("hello world", "world hello");           // 100.0 (order ignored)
Fuzz.PartialRatio("new york", "the wonderful new york mets");// 100.0 (substring)
Fuzz.WRatio("fuzzy wuzzy was a bear", "wuzzy fuzzy was a bear"); // 95.0

Pitfall #1: fuzz.ratio is not Levenshtein — it's the Indel similarity ×100. Lodestar.Fuzzy builds on Lodestar.Text's Indel.

Finding the best candidate — process

string[] choices = ["new york mets", "new york yankees", "boston red sox"];

// best candidates (default WRatio scorer), sorted, with a cutoff
IReadOnlyList<ExtractResult> top = Process.Extract("new york", choices, limit: 3, scoreCutoff: 50);

// the single best (or null)
ExtractResult? best = Process.ExtractOne("new york mets", choices);

Process.ExtractOne returns null when nothing clears scoreCutoff, which is the difference from rapidfuzz's extractOne worth knowing before porting a call that assumes a result.

Any scorer can be supplied: Process.Extract takes one — Process.Extract(q, choices, scorer: Fuzz.Ratio).

Deduplicating records (with blocking)

To avoid quadratic comparison, first partition by a blocking key (initial, Soundex code, postal code…), then compare only within each block. That is what Deduplicator.FindClusters does, and the key is the caller's to choose: two records in different blocks are never compared, whatever their similarity.

string[] records = ["John Smith", "Jon Smith", "Jane Doe", "Jayne Doe", "Bob Brown"];

IReadOnlyList<IReadOnlyList<int>> clusters = Deduplicator.FindClusters(
    records,
    blockingKey: r => r[..1],                  // block = first letter
    similarity: Fuzz.TokenSetRatio,
    threshold: 80);

// clusters: { {0,1}, {2,3}, {4} } — indices of duplicates grouped together

Grouping is the transitive closure (union-find): if ab and bc, all three are in the same cluster. Accepted trade-off: two true duplicates in different blocks are never compared (recall ↓, speed ↑).

Lodestar

Project

Clone this wiki locally