-
-
Notifications
You must be signed in to change notification settings - Fork 13
Benchmarks
GoldenMatch is benchmarked against the University of Leipzig entity resolution benchmark datasets.
| Dataset | Records | Best Strategy | Precision | Recall | F1 | Time |
|---|---|---|---|---|---|---|
| DBLP-ACM | 2.6K vs 2.3K | Vertex AI embeddings | 97.5% | 97.3% | 97.4% | 119s |
| DBLP-Scholar | 2.6K vs 64K | multi-pass + fuzzy | 67.2% | 84.1% | 74.7% | 83.9s |
| Abt-Buy | 1K vs 1K | Vertex AI embeddings | 85.5% | 83.9% | 84.7% | 53s |
| Amazon-Google | 1.4K vs 3.2K | Vertex AI embeddings | 60.6% | 56.8% | 58.6% | 110s |
Previous bests (without Vertex AI): DBLP-ACM 97.2% (multi-pass), Abt-Buy 59.5% (LLM boost), Amazon-Google 40.5% (rec_emb). Vertex AI's text-embedding-004 provides dramatically better embeddings for product matching.
| Tool | DBLP-ACM | Abt-Buy | Approach | Training Required |
|---|---|---|---|---|
| GoldenMatch | 97.4% | 84.7% | Vertex AI embeddings (zero-config) | No |
| Ditto | 99.0% | 89.3% | Fine-tuned DistilBERT | Yes (1000+ labels) |
| DeepMatcher | 98.4% | 62.8% | Deep learning | Yes |
| Splink | ~95% | ~70% | Fellegi-Sunter (Spark) | Yes (labels) |
| dedupe | ~96% | ~75% | Active learning | Yes (200+ labels) |
| Zingg | ~96% | ~80% | Active learning (Spark) | Yes (labels) |
- DBLP-ACM (97.4%): Within 1.6pts of Ditto with zero training — competitive with state-of-the-art.
-
Abt-Buy (84.7%): Vertex AI's
text-embedding-004closed most of the gap with Ditto (89.3%). Previously 59.5% with LLM boost on local MiniLM. - Amazon-Google (58.6%): 45% relative improvement over previous best (40.5%). Product matching with very different naming conventions remains hard across all tools.
- DBLP-Scholar (74.7%): Multi-pass blocking + fuzzy scoring. Not yet tested with Vertex AI.
See Comparison with Other Tools for a full feature-by-feature breakdown.
Measured on a laptop (17GB RAM, no GPU) with exact + fuzzy matching, blocking, clustering, and golden record generation:
| Records | Time | Throughput | Pairs Found | Memory | Notes |
|---|---|---|---|---|---|
| 1,000 | 0.2s | 5,500 rec/s | 210 | 101 MB | laptop |
| 10,000 | 1.4s | 7,300 rec/s | 7,000 | 123 MB | laptop |
| 100,000 | 12s | 8,200 rec/s | 571,000 | 544 MB | laptop |
| 1,000,000 | ~43 min | ~390 rec/s | 836K clusters | 9.98 GB |
ubuntu-latest 4c/16GB (v1.x scale audit) |
| 5,000,000 | 9.94 min | ~8,400 rec/s | 1.67M multi-member clusters | 6.4 GB |
large-new-64GB 16c/64GB, backend="bucket" (v1.16.0) |
Near-linear scaling: throughput stays consistent as data grows. Memory usage scales linearly.
v1.16.0 5M-on-one-node breakthrough. The bucket backend completes 5M dedupe in under 10 minutes at 6.4 GB peak RSS where the pre-v1.16 chunked path was hanging at 63 GB plateau on the same fixture. Pass backend="bucket" or let the v3 planner auto-pick it on 16-core / 32+ GB Linux nodes. See CHANGELOG.md and examples/at_scale_bucket_backend.py.
| Stage | Time | % of Total |
|---|---|---|
| Fuzzy matching | 5.6s | 45% |
| Golden records | 4.0s | 33% |
| Blocking | 2.0s | 16% |
| Clustering | 0.5s | 4% |
| Auto-fix + standardize | 0.1s | 1% |
| Matchkeys + exact match | 0.1s | 1% |
The bottleneck is fuzzy NxN scoring within blocks (RapidFuzz cdist). Coarser blocking keys = faster but lower recall. Fine blocking keys = slower but higher recall.
With exact matching only (no fuzzy), 1M records process in ~15 seconds:
138,730 duplicate clusters found with 100% precision and 100% recall.
For datasets exceeding available memory, use database sync mode:
goldenmatch sync --table customers --connection-string "$DB" --config config.yamlProcesses in chunks, maintains persistent ANN index, matches incrementally. Tested to 10M+ records in Postgres.
| Tool | Throughput | Hardware Required |
|---|---|---|
| GoldenMatch | 8,200 rec/s | Laptop (no GPU) |
| dedupe | ~500 rec/s | Laptop |
| Splink | ~50,000 rec/s | Spark cluster |
| Zingg | ~30,000 rec/s | Spark cluster |
GoldenMatch is the fastest single-machine deduplication tool. Splink and Zingg are faster but require distributed Spark clusters.
Simulated with ground truth labels (5% noise to approximate LLM accuracy):
| Dataset | Zero-Shot | LLM Boost (300 labels) | Improvement | Cost |
|---|---|---|---|---|
| DBLP-ACM | 94.8% | 96.6% | +1.8pts | ~$0.30 |
| Abt-Buy | 44.5% | 59.5% | +15pts | ~$0.30 |
The optimal configuration: MiniLM base model, 300 labels, 3 epochs, train on multi-pass pairs, score on ANN pairs.
⚡ GoldenMatch — Entity resolution toolkit | PyPI | GitHub | Open in Colab | MIT License
🟡 Golden Suite (Monorepo)
Suite Packages
- GoldenCheck · data quality
- GoldenFlow · transforms
- GoldenPipe · orchestrator
- InferMap · schema mapping
Getting Started
- Installation
- Quick Start
- Auto-Config Controller · enhanced through v1.12
- Configuration
- Verification · new in v1.5
- CLI Reference
Core Concepts
AI Integration
Advanced
- PPRL
- Domain Packs
- Streaming / CDC
- Database Integration
- GPU & Vertex AI
- REST API
- Interactive TUI
- Web UI · new in v1.7
- Evaluation
Reference
pip install goldenmatch
npm install goldenmatch