-
Notifications
You must be signed in to change notification settings - Fork 0
migration
This page is Lodestar's migration hub. It answers a simple question: "I do this in Python, what do I do in C#?"
The project's guiding principle (see the rationale) is honest: we don't rewrite Python's data-science ecosystem. Most of it already exists in .NET, and Python's dense linear algebra relies on Fortran BLAS/LAPACK kernels there's no point reimplementing. We use what exists, and only write native code where .NET has a real gap: text (similarity, vectorization).
| Python | Role | .NET recommendation | Verdict |
|---|---|---|---|
| PyTorch | tensors, autograd, training, GPU | TorchSharp (= libtorch); ONNX Runtime for inference only | ✅ Use |
| matplotlib | plotting | ScottPlot, Plotly.NET, OxyPlot | ✅ Use |
| NumPy | N-dim arrays, dense algebra |
Math.NET Numerics (+ native MKL/OpenBLAS provider); System.Numerics.Tensors
|
✅ Use |
| scikit-learn | classical ML, pipelines, metrics | ML.NET; SharpLearning | ✅ Use except text vectorization → Lodestar.Text and classification metrics → Lodestar.Metrics |
| pandas | DataFrame, groupby, IO |
Microsoft.Data.Analysis; Deedle
|
🟡 Use (rougher) |
| statsmodels | econometric regression, time series, tests | Math.NET (basics); Accord.NET | 🟠 Decide — rich econometrics is a gap |
| seaborn | tidy statistical viz | ScottPlot / Plotly.NET (charts rebuilt) | 🟠 Decide — statistical presets missing |
Legend. ✅ a solid equivalent exists, use it as is. 🟡 an equivalent exists but is less mature than Python; expect some glue. 🟠 the foundation exists but a whole area is missing: a candidate for native code if your usage justifies it.
One area truly justifies native code — text — and one more turned out to be
a gap the .NET options do not fill honestly: the evaluation metrics every
sklearn user reaches for. That's
Lodestar.Text
and its siblings, delivered as lots (see the brief):
- String distances & similarity — Levenshtein, Damerau-Levenshtein, Jaro-Winkler, Jaccard, Ratcliff-Obershelp, phonetics… (done)
-
Tokenization & sparse vectorization —
CountVectorizer,TfidfVectorizer(exact sklearn semantics), home-grown CSR matrix. (done) - Embeddings & semantic search — ONNX Runtime + sub-word tokenizers. (done)
-
Applied fuzzy matching —
rapidfuzz.fuzz/processequivalents. (done) - Classification metrics — sklearn-parity precision, recall, F1, confusion matrix, report and ROC-AUC. (done)
| Guide | Status |
|---|---|
| NumPy → .NET | draft |
| pandas → .NET | draft |
| scikit-learn → .NET | draft |
| statsmodels → .NET | draft |
| PyTorch → .NET | draft |
| matplotlib → .NET | draft |
| seaborn → .NET | draft |
The detailed equivalence table (Python call → C# call, behavioral differences,
performance notes), filled in as we go, is in equivalence.md.
This document is not legal advice; third-party dependency licenses are recorded in
THIRD-PARTY-NOTICES.md.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels