-
Notifications
You must be signed in to change notification settings - Fork 0
pandas
Verdict: use it, accepting more roughness. The equivalent exists but is less mature and less ergonomic than pandas; expect some glue.
| pandas need | Recommended .NET |
|---|---|
General DataFrame, CSV IO, typed columns |
Microsoft.Data.Analysis |
| Time series, rich indices | Deedle (F# origin, excellent at time series) |
dotnet add package Microsoft.Data.Analysisusing Microsoft.Data.Analysis;
DataFrame df = DataFrame.LoadCsv("data.csv");
df["price"] = df["price"].Multiply(1.2);
DataFrame expensive = df.Filter(df["price"].ElementwiseGreaterThan(100));-
groupby/pivot. Less complete and less fluent than pandas; sometimes it's simpler to group with LINQ over the columns. -
Index. No label index like pandas in
Microsoft.Data.Analysis(positional indexing). Deedle is closer. -
Missing values.
null/NaN handling differs from pandas; check column by column. -
Lodestar glue. There is no
DataFrame↔ sparse-matrix bridge and none is planned; the join is a LINQ expression, shown below.
The vectorizers take IEnumerable<string>, so a column reaches them without anything
in between. Both of these work — the typed column already enumerates as strings, and
the untyped indexer needs a cast:
using Microsoft.Data.Analysis;
using Lodestar.Text.Vectorization;
DataFrame df = DataFrame.LoadCsv("reviews.csv");
// The column is typed: enumerate it directly.
IEnumerable<string> documents =
((StringDataFrameColumn)df["review"]).Select(v => v ?? string.Empty);
// Or from the untyped indexer, which enumerates as object.
IEnumerable<string> viaCast =
df["review"].Cast<string?>().Select(v => v ?? string.Empty);
CsrMatrix counts = new TfidfVectorizer().FitTransform(documents);Decide what an empty cell means before you write ?? string.Empty. pandas keeps
NaN distinct from ""; a vectorizer cannot, and a missing review scored as an empty
document is a row of zeros that looks like a legitimate answer. Filter the nulls out if
that is what you mean.
Going back is reading, not converting: CsrMatrix
exposes Values, ColumnIndices and RowPointers, and ToDense() if the shape is
small enough to want a rectangle. Materialising a document-term matrix into a
DataFrame is usually the wrong move — a vocabulary of thirty thousand terms is
thirty thousand columns, almost all zero, which is exactly the density the sparse
format exists to avoid.
The doctrine in CLAUDE.md is native code only where .NET has a real gap, and this is
not one: both sides already exist and the join above is one expression. A bridge would
also cost Lodestar.Text its no-third-party-dependency property, or need a fifth
package to hold a Select.
Guide to be expanded as real needs arise.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels