Skip to content

v0.2.0

Choose a tag to compare

@ihb2032 ihb2032 released this 05 Jun 08:50
· 431 commits to main since this release
8d6a2ae

MoonFrame v0.2.0 — Method-chain API, GroupBy, Join & NDJSON

MoonFrame v0.2 moves the entire v0.1 surface to a method-chain + raise API, then grows split-apply-combine (group_by), relational join, and NDJSON I/O on the new foundation. This is a breaking release (pre-1.0, so it ships as a minor bump) — see Breaking changes & migration.

📖 Public API reference · 🧪 Runnable quickstart

Highlights

🔗 Method-chain + raise API

Pipelines read top-to-bottom like pandas / Polars. Every fallible verb is now a method on DataFrame that raise DataError instead of returning a Result — bridge back to a value with try?.

📊 GroupBy

group_by(keys).agg([...]) returns one row per group with Count / Sum / Mean / Min / Max reductions, optional with_alias, deterministic first-appearance group order (Polars maintain_order = true), and null keys kept as their own group.

🔗 Join

Hash equi-join with inner / left / cross variants and Polars-aligned semantics: a null key matches nothing (null != null), a NaN key matches other NaNs, colliding right columns gain a "_right" suffix, and key columns coalesce by Polars' per-how default (overridable via with_coalesce).

📄 NDJSON I/O

Read and write the JSON Lines format (one JSON object per line), reusing the JSON-records type inference. Reading is lenient (blank lines skipped, CRLF tolerated); writing emits one compact object per row.

🧪 Quality

  • Runnable quickstart.mbt.md — doc tests compiled and run on all four backends in CI, so the examples can't go stale.
  • QuickCheck property tests for the frame and io packages.
  • Strict CSV column validation — CsvReadOptions::strict_column_count rejects ragged rows instead of silently padding / truncating.
  • Cross-platform CI matrix + a line-coverage gate.

⚠️ Breaking changes & migration

v0.1 v0.2
@ops.select(df, names) (free function) df.select(names) (method)
op(df, ...) -> Result[T, DataError] + .bind / .unwrap df.op(...) -> T raise DataError, chained directly
pattern-match Ok(x) / Err(e) call directly in a raise context, or try? expr for a Result
filter_try(df, row => row.get_int("x").map(v => v > 0)) df.filter(row => row.get_int("x") > 0)
sort_by(df, spec) / sort_by_many(df, specs) df.sort_by([(col, order, nulls), ...])
Series::min() / max() (Result-wrapped) Series::min_value() / max_value() (total)
@io.to_markdown(df) df.to_markdown()
enum DataError pub(all) suberror DataError (same variants / message() / Show)
import ... @ops gone — verbs live on DataFrame in @frame

Two more behavioural notes:

  • sum / mean now propagate NaN (Polars-aligned: null is missing, but NaN is a value). Min / Max / Count / unique counts were already aligned.
  • The ops package, filter_try, sort_by_many, and the Result-wrapped Series::min() / max() are removed.

read_* / write_* / parse_* I/O now raise; the format_* serialisers stay total and return a String. Full per-symbol detail in docs/api.md.

What's Changed

  • Fix v0.1 review findings by @ihb2032 in #15
  • Migrate public API to method-chain + raise (v0.2 Stage 1) by @ihb2032 in #16
  • Fix JSON non-finite floats + from_rows empty-schema invariant by @ihb2032 in #17
  • Refactor loops to closures by @ihb2032 in #18
  • feat(frame): add group_by by @ihb2032 in #19
  • feat(frame): add join by @ihb2032 in #20
  • feat(frame): align sum / mean NaN handling with Polars by @ihb2032 in #21
  • test: add quickcheck property tests, runnable quickstart, all-backend CI by @ihb2032 in #22
  • feat(io): add ndjson support by @ihb2032 in #23
  • chore: add strict CSV column validation and strengthen quality gates by @ihb2032 in #24