Skip to content

Releases: ihb2032/MoonFrame

v0.6.0

Choose a tag to compare

@ihb2032 ihb2032 released this 31 Jul 02:22
f272ce0

MoonFrame v0.6.0 — API convergence

Breaking

  • One constructor per type, spelled as the type's own name:
    Field::Field, Schema::Schema, CsvReadOptions::CsvReadOptions,
    HtmlOptions::HtmlOptions, … The new / default / with_* builder
    chains are gone — thirteen builder methods become zero.
  • Every *_with_options entry point folds into its plain form as an optional
    argument: read_csv(path, options?), scan_ndjson(path, options?),
    parse_csv_str(text, options?). NdjsonReadOptions folds into
    JsonReadOptions, and the JSON verbs are now parse_json_str /
    format_json / write_json.
  • Expr is opaque — its AST moved into an internal package — and Expr and
    JoinOptions lose ==. Compare rendered expressions instead:
    (col("a") + lit_int(1)).to_string().
  • The storage-backend surface leaves the public API, and the read-side fields
    of DataFrame / Series / Field / JoinOptions go private behind
    copying accessors, so a validated value cannot be mutated into an
    inconsistent one.
  • TypeMismatch / ParseError carry structured detail instead of a message
    string. DataFrame::take is renamed gather. SortOrder / NullOrder
    move from frame to types. A column-less frame keeps its row count
    (N × 0) rather than collapsing to 0 × 0, and a declared
    nullable = false now survives the projections.

Every source-level upgrade step, with before/after: docs/migration.md

Added

  • Numeric expressions: abs / floor / ceil / sign, floor_div /
    modulo / pow, round(decimals?), is_in, is_between(closed?).
  • Nine string ops: str_reverse, str_pad_start / str_pad_end /
    str_zfill, str_slice, str_split_get, str_len_bytes,
    str_extract, str_count_matches.
  • Predicate push-down into scan_csv / scan_ndjson: the reader parses only
    the rows the filter keeps.
  • Column selectors — numeric_cols, cols_of_dtype, cols_matching,
    cols_starts_with / cols_ends_with / cols_contains.
  • Expr::map_batches (a whole-column escape hatch, usable as a custom
    aggregation), rename_with, DataFrame::reverse / with_row_index,
    unique(subset?, keep?).
  • Series gains sort / head / tail / reverse and the statistics
    std / variance / median / first / last; LazyFrame gains the
    whole-frame reductions sum / mean / min / max / count /
    null_count.

Full history: docs/changelog.md

v0.5.0

Choose a tag to compare

@ihb2032 ihb2032 released this 26 Jun 00:50
c455821

MoonFrame v0.5.0

Highlights

  • One engine for the four verbs. select / filter / agg / with_columns
    all take Exprs now, on both DataFrame and LazyFrame. The v0.4 *_exprs /
    *_where twins, the AggSpec reduction specs, and the closure filter are
    gone — the plain verb is the expression form. A reduction can run over a
    derived column: (col("revenue") - col("cost")).sum().
  • Expression keys everywhere. sort (renamed from sort_by), group_by,
    join, and the drop family name their keys with Expr, so a key can be
    derived, not just a column name.
  • One join. The per-type inner_join / left_join / right_join /
    outer_join / cross_join methods collapse into the single
    join(other, JoinOptions) (Polars has no *_join), with left_on / right_on
    for differently-named keys.
  • A wider vocabulary. New aggregations (std / variance / median /
    n_unique / first / last), a str_* string namespace, fill_null /
    fill_nan / NaN probes on the expression layer, lit_series, and cols.
  • Lazy file sources. scan_csv / scan_ndjson start a plan straight from a
    file, and projection pushdown reaches into the scan — a column the plan
    never reads is never parsed.
  • Series in its own package, and the last non-Polars names are aligned.
  • Robustness. A canonical storage backend closes an optimizer-soundness gap,
    an iterative Expr walk removes a stack-overflow class, and a new differential
    fuzzer permanently guards collect ≡ execute.

⚠️ Breaking changes

v0.5 removes the duplicate verbs and the tail of non-Polars names outright —
there are no deprecated aliases. Every change below is a pure rename or a
mechanical rewrite. The complete before/after tables live in
docs/migration.md; the headlines:

Series moved to its own package

Series is extracted from frame into a new series package so the expression
layer can build on the per-column unit. Facade users are unaffected —
@moonframe.Series still names the same type. Only code importing the type
directly from frame needs the new path
(ihb2032/MoonFrame/series.Series).

Smaller behavioural breaks

  • A CSV delimiter that collides with the quote or a line terminator is now
    rejected up front (#118).
  • The unused EmptyDataFrame error variant is removed (#107) — only affects an
    exhaustive match over DataError.

→ The full, mechanical upgrade is in docs/migration.md.

New & additive

  • Aggregations: std / variance (sample statistics) / median (skips
    NaN) / n_unique / first / last (positional), joining sum / mean /
    min / max / count.
  • String namespace: str_to_uppercase / str_to_lowercase /
    str_strip_chars / str_len_chars / str_contains / str_starts_with /
    str_ends_with / str_replace / str_replace_all (literal matching, no
    regex), each a first-class, introspectable node.
  • fill_null on expressions: col("x").fill_null(value) where value is
    any Expr (a literal, another column for a coalesce, or a tree), plus a
    whole-frame df.fill_null(value).
  • NaN handling: is_nan / is_not_nan probes and fill_nan(value) — the
    dual of fill_null that replaces a Float NaN while leaving nulls in place.
  • lit_series embeds a (broadcasting) Series as an expression, and
    cols(["a", "b"]) expands names to col expressions.
  • map_elements / map_many — the host-closure escape hatch over a row's
    cells as Scalars, carried by an inspectable, pushdown-able Expr.
  • Lazy file sources: scan_csv / scan_ndjson (and _with_options
    variants) with projection pushdown into the scan, so
    scan_csv("sales.csv").select(cols(["region", "revenue"])).collect() reads
    only those two columns. (An array-shaped JSON document has no row-wise scan, so
    there is no scan_json.)
  • DataFrame::unique() drops duplicate rows, keeping first-appearance order.
  • ChartSpec::with_color_type(VegaType) overrides the Vega-Lite type of a
    chart's color channel, so a numeric grouping column renders as distinct
    per-group colours instead of a continuous gradient.

Correctness, performance & internals

  • A canonical storage backend. A column's backend (the unboxed Numeric
    fast path vs. the general Builtin) is now a function of its content, not of
    how it was built: any row gather that leaves an all-valid Int/Float column
    re-converges it onto Numeric. This closes a predicate-pushdown soundness
    gap
    — sinking a Filter below a stage carrying a derived column or group key
    could otherwise land on a different backend than the eager chain, which
    Series equality observes, making collect diverge from execute (#93, #66).
  • Optimizer soundness fixes: keep filters above an aggregate with a
    non-row-stable key (#89), keep sign-sensitive predicates above an aggregate on
    a Float key (#95), and a batch of audit fixes covering optimizer soundness,
    CSV/JSON round-trip, and BOM handling (#92).
  • A differential fuzzer for optimizer/executor equivalence permanently
    guards collect ≡ execute across the widened domain — null / NaN / ±0.0 /
    string / multi-column (#122).
  • Iterative deep-tree walks. Expr trees are now walked iteratively, so a
    pathologically deep expression no longer risks a stack-overflow abort (#121).
  • Other fixes: grouped agg no longer nulls map-wrapped reductions (#112);
    a grouped/whole-frame map whose every result cell is null now falls back to the
    input dtype and yields an all-null column instead of raising; large integers in
    a JSON Float column recover a finite Double (#94); a repeated left_on key
    is allowed for multi-condition joins (#100).
  • Performance: on-demand validity gathers keep grouped aggregation linear
    (#96), the Numeric fast path is taken in scalarization and join gathers
    (#97), truncated rendering is bounded to shown rows and rename is linearized
    (#99), and an identity scope short-circuits a column read (#98).
  • Internals: migration off the deprecated try? expression for MoonBit
    0.10.0 (#102), plus a sweep of dedup/refactor work across column / series /
    frame / io / lazy with no API or behaviour change.

v0.4.0

Choose a tag to compare

@ihb2032 ihb2032 released this 14 Jun 08:43
ab76614

MoonFrame v0.4.0

Highlights

  • Expression engine — build column computations as data: (col("revenue") - col("cost")).with_alias("profit"), comparisons, Kleene-logical & / |, aggregations, cast, and when / then / otherwise. Composable, introspectable (explain()), and vectorized.
  • Eager consumers — with_columns, select_exprs, filter_where, and agg_exprs bring expressions to DataFrame / GroupedDataFrame directly.
  • Lazy query layer — lazy_frame(df). … .collect(), with explain() and an optimizer that does predicate pushdown and projection pushdown.
  • Join hardening — duplicate-key rejection and storage-backend preservation.

Expression engine — the expr package

A reified, composable column expression. col("name") and lit_int / lit_float / lit_str / lit_bool / lit build the leaves; the overloaded operators compose them; and methods extend them:

  • Operators: + - * / (arithmetic — / is always Float, and dividing by zero yields IEEE ±inf / NaN rather than trapping), & / | (Kleene-logical, not bitwise), and unary -.
  • Methods: comparisons eq / ne / lt / le / gt / ge (→ Bool), not / is_null / is_not_null, aggregations sum / mean / min / max / count, cast, and with_alias.
  • Conditional: when(cond).then(a).otherwise(b), row-wise.

Building a tree is total — it never fails; evaluation errors (unknown column, type mismatch) surface later, at the point of use. explain() and the Show impl render the operator form, and referenced_columns / output_name introspect it. An Expr is read-only outside expr: you build it through the surface above, never by naming a variant.


Eager expression consumers — the frame package

The whole-frame evaluator and its DataFrame / GroupedDataFrame consumers:

  • with_columns — derive or replace columns from expressions.
  • select_exprs — project to the evaluated expressions (an all-aggregation selection collapses to a single row).
  • filter_where — vectorized boolean row selection; a reified, pushdown-able alternative to the closure filter.
  • agg_exprs — the expression form of agg, generalising AggSpec to compound reductions.

Evaluation is vectorized, with Int / Float promotion, null propagation, Kleene logic, and the Series reduction's NaN rules (NaN propagates through sum / mean, is skipped by min / max).


Lazy query layer + optimizer — the lazy package

lazy_frame(df) (or LazyFrame::from(df)) starts a deferred plan; total builder methods that mirror the eager verbs grow it; explain() prints it; and collect() runs it.

With no optimizer in front, a collect() is bitwise-equal to the equivalent eager pipeline. The optimizer then adds two result-preserving rewrites:

  • Predicate pushdown — sink each filter toward the scan, past the stages it provably commutes with.
  • Projection pushdown — insert a narrowing selection over a scan whose consumers read only a subset of its columns.

So explain() versus explain(optimized=true) is a before / after view of exactly what the optimizer moved and pruned. LazyFrame::group_by(keys).agg(exprs) is the lazy mirror of eager grouping.


Join — duplicate-key check and backend preservation

  • join now rejects a key repeated in on with DuplicateColumn, matching the "no duplicate keys" contract of group_by and select (previously on = ["id", "id"] silently collapsed to the single key ["id"]). A missing repeated key still surfaces as ColumnNotFound at its first appearance.
  • Join output columns now preserve the storage backend of their source where they pick up no unmatched-row null — an all-valid Numeric source column stays Numeric instead of demoting to Builtin, matching filter / sort_by / take / drop_nulls / fill_null. Only the representation changes; values and dtypes are identical.

v0.3.0

Choose a tag to compare

@ihb2032 ihb2032 released this 09 Jun 08:46
77b70b2

MoonFrame v0.3.0 — Release Notes

MoonFrame v0.3.0 grows the library's user-facing output surface, completes the join matrix, and adds engineering depth with a pluggable column-storage backend — all on top of the v0.2 method-chain core. Pre-1.0, breaking changes ride the minor version (see Breaking changes).

Highlights

  • HTML & Vega-Lite export — render a frame as a styled <table> or a Vega-Lite v5 chart spec you can paste straight into the Vega editor.
  • Full join matrix — right / outer complete the inner / left / right / outer / cross set.
  • Read resilience — the CSV / JSON / NDJSON readers gained escape hatches for messy, real-world data.
  • Pluggable column storage — an all-valid Numeric fast path alongside the general-purpose Arrow Builtin column, behind a closed ColumnStorage seam.
  • Faster group-by / join and O(1) slicing under the hood.

What's new

Output formats

  • DataFrame::to_html() / to_html_with_options(...) — a pure, dependency-free renderer (CSS class, <caption>, a max_rows cap, and HTML escaping by default), parallel to to_markdown.
  • format_vega_lite(df, spec) / write_vega_lite(...) — a complete Vega-Lite v5 spec ($schema + mark + encoding + inline data.values). ChartSpec::bar / line / point / area, with with_color / with_title.

Join — Right / Outer

JoinType gained Right and Outer, completing the matrix, with right_join / outer_join convenience methods and Polars-aligned semantics (null keys match nothing; NaN keys match; per-how coalesce defaults, overridable with with_coalesce).

Read resilience

infer_schema_rows = 0 scans every row instead of a leading window; on_parse_error (Raise / Null) chooses whether an off-type cell past the window fails the read or downgrades to null; CSV's allow_nonfinite_floats controls whether nan / inf tokens infer as Float or fall back to String.

Pluggable column storage

A Series now holds a ColumnStorage — a closed { Builtin; Numeric } seam. Numeric is an all-valid, unboxed Int64 / Double column with no validity bitmap (the null_count == 0 fast path); from_ints / from_floats pick it automatically, and structural transforms keep a column on the fast path. storage_kind() / to_numeric() / to_builtin() inspect and convert.

Performance

  • group-by / join keys are now hashed on native cell values — no per-row string key, and no decimal-string round-trip for numeric keys.
  • Bitmap slicing is O(1) — head / tail / slice share the parent's validity buffer (a zero-copy view) instead of repacking it.

Breaking changes

Pre-1.0, breaking changes ride the minor version. The source-level breaks:

  • Series::storage() now returns @column.ColumnStorage (was @column.BuiltinColumn). The .data() / .validity() reading surface is unchanged, so column-reading call sites still compile; use .to_builtin() for the concrete BuiltinColumn.
  • Series::new(name, ...) takes a ColumnStorage (pass ColumnStorage::from_builtin(col), or use the unchanged Series::from_builtin).
  • JoinType gained Right / Outer — an exhaustive match over it must now handle the two new variants.
  • The CSV / JSON / NDJSON *ReadOptions structs gained pub(all) fields; a full struct literal must add them or switch to ::default() (the defaults reproduce the prior behaviour exactly).

Step-by-step upgrade notes are in migration.md.

v0.2.0

Choose a tag to compare

@ihb2032 ihb2032 released this 05 Jun 08:50
8d6a2ae

MoonFrame v0.2.0 — Method-chain API, GroupBy, Join & NDJSON

MoonFrame v0.2 moves the entire v0.1 surface to a method-chain + raise API, then grows split-apply-combine (group_by), relational join, and NDJSON I/O on the new foundation. This is a breaking release (pre-1.0, so it ships as a minor bump) — see Breaking changes & migration.

📖 Public API reference · 🧪 Runnable quickstart

Highlights

🔗 Method-chain + raise API

Pipelines read top-to-bottom like pandas / Polars. Every fallible verb is now a method on DataFrame that raise DataError instead of returning a Result — bridge back to a value with try?.

📊 GroupBy

group_by(keys).agg([...]) returns one row per group with Count / Sum / Mean / Min / Max reductions, optional with_alias, deterministic first-appearance group order (Polars maintain_order = true), and null keys kept as their own group.

🔗 Join

Hash equi-join with inner / left / cross variants and Polars-aligned semantics: a null key matches nothing (null != null), a NaN key matches other NaNs, colliding right columns gain a "_right" suffix, and key columns coalesce by Polars' per-how default (overridable via with_coalesce).

📄 NDJSON I/O

Read and write the JSON Lines format (one JSON object per line), reusing the JSON-records type inference. Reading is lenient (blank lines skipped, CRLF tolerated); writing emits one compact object per row.

🧪 Quality

  • Runnable quickstart.mbt.md — doc tests compiled and run on all four backends in CI, so the examples can't go stale.
  • QuickCheck property tests for the frame and io packages.
  • Strict CSV column validation — CsvReadOptions::strict_column_count rejects ragged rows instead of silently padding / truncating.
  • Cross-platform CI matrix + a line-coverage gate.

⚠️ Breaking changes & migration

v0.1 v0.2
@ops.select(df, names) (free function) df.select(names) (method)
op(df, ...) -> Result[T, DataError] + .bind / .unwrap df.op(...) -> T raise DataError, chained directly
pattern-match Ok(x) / Err(e) call directly in a raise context, or try? expr for a Result
filter_try(df, row => row.get_int("x").map(v => v > 0)) df.filter(row => row.get_int("x") > 0)
sort_by(df, spec) / sort_by_many(df, specs) df.sort_by([(col, order, nulls), ...])
Series::min() / max() (Result-wrapped) Series::min_value() / max_value() (total)
@io.to_markdown(df) df.to_markdown()
enum DataError pub(all) suberror DataError (same variants / message() / Show)
import ... @ops gone — verbs live on DataFrame in @frame

Two more behavioural notes:

  • sum / mean now propagate NaN (Polars-aligned: null is missing, but NaN is a value). Min / Max / Count / unique counts were already aligned.
  • The ops package, filter_try, sort_by_many, and the Result-wrapped Series::min() / max() are removed.

read_* / write_* / parse_* I/O now raise; the format_* serialisers stay total and return a String. Full per-symbol detail in docs/api.md.

What's Changed

  • Fix v0.1 review findings by @ihb2032 in #15
  • Migrate public API to method-chain + raise (v0.2 Stage 1) by @ihb2032 in #16
  • Fix JSON non-finite floats + from_rows empty-schema invariant by @ihb2032 in #17
  • Refactor loops to closures by @ihb2032 in #18
  • feat(frame): add group_by by @ihb2032 in #19
  • feat(frame): add join by @ihb2032 in #20
  • feat(frame): align sum / mean NaN handling with Polars by @ihb2032 in #21
  • test: add quickcheck property tests, runnable quickstart, all-backend CI by @ihb2032 in #22
  • feat(io): add ndjson support by @ihb2032 in #23
  • chore: add strict CSV column validation and strengthen quality gates by @ihb2032 in #24

v0.1.0

Choose a tag to compare

@ihb2032 ihb2032 released this 30 May 03:11
07f1515

First public release — a lightweight, column-oriented DataFrame /
tabular-data library for the MoonBit
ecosystem.

v0.1 covers the core value types, an Apache-Arrow-style column backend,
Series / DataFrame / RowView, the full operator set (select /
filter / sort / null-handling / stats / describe), and three I/O formats
(CSV / Markdown / JSON) — usable both as sub-packages and through a
single facade import.

Highlights

  • Total-by-construction error model — never aborts the host process;
    no hidden unwrap panics. Fallible work returns Result[_, DataError],
    provably-total work returns its value directly.
  • Apache Arrow style column layout — separate value buffer + packed
    validity Bitmap (1 = valid); 64-bit numerics (Int64 / Double).
  • DataFrame with O(1) name lookup and a check_invariants()
    structural spec asserted by every op test.
  • Stable mergesort with explicit NullOrder; SQL-style lexicographic
    String ordering.
  • I/O — NyaCSV-backed CSV with strict type inference, GFM Markdown
    tables, and records-shape JSON round-trip.
  • Facade package — import "ihb2032/MoonFrame" @moonframe reaches the
    whole surface; sub-package imports remain supported.

Quality

  • 514 tests · 100% line coverage · 0 warnings on moon check
  • No unwrap / abort in library code; 0 panics, 0 stubs

Documentation

License

Apache-2.0