Skip to content

Nutmeg 0.1.0

Latest

Choose a tag to compare

@alexy alexy released this 24 Sep 03:47
· 4 commits to main since this release

Nutmeg 0.1.0

Graph analytics inside Sail: Grust's graph kernels as a Spark data source and
as SQL table functions, in the Sail server's own process.

This is a tagged source release, not a crates.io publish. Nutmeg's
workspace is publish = false and stays that way: nutmeg-server compiles
Sail's crates into itself, and Sail does not publish to crates.io ("There is
no plan to publish it as Rust crates for use in other Rust projects",
lakehq/sail#1991). Clone the tag and build.

A clean clone builds

Both dependencies are pinned in the workspace manifest, so there are no
sibling checkouts any more:

dependency pin
grust-algorithms (arrow), grust-algorithm-procedures (arrow), grust-procedures, grust-core crates.io 0.23.0 ("Langoustine"), cut from querygraph/grust v0.23.0 = 6504c0c050c01071ffc67724e80f124e07ffa5dc
sail-common, sail-common-datafusion, sail-telemetry, sail-session, sail-spark-connect git lakehq/sail rev f1cf1729b1d083f2b97f1ce6e68a0d92c5ccee8f — on lakehq/sail main, carrying the session-factory hook from #2630

DataFusion 55.1, Arrow 59, Sail 0.7.1 follow from the lock file.

git clone --branch v0.1.0 https://github.com/querygraph/nutmeg
cd nutmeg && cargo build --release -p nutmeg-server

Which of the two shapes you want

Grust reaches Sail two ways, and the README now opens with the choice.

  • Nutmeg — embedded. nutmeg-server is the Sail Spark Connect server,
    with Nutmeg's data source and table functions in every session. A staged
    graph is Arrow memory in that process and a kernel reads it in place:
    nothing crosses a process boundary. One machine and speed → this. The
    costs are honest ones: you are bounded by that machine, and staged graphs
    are gone when the server restarts. It does not work in Sail's
    local-cluster or kubernetes-cluster modes.
  • grust-sail — the client, in Grust. A Spark Connect client that links
    no Sail crate, talks to a stock Sail server of any topology, and keeps
    graphs in ordinary grust_nodes/grust_edges Delta tables. An existing
    Sail cluster → this.
    You pay a round trip of the edge list, and two
    things are stated exactly rather than promised: grust-sail runs no graph
    kernel and pushes none into Sail (it pushes degree aggregates, triplet
    joins, lowered traversals and the pushable subset of read-only Cypher), so
    running grust-algorithms over Sail-held data means reading the graph into
    your own process; and while nothing in it is deployment-specific, Grust
    records no cluster-mode run of it, so treat cluster support as unobstructed
    rather than verified.

Score precision on the rank kernels

Grust 0.23's precision option — f64 (the default) or f32 — is served on
pagerank and articleRank, and the schema a read reports follows it:
Spark double at f64, float at f32. Result schemas are observed by
running a kernel once on a three-node probe graph, and that observation is
now cached per value of precision instead of per algorithm.

nm.pagerank.stream(g, precision="f32")
SELECT score FROM nutmeg_pagerank('g', '{"precision": "f32"}')

It is Grust's own option, so a bad value is refused at planning with a
message naming it, exactly as a misspelled dampng is.

What is in it

Everything in CHANGELOG.md under 0.1.0: streaming reads that an interrupt
can stop, per-read executions with their own timeout, work limit, memory
ceiling and thread count, one memory budget over staged rows, sorts,
projections and kernels, canonical staging order, joined-DataFrame staging,
declared nullability, no unsigned column, GDS column-name aliases, kernels
that read node properties, and the Citi Bike example.

Nutmeg serves whatever Grust's registry holds, through Grust's own runner: 41
kernels at Grust 0.23.0, each reachable as a data source read and as a SQL
table function. A test reads that list back out of the README, so it cannot
go stale.

The Citi Bike example, recaptured

examples/citibike follows Neo4j's Aura Graph Analytics with Spark
tutorial step for step and then goes past it. Its results/ were recaptured
on this release's pins: three consecutive runs gave the same output.md
byte for byte and the same community map, and results.json differed only
in the last digits of mean_minutes_in, a Sail SQL AVG.

Against the previous capture, every value Nutmeg produces is unchanged —
Leiden's 8 communities and modularity 0.44082808748357094 (NetworkX agreeing
to 0.000e+00), PageRank against the NumPy reference at both tolerances,
betweenness at 3.638e-12, A*/Dijkstra at 12914.980413 m. One Nutmeg-reported
number moved, and it is a real Grust 0.23 change: projectionStats's
csrBytes fell from 25531544 to 19147108 for the tutorial's graph — the CSR
is about a quarter smaller.

Not supported

Sail's cluster modes. GRUST-SAIL.md §7 in Grust has the reasoning: a
cluster-mode driver serialises each stage through a hard-wired codec that
refuses a node it does not know, and places stages on workers that hold none
of the driver's graphs. No Sail change is proposed for it.