Repository navigation
Nutmeg 0.1.0
Graph analytics inside Sail: Grust's graph kernels as a Spark data source and
as SQL table functions, in the Sail server's own process.
This is a tagged source release, not a crates.io publish. Nutmeg's
workspace is publish = false and stays that way: nutmeg-server compiles
Sail's crates into itself, and Sail does not publish to crates.io ("There is
no plan to publish it as Rust crates for use in other Rust projects",
lakehq/sail#1991). Clone the tag and build.
A clean clone builds
Both dependencies are pinned in the workspace manifest, so there are no
sibling checkouts any more:
| dependency | pin |
|---|---|
grust-algorithms (arrow), grust-algorithm-procedures (arrow), grust-procedures, grust-core |
crates.io 0.23.0 ("Langoustine"), cut from querygraph/grust v0.23.0 = 6504c0c050c01071ffc67724e80f124e07ffa5dc |
sail-common, sail-common-datafusion, sail-telemetry, sail-session, sail-spark-connect |
git lakehq/sail rev f1cf1729b1d083f2b97f1ce6e68a0d92c5ccee8f — on lakehq/sail main, carrying the session-factory hook from #2630 |
DataFusion 55.1, Arrow 59, Sail 0.7.1 follow from the lock file.
git clone --branch v0.1.0 https://github.com/querygraph/nutmeg
cd nutmeg && cargo build --release -p nutmeg-serverWhich of the two shapes you want
Grust reaches Sail two ways, and the README now opens with the choice.
- Nutmeg — embedded.
nutmeg-serveris the Sail Spark Connect server,
with Nutmeg's data source and table functions in every session. A staged
graph is Arrow memory in that process and a kernel reads it in place:
nothing crosses a process boundary. One machine and speed → this. The
costs are honest ones: you are bounded by that machine, and staged graphs
are gone when the server restarts. It does not work in Sail's
local-clusterorkubernetes-clustermodes. grust-sail— the client, in Grust. A Spark Connect client that links
no Sail crate, talks to a stock Sail server of any topology, and keeps
graphs in ordinarygrust_nodes/grust_edgesDelta tables. An existing
Sail cluster → this. You pay a round trip of the edge list, and two
things are stated exactly rather than promised:grust-sailruns no graph
kernel and pushes none into Sail (it pushes degree aggregates, triplet
joins, lowered traversals and the pushable subset of read-only Cypher), so
runninggrust-algorithmsover Sail-held data means reading the graph into
your own process; and while nothing in it is deployment-specific, Grust
records no cluster-mode run of it, so treat cluster support as unobstructed
rather than verified.
Score precision on the rank kernels
Grust 0.23's precision option — f64 (the default) or f32 — is served on
pagerank and articleRank, and the schema a read reports follows it:
Spark double at f64, float at f32. Result schemas are observed by
running a kernel once on a three-node probe graph, and that observation is
now cached per value of precision instead of per algorithm.
nm.pagerank.stream(g, precision="f32")SELECT score FROM nutmeg_pagerank('g', '{"precision": "f32"}')It is Grust's own option, so a bad value is refused at planning with a
message naming it, exactly as a misspelled dampng is.
What is in it
Everything in CHANGELOG.md under 0.1.0: streaming reads that an interrupt
can stop, per-read executions with their own timeout, work limit, memory
ceiling and thread count, one memory budget over staged rows, sorts,
projections and kernels, canonical staging order, joined-DataFrame staging,
declared nullability, no unsigned column, GDS column-name aliases, kernels
that read node properties, and the Citi Bike example.
Nutmeg serves whatever Grust's registry holds, through Grust's own runner: 41
kernels at Grust 0.23.0, each reachable as a data source read and as a SQL
table function. A test reads that list back out of the README, so it cannot
go stale.
The Citi Bike example, recaptured
examples/citibike follows Neo4j's Aura Graph Analytics with Spark
tutorial step for step and then goes past it. Its results/ were recaptured
on this release's pins: three consecutive runs gave the same output.md
byte for byte and the same community map, and results.json differed only
in the last digits of mean_minutes_in, a Sail SQL AVG.
Against the previous capture, every value Nutmeg produces is unchanged —
Leiden's 8 communities and modularity 0.44082808748357094 (NetworkX agreeing
to 0.000e+00), PageRank against the NumPy reference at both tolerances,
betweenness at 3.638e-12, A*/Dijkstra at 12914.980413 m. One Nutmeg-reported
number moved, and it is a real Grust 0.23 change: projectionStats's
csrBytes fell from 25531544 to 19147108 for the tutorial's graph — the CSR
is about a quarter smaller.
Not supported
Sail's cluster modes. GRUST-SAIL.md §7 in Grust has the reasoning: a
cluster-mode driver serialises each stage through a hard-wired codec that
refuses a node it does not know, and places stages on workers that hold none
of the driver's graphs. No Sail change is proposed for it.