Skip to content

ML.NET learner

H.P. Gansevoort edited this page Oct 5, 2026 · 2 revisions

The ML.NET learner

A table is often learned better by a tree than by a network, and a table that a tree learns better does not have to become a network. DeepSharp's learner seam takes any learner, so a trainer from ML.NET stands exactly where a network stands: the pipeline prepares the rows, the learner learns from them, and the pipeline's own report measures what comes back. That last part is the point — a tree and a network measured by one report on the same rows is the only comparison between them that means anything.

Two packages, and which one you need

what it is brings
DeepSharp.Learners.ML The verb: it says which trainer a pipeline's rows are prepared for, and it reads and writes the model file. nothing of ML.NET
DeepSharp.Learners.MLNet The training: it hands the rows to Microsoft.ML, fits the tree and writes the model. Microsoft.ML and Microsoft.ML.FastTree

The split is the dependency, not the design. A notebook, a server or an application that only reads a model is forced to carry whichever package brings a verb — so the verb's package carries no ML.NET at all, and the library travels only to whoever asks to train. A check of the packages measures exactly that on every build: one application trains and writes a model file, another reads that file back while carrying nothing of ML.NET.

Declaring it

using DeepSharp.Learners.ML;
using DeepSharp.Pipelines;

var pipeline = Pdd.Create()
    .ReadCsv("titanic.csv")
    .Declare(schema => schema.Integer("survived", "sibsp").Category("sex").Optional("age", ColumnKind.Number).Number("fare"))
    .SplitStratified("survived", train: 0.70, validation: 0.15)
    .Target("survived")
    .FillMissing(fill => fill.Median("age"))
    .EncodeCategories()
    .Normalise("age", "fare")                       // declared, and left out of the run a tree is made for
    .Report(report => report.Measure(Metric.Accuracy, Metric.Precision).On(Part.Validation, Part.Test).As(Shown.Numbers))
    .WithML(trainer => trainer.FastTree(trees: 50)) // or .FastForest(…)
    .Build();

That is the declaration, and it needs only DeepSharp.Learners.ML. Training it needs the other package, and reads as one line more:

// using DeepSharp.Learners.MLNet;
// var trained = pipeline.TrainWithML();
// File.WriteAllText("titanic.mlmodel.json", trained.ToJson());

TrainWithML rather than Train, because the networks package already offers Train on a pipeline and two packages offering one name on one type either refuse to compile or quietly take each other's calls.

What a tree does without, and why it still compares

A tree takes numbers of any size, so the run made for it leaves the scalings out — Needs.NoScale — and writes down which steps it left out, in the file. Everything else is the same run over the same rows in the same parts, which is why the report can measure a tree and a network against each other at all. Declare the scalings anyway: the same declaration is then the one a network is trained from, and the two differ in a single line.

Which trainers, and why so few

Boosted trees and a forest. No linear trainer, no LightGBM.

A declaration is a promise that running it again gives the same model, and that promise is what decides the list. Measured on this library's own rows: at its defaults ML.NET's stochastic dual coordinate ascent trainer gave five different models in five runs, even under one seeded context; pinning it to a single thread makes it repeat on one machine, and the last bits still move when the processor takes another instruction path. Boosted trees do not move. So a tree repeats from the seed the declaration carries, and a trainer that cannot be replayed is refused by name rather than offered with a warning nobody reads.

LightGBM is left out for a plainer reason: its native side is built for x64 only and wants an OpenMP runtime the package does not bring.

The seed, and what it actually changes

The declaration carries a seed, and the trainers read it from there — ML.NET's own context seed never reaches them. Measured: a forest bags its rows, so another seed gives another model; boosted trees at these settings draw nothing at random, so the seed changes nothing about them. Both are replayable, which is what the seed is there to promise.

The model file

One file, as a network's is: the pipeline's text exactly as it was written, the model bound to that text, and a model read beside another fit of those steps refused rather than quietly answering from numbers learned somewhere else.

It says two things a network's file does not have to. The model inside it is ML.NET's own archive, which only ML.NET opens — so the file names the version of ML.NET that wrote it and the processor it was written on, and a reader that cannot open it can say which package and which version would. The archive's own bytes are never what ties the file to anything: ML.NET stamps it with the clock it was saved at, so two saves of one model differ while the model does not.

Next: every verb in the order you write them: Pipeline · the other learner: Networks · why the course is shaped this way: PDD

Clone this wiki locally