Repository navigation
Measuring and drawing
The measures are declared with the pipeline, before a single number exists. That ordering is the point: a measure chosen after seeing the result is a measure chosen because of the result, and every run is then measured the same way whether it turned out well or not.
using DeepSharp.Charts;
using DeepSharp.Learners.Networks;
using DeepSharp.Pipelines;
var pipeline = Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema.Integer("survived").Category("sex").Optional("age", ColumnKind.Number).Number("fare"))
.SplitStratified("survived", train: 0.70, validation: 0.15)
.Target("survived")
.FillMissing(fill => fill.Median("age"))
.EncodeCategories()
.Normalise("age", "fare")
.Report(report => report // declared before anything learns
.Measure(Metric.Accuracy, Metric.Precision, Metric.Recall, Metric.ConfusionMatrix)
.On(Part.Train, Part.Validation, Part.Test)
.As(Shown.Numbers, Shown.Drawn))
.WithTensorflow(network => network
.Dense(16).Relu().Dense(1).Adam(0.01).BinaryCrossEntropy().Run(seed: 20260929, epochs: 100))
.Build();
var trained = pipeline.Train();
File.WriteAllText("titanic-report.html", trained.Measures!.Report().ToHtml());
File.WriteAllText("titanic-loss.svg", trained.History!.LossCurve());Every measure comes with the average's beside it. A model that beats the average by nothing is a model that learned nothing, and that comparison is worth more than any single number. The test rows' measures stand last, because they are the ones you are not allowed to tune against.
Measures!.Report() is the report as the declaration said its measures are shown, and ToHtml() is that page —
which is also what a notebook block shows when you ask it for the report.
The charts are drawn from what the run already kept. The loss curve, the learning rate, the confusion matrix,
predicted against actual, the residuals, the measures as bars — all SVG, no viewer to start, nothing recomputed. That is
DeepSharp.Charts, which draws with MatPlotLibNet and is a package of its own: a project that serves a model and draws
nothing never carries it.
A tree and a network are measured by one report, and this is how. One learner a declaration, so the comparison is
two declarations that differ in a single line — .WithML(trainer => trainer.FastTree()) where the other says
.WithTensorflow(…). Everything above that line is the same steps over the same rows in the same parts, and the run made
for a tree leaves out only the scalings a tree does without, writing down which. The report then reads the rows' keys and
their answers and nothing of the features, so the two are measured on the same rows, by the same metrics, in the answer's
own units. See ML.NET learner.
A learner that does without some steps is measured all the same. A pipeline run for a learner that does without some steps writes down which it left out, so two learners fed the same rows are compared rather than guessed between.
Next: Naming the network · Every measure and every chart: Networks