-
Notifications
You must be signed in to change notification settings - Fork 0
Networks
The learning half: the layers, the losses, the optimizers and the loop that trains them; a network described in
Keras's words or written as code, which is the same network either way; and a network trained behind a pipeline,
measured by its report, drawn, and kept with that pipeline as one file. The layers and the loop live in DeepSharp,
beside the tensors. Where a network meets a pipeline is DeepSharp.Learners.Networks, and the charts are
DeepSharp.Charts — each a package of its own, so a project carries only what it uses. The arithmetic runs on the light
engine unless you hand it another, such as libtorch — TorchSharp backend — and a network PyTorch, Keras or an ONNX
exporter saved is read into the same network — Importing a model.
dotnet add package DeepSharp
dotnet add package DeepSharp.Pipelines
dotnet add package DeepSharp.Learners.Networks
dotnet add package DeepSharp.Charts
using DeepSharp.Charts;
using DeepSharp.Learners.Networks;
using DeepSharp.Networks;
using DeepSharp.Pipelines;
var prepared = Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema.Integer("survived", "sibsp", "parch").Category("pclass", "sex").Optional("age", ColumnKind.Number).Number("fare"))
.SplitStratified("survived", train: 0.70, validation: 0.15)
// ---- nothing above this line learns from the rows ----
.FillMissing("age", With.Median)
.EncodeCategories()
.Normalise("age", Scale.MidRange) // a network takes every feature between -1 and 1
.Normalise("fare", Scale.MidRange)
.Normalise("sibsp", Scale.MidRange)
.Normalise("parch", Scale.MidRange)
.Target("survived")
.Report(report => report
.Measure(Metric.Accuracy, Metric.Precision, Metric.Recall, Metric.ConfusionMatrix)
.On(Part.Train, Part.Validation, Part.Test)
.As(Shown.Numbers, Shown.Drawn))
.Build()
.Run();
var trained = new Sequential().Dense(16).Relu().Dense(1)
.Compile(new Adam(0.01), new BinaryCrossEntropy())
.Fit(prepared, new FitOptions(seed: 20260929)
{
Epochs = 100,
EarlyStopping = new EarlyStopping { Patience = 10, RestoreBest = true },
});
var test = trained.Measures!.Parts[^1]; // the test rows' measures, each beside the average's
File.WriteAllText("titanic.network.json", trained.ToJson());
File.WriteAllText("titanic-loss.svg", trained.History!.LossCurve());The network learns from the training rows and is judged by the validation rows once an epoch; the test rows never reach the loop, and the report measures them once it has learned. Run on the published file, it stopped after 46 epochs, kept the network of its 36th, and got 0.815 of the test passengers right, where predicting the training rows' most common answer gets 0.615. The networks sample in the repository trains it, a price five days on and a day's bikes hour by hour, and prints what each did.
In Keras's words. new Sequential() and then the layers — .Dense(units), .Relu(), .Tanh(), .Sigmoid(),
.Dropout(rate), .BatchNorm(), .LayerNorm(), .Reshape(shape), .Conv2D(filters, window), .Flatten() — and
nothing about widths: each is worked out from the shape of an example as it reaches the layer, a row of numbers for a
dense layer, an image of rows, columns and channels for a convolution. A word an example cannot reach is refused,
naming the word, where it stands and what reached it:
This network cannot be lowered at 36: word 2, dense, takes each example as a row of numbers, and each reaching it is 6x6x1: flatten it first.
Compile takes the description as it stands and draws nothing. The network is built at its first fit, from the shape
of the rows it learns from and the run's seed, so the one seed the history records reproduces the whole run — what
the layers start at, the order of every epoch's rows, what every dropout leaves out. .Input(shape) states the shape of
an example, so every word is checked where the description is compiled; to hold the network before any rows, lower it
by hand — description.Lower(new Shape(14), new RandomStream(42)) — and compile that.
The words take Keras's settings where Keras's differ: .BatchNorm(momentum, epsilon) takes Keras's momentum — the
share of the running statistics a training batch leaves as they were, 0.99 in Keras — and keeps its complement, with
Keras's epsilon, and .LayerNorm(epsilon) takes Keras's epsilon; without settings, both are PyTorch's. A window padded
as TensorFlow's padding='same' is new Window(3, 3) { Stride = 2, PaddingMode = PaddingMode.Same }. And a last
.Sigmoid() before BinaryCrossEntropy, Keras's habit, is lifted into the loss, which applies the sigmoid itself: the
network ends at the layer before it, so every prediction goes through one sigmoid, not two. A stack that ends in a
Sigmoid layer is refused where it is compiled with such a loss, naming the layer.
As code. A network is a layer that holds layers, and a LayerStack holds them in order:
using DeepSharp.Networks;
var stream = new RandomStream(42);
var network = new LayerStack(
new Dense(14, 16, stream.Draw("initialise:0", 0, 0)),
new Relu(),
new Dense(16, 1, stream.Draw("initialise:2", 0, 0)));That is exactly what the words above lower to under the same stream: the same numbers, bit for bit, and the same
file. A network with a forward pass of its own derives from Network, adds its layers by name with AddLayer and
says what it does in Compute; to be saved in its file, it says the name the file knows it by and how it is built
again from what it wrote:
using System.Text.Json;
using DeepSharp.Networks;
using DeepSharp.Tensors;
public sealed class Passenger : Network, ISaved<Passenger>
{
private readonly Dense _hidden;
private readonly Dense _out;
public Passenger(Draws draws)
{
_hidden = AddLayer("hidden", new Dense(14, 16, draws));
_out = AddLayer("out", new Dense(16, 1, draws));
}
public static string Name => "passenger"; // what its file calls it
public static Passenger Rebuild(JsonElement settings, Rebuilding rebuilding) => new(rebuilding.Draws);
public void WriteSettings(Utf8JsonWriter writer) // nothing: its layers write their own
{
}
protected override Tensor Compute(Tensor input, Pass pass) =>
_out.Forward(pass.Backend.Relu(_hidden.Forward(input, pass)), pass);
}A program that reads its file back registers it with the catalog it reads with:
TrainedNetwork.FromJson(text, NetworkCatalog.BuiltIn().Register<Passenger>(), StepCatalog.BuiltIn()). One that is not
registered is refused, naming it, rather than read as something else.
Every number a network holds is a slot with a dotted path — hidden.weight, 0.bias, 2.running_mean — its own
first, then those of the layers it holds; Slots() lists them, and HeldLayers() every layer a layer holds, each with
its path, so the layer holding any slot is named by the slot's path. A Parameter is moved by the optimizer; a
RunningStatistic, such as a batch normalisation's mean, only by a training pass, so the validation rows never shape
what the network keeps. A pass says whether it trains or evaluates, which is why no layer has a mode to switch.
Dense(inputs, outputs, draws) |
Every value an example holds, weighed into each of so many. |
Relu, Tanh, Sigmoid
|
The activations. |
Dropout(rate) |
Leaves out a share of the values while training, and scales what it keeps, so evaluation lets everything through. |
BatchNorm(features) |
Each feature normalised over the batch while training, and by the means and spreads it kept while evaluating; its Momentum is PyTorch's, the share a new batch takes. |
LayerNorm(features) |
Each example normalised over its own features. |
Conv2D(inChannels, outChannels, window, draws) |
A Window slid over each image, channels last, as Keras keeps them; padded by the border it states, or as TensorFlow's 'same'. |
Flatten, Reshape(each)
|
An example laid out as one row, or in another shape holding as many values. |
Each starts as PyTorch's does — weights uniform by the number of values that reach them, biases within one over its square root — and each forward pass and its gradients were matched against PyTorch's on the same rows.
MeanSquaredError |
For amounts: the numbers as they are. |
CrossEntropy |
For shares of a whole: the numbers through a softmax. |
BinaryCrossEntropy |
For chances: each number through a sigmoid. |
Sgd(rate) { Momentum } |
PyTorch's stochastic gradient descent. |
Adam(rate) { Betas, Epsilon } |
PyTorch's Adam. |
ConstantRate, StepDecay, ExponentialDecay, CosineDecay
|
How the rate changes from epoch to epoch. |
A loss takes the network's raw numbers and says how they are read, so there is no softmax layer to forget or to apply twice, and a prediction goes through the same activation. It is always named, and it refuses a row whose answers it could not have meant — shares that do not sum to one, a chance above one — naming the row as it was read, before anything is trained on it.
compiled.Fit(train, validation, new FitOptions(seed)) trains on rows of tensors, and compiled.Fit(prepared, options) on the rows a pipeline hands over. The seed is asked for and never assumed; the rest is Keras's fit: one
epoch and batches of thirty-two unless said, the rows shuffled afresh every epoch and the last batch kept short.
EarlyStopping { Patience, MinDelta, RestoreBest } stops the run as Keras's does, and restoring the best epoch brings
back every slot. A loss that is not a finite number stops the run with the epoch and the batch it came from. The options
hold the light engine unless they name another — new FitOptions(seed) { Backend = engine } — and that engine does the
whole run: the training passes, the looks at the validation rows and the report's measures. Every batch's gradients are
worked out by a recording of its pass, and only towards the slots the optimizer moves: nobody asks how a batch's own rows
moved the loss, so that way back is never worked out.
History holds what the run did: every epoch's training loss, validation loss and rate, the seed, the best epoch and
why it stopped.
Checkpoints are taken every epoch, or only when the validation loss improves, and a run goes on from one exactly as
if it had never stopped — on the engine it was taken on, bit for bit. A checkpoint records what its run went under: the
seed, the batch size, the early stopping, and the engine, with the version and the device an engine that implements
INamesItsVersionAndDevice names. Handed another engine — or another seed, batch size or early stopping — a run going on
from it is refused, naming each difference, before anything is put back, since it would be another run under the same
seed: gone on from on libtorch, the networks sample's Titanic run ended with weights up to 3.0e−7 from those of the run
it was taken of. The epochs are not among them, since going on to more of them is what a resume is for:
var compiled = new Sequential().Dense(16).Relu().Dense(1).Compile(new Adam(0.01), new BinaryCrossEntropy());
compiled.Fit(prepared, new FitOptions(seed: 7)
{
Epochs = 50,
Checkpoints = new Checkpoints(checkpoint => File.WriteAllText("run.json", CheckpointFile.Write(compiled, prepared, checkpoint))),
});
var resumed = CheckpointFile.Read(File.ReadAllText("run.json"), NetworkCatalog.BuiltIn(), prepared);
var more = resumed.Compiled.Fit(prepared, new FitOptions(seed: 7) { Epochs = 100, ResumeFrom = resumed.Checkpoint });CheckpointFile.Read reads the file once, whole, before it holds it to the pipeline handed over: its network is held to
the pipeline the file carries, and that to the one handed over. A checkpoint kept under keys of your own is read with
NetworkDocument.ReadCheckpoint(json, network, training, catalog), its network and its run from one reading of the
text. A checkpoint 0.4.0 wrote records neither the batch size, nor the early stopping, nor the engine, and goes on under
whatever it is handed, as it did.
What Fit(prepared, options) gives back is a TrainedNetwork: the network, its history, how the pipeline's report
measured it, and what it predicts for rows that arrive later — in the answer's own units, through the loss's activation
and the pipeline's way back:
var passengers = new InMemoryRowSource(
["pclass", "sex", "age", "sibsp", "parch", "fare"],
[["3", "male", "22", "1", "0", "7.25"], ["3", "Male", "22", "1", "0", "7.25"]]);
var predictions = trained.Predict(passengers); // the chance each passenger survived: 0.16 for the first, as it ran here
var unfamiliar = predictions.Unfamiliar; // nothing for the first; sex_other for the secondA target comes back as a value or the chance of one, a label as its chance, a share of a whole as a count by the row's
own total, and a return as a price by the row's own price. A class is never guessed: where the line between classes
lies is for the use to say. trained.Predict(rows, engine) serves on the engine handed in, for that call alone, and
trained.Predict(rows) on a light engine of its own; either asks the engine for thirty-two rows at a time, Keras's own
default, and no answer depends on which rows share its pass.
What the network learned nothing about. A feature that held one value on every training row taught the network
nothing. The encoder keeps a place for a category the training rows never held — sex_other — and marks a cell that was
empty — sex_was_missing — and on the Titanic passengers those are nought on all 623 training rows, so the first
layer's weights for them are still the random start's after training. A passenger written 'Male' is answered from there,
by the seed rather than by anything learned: across six seeds the same passenger was given 0.257 to 0.495 as 'Male',
against 0.162 to 0.186 as 'male'. Unfamiliar names, for each served row, the features it moves away from the one value
the training rows held, and the report counts the rows of each part that do, as UnfamiliarRows beside Rows. The
answer itself is left as it is: a pipeline that should answer no category its training rows never held declares
EncodeCategories(unseen: Unseen.Refuse).
A network takes every feature on one scale, so it is refused a run made for a learner that does without the scales or
takes categories itself — RunFor(Needs.NoScale) or RunFor(Needs.Categories), which Pipeline has — and fits behind
a run of every step. A pipeline read from its file holds no rows, and fitting behind it says so, and how to have them:
run its declaration over the rows, new Pipeline(declaration, rows, folder).Run().
One file. trained.ToJson() writes the network and the pipeline it was trained behind together, and
TrainedNetwork.FromJson(text, NetworkCatalog.BuiltIn(), StepCatalog.BuiltIn()) reads them back in a program that has
never seen the data. The network's part says what it was trained on: its features, its answers and the output that
named them, the seed, the epoch its numbers come from, each feature that held one value on every training row, and the
SHA-256 of the pipeline's own file as it was fitted. The same names are not the same fit — fitted again on a few more
rows, a pipeline keeps every feature's name and moves the numbers it learned — so the network is read only beside the
very fit it was trained behind, and refused beside any other, when it is loaded and when a run goes on. A checkpoint is
the same file with what the run needs to go on, and serves as a trained network too. The file names no engine, so a
network trained on one engine is read back and served on another, and every line of it ends with a line feed, on every
system.
The file is one text, and .NET makes no text longer than 1,073,741,791 characters, so the one file holds a network of up to about 45 million parameters — its numbers take about 24 characters each — and a checkpoint under Adam, which keeps two more numbers a parameter, one of about 14 million, or about 10 million keeping the best epoch's. Reading a network's file takes 36 bytes a parameter — the text's own copy and the numbers — and a checkpoint under Adam 96, or 128 keeping the best epoch's.
The measures a network is held to are the pipeline's, declared before any number exists: .Report(...) names them —
RMSE, MAE and R² for amounts; accuracy, precision, recall and the confusion matrix for classes — the parts they are taken
on, and how they are shown. Pipeline has the rules. Each measure is in the answer's own units, beside what predicting
the training rows' average would score, by scikit-learn's definitions; the network's measures are taken on the engine its
run was handed, by the same evaluation it serves with.
trained.Measures!.Report(), in DeepSharp.Charts, is the report as it says its measures are shown — the numbers, the
charts or both — as one value whose ToHtml() gives its HTML, so a notebook's C# cell that ends with it shows the
report. What the report measured can also cross as text: trained.Measures!.PredictionsToJson() writes, for each part,
the key of each row and what was predicted for it, with the fit the predictions were made behind, and
prepared.MeasureAgain(text) measures that text again on another run of the same pipeline — refused beside any other
fit, as the network's file is. That is how a notebook's report block draws a network a C# cell trained: Notebook.
DeepSharp.Charts draws them, as the text of an SVG, drawn with MatPlotLibNet:
history.LossCurve() |
The training loss of every epoch, and the validation loss beside it. |
history.LearningRates() |
The rate every epoch took: the schedule. |
measures.Bars() |
Every measure, each part's beside the average's. |
measures.ConfusionMatrices() |
Each confusion matrix as a heatmap of counts, a row to each class the rows held. |
measures.PredictedAgainstActual() |
What was predicted against what was there, part by part, beside the line where they agree. |
measures.Residuals() |
What was left over against what was predicted. |
correlation.Heatmap() |
A correlation, on the whole of its scale. |
A network's file holds the kinds it is made of — each by the name it is registered under, with the settings it is
rebuilt from — every slot's numbers by its path, and its loss. NetworkDocument writes and reads it, through a
NetworkCatalog of the kinds a reader knows; a network written as code is saved by a name of its own — it implements
ISaved<T> and is registered with NetworkCatalog.BuiltIn().Register<Passenger>() — and nothing is ever made from a
type looked up by its name. Read back, it is the same network to the last bit. Everything wrong with a file is said at
once, each at its line and column:
(9,13): 'dens' is not a kind this catalog knows. The nearest one it knows is 'dense'.
(15,7): '0.bias' is missing: every slot of the network is written.
(17,9): '0.weight' is a 4x1 slot here, and is written as 2x4x1.
Numbers trained somewhere else go into a network the same way: network.Load(entries) puts each into the slot its path
names, all of them or none, as PyTorch's load_state_dict does with strict=True, and refuses a missing, extra or
misshapen slot in the same words as the file does. A model PyTorch, Keras or an ONNX exporter saved is read so, into a
network written here or one its file describes — Importing a model has each reader.
- A schedule that watches the validation loss. The rows a model is chosen on would then shape the weights it is chosen for.
- Weight decay, Nesterov momentum, AMSGrad, AdamW. Each does nothing by default in PyTorch, and nothing asks for one yet.
- Hand-written GPU code. A graphics card is reached through an engine that already reaches one: libtorch, through the TorchSharp backend.