Repository navigation
Searching for a network
Which learning rate, how many units, which activation — a network has numbers nobody can work out in advance, and the usual way to find them is to try several and keep the one that did best. A study does that, and does it by the rule this library is built on: every candidate is trained on the rows the pipeline prepared, chosen between by the validation rows, and never looked at on the rows it will be measured on until one has been chosen.
using DeepSharp.Learners.Networks;
using DeepSharp.Pipelines;
var space = new SearchSpace(
[
new LogRange("rate", 1e-3, 1e-1), // spread across the decades, as a learning rate is
new WholeRange("units", 4, 32),
new Choices("activation", ["relu", "tanh"]),
]);
var study = new Study(space, (candidate, network) =>
{
var layered = network.Dense(candidate.Whole("units"));
layered = candidate.Choice("activation") == "relu" ? layered.Relu() : layered.Tanh();
return layered.Dense(1).Adam(candidate.Number("rate")).BinaryCrossEntropy().Run(seed: 20260929, epochs: 10);
})
{
Trials = 12,
Deals = 3,
Sampler = new RandomSampler(11),
};
var found = study.Run(deal => Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema.Integer("survived", "sibsp", "parch").Category("sex").Optional("age", ColumnKind.Number).Number("fare"))
.SplitAtRandom(0.70, 0.15, seed: 100 + deal, testSeed: 7) // another deal of train and validation, the same test rows
.Target("survived")
.FillMissing("age", With.Median)
.EncodeCategories()
.Normalise("age", Scale.MidRange)
.Normalise("fare", Scale.MidRange)
.Normalise("sibsp", Scale.MidRange)
.Normalise("parch", Scale.MidRange)
.Report(report => report.Measure(Metric.Accuracy).On(Part.Validation, Part.Test).As(Shown.Numbers))
.Build());
foreach (var trial in found.Trials)
{
Console.WriteLine($"{trial.Number,2} {trial.Score:0.000} ± {trial.Spread:0.000} rate {trial.Candidate.Number("rate"):0.0000}, {trial.Candidate.Whole("units")} units");
}
var best = found.Best.Declared; // the winner, as a step a pipeline keeps in its file
var onTest = found.Winner[0].Measures!.Parts.Single(part => part.Part == Part.Test);
Console.WriteLine($"{best.Epochs} epochs, accuracy {onTest.Values[0].Value:0.000} on {onTest.Rows} rows nobody chose by");A candidate is a point of a space, and the network it declares is written by you. The space names what may vary — a number
between two bounds (NumberRange), a number that spreads across decades (LogRange), a whole number (WholeRange), one of a
few words (Choices) — and the lambda turns the values a trial drew into the network, in the same chain .WithTorch(…) uses.
Nothing about the network's shape is fixed by the study: a layer you add when units is above sixteen is only code. A
network of your own, written as a step, is a candidate too: give the study a function from a candidate to a LearnNetworkStep.
Validation chooses; the test rows are for the winner alone. Each candidate's score is its validation loss, the mean of the
deals when there are several, with how much it moved from one deal to the next. The study never trains on the validation
rows and never reads a trial's test rows: a trial keeps no network, and only the winner's networks — found.Winner, one for
each deal — are kept, with the report's measures of the test rows. A study that chose by the test rows, even once, would
have measured a number it had picked the best of, and a best of twelve is flattered.
A deal is another division of the rows into training and validation, and the test rows must not move. One division
cannot separate a better network from a luckier split: the same network scores a few percent differently on two divisions
of the same rows. So the study can prepare the pipeline again, dealt another way, for each deal, and judge a candidate by the
mean over them. But a split at random deals every part from one seed, so a second seed would also deal another test part, and
a candidate would be chosen on rows another deal learned from. SplitAtRandom(train, validation, seed, testSeed) fixes the
test rows with testSeed and lets seed deal the rest; with more than one deal, a study whose deals do not share one test
part is refused, saying so.
A study that takes hours can be watched and stopped. OnTrial is handed every trial as it ends, in order, with how it
was judged; Run(pipelineFor, token) stops before the next trial, and a trial before its next batch, by throwing an
OperationCanceledException that carries the token — nothing is returned, because a study cut short has not found
what it set out to find. Hearing the trials changes no number.
The same study run again is the same numbers. Every number a trial draws is worked out from the sampler's seed and the trial's place, so trial five is the same candidate whether the study has ten trials or a hundred, and a dimension added to the space leaves the values of the others where they were. Trials run one after another and share nothing they change.
Random search is what ships, and a smarter one plugs in. A study asks an ISampler for the candidate of each trial and
hands it the trials that finished before, so a sampler that learns from them — a tree of Parzen estimators, say — can be
written in any assembly and given to the study. Random search is the baseline a cleverer one has to beat, and over a few
dozen trials it often does not. There is no early stopping of a weak trial, because a trial's loss halfway is not what it
ends at, and none of the study chooses by anything but a finished network.
A second pipeline is compared with the first in pairs. Whether a change helped is the same question as which
network is best, one level up, and it has the same trap: a difference measured on one division is as large as the noise of
the division. A Comparison trains one network behind two pipelines on each of many deals and reports, deal by deal, how
far the second sits from the first.
using DeepSharp.Learners.Networks;
using DeepSharp.Pipelines;
Pipeline Passengers(int deal, bool stratified)
{
var declared = Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema.Integer("survived", "sibsp", "parch").Category("sex").Optional("age", ColumnKind.Number).Number("fare"));
var split = stratified
? declared.SplitStratified("survived", 0.70, 0.15, seed: 100 + deal)
: declared.SplitAtRandom(0.70, 0.15, seed: 100 + deal);
return split
.Target("survived")
.FillMissing("age", With.Median)
.EncodeCategories()
.Normalise("age", Scale.MidRange)
.Normalise("fare", Scale.MidRange)
.Normalise("sibsp", Scale.MidRange)
.Normalise("parch", Scale.MidRange)
.Build();
}
var comparison = new Comparison(network => network.Dense(8).Relu().Dense(1).Adam(0.01).BinaryCrossEntropy().Run(seed: 1, epochs: 10))
{
Deals = 10,
};
var effect = comparison.Run(deal => Passengers(deal, stratified: false), deal => Passengers(deal, stratified: true));
Console.WriteLine($"{effect.Mean:0.0000} ± {effect.Spread:0.0000} over {effect.Deals} deals (standard error {effect.StandardError:0.0000})");On the passenger list the tutorial reads, that prints a mean of −0.0069 with a spread of 0.0378 over ten deals: whether the share of survivors is kept the same in every part changes nothing a network of this size can show.
The difference is the second pipeline's score less the first's, so with a loss for a score a positive mean says the second is worse. It is read against its own spread: a mean of 0.01 with a spread of 0.05 is a deal's worth of luck, and a mean of 0.01 with a spread of 0.002 is a difference. The two pipelines may differ in anything — a feature left out, a gap filled another way, a split at random against a split by time — and a change that helps on one split and hurts on another shows as a spread larger than its mean, which is what a feature that knows the future of the rows it is measured on does.
Next: Networks has the layers, the optimizers and the loop a candidate is made of · Splitting the rows has the test seed