Repository navigation
Splitting the rows
The split is the line the library is built around. Above it: everything that is the same for every row, whoever reads it. Below it: everything learned from data, and only the training rows may teach it.
using DeepSharp.Pipelines;
var passengers = Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema.Integer("survived").Optional("age", ColumnKind.Number).Number("fare"))
.SplitStratified("survived", train: 0.70, validation: 0.15) // test is what is left
// ---- nothing above this line learns from the rows ----
.Target("survived") // the answer, named right below the line
.FillMissing(fill => fill.Median("age"))
.Normalise("age", "fare")
.Build()
.Run();
Console.WriteLine($"{passengers.CountIn(Part.Train)} train, {passengers.CountIn(Part.Validation)} validation, {passengers.CountIn(Part.Test)} test");Seventy to learn from, fifteen to choose with, and what is measured on is never written down — it is what is left. Naming the test share as well is allowed and says the same thing twice; nothing can ask for more rows than there are.
Which split, and why it matters more than it looks. A stratified split keeps the shares of an answer the same in every part, which is what you want when one answer is rare. A split by time gives the earliest rows to learn from and the latest to be measured on, and takes a gap of days between the parts so that a row just before the boundary cannot leak into the part just after it through a feature that looks ahead. Using a random split on a series in time is the most common way to get a wonderful validation number and a worthless model.
What the split writes down. How many rows went to each part, and a digest of the rows it divided that does not depend on the order they arrived in — so what a fit saw travels with what it learned, and a replay can be held to it.
A file's own order is not a mixture. Rows written every survivor first, or every month in turn, give parts that are
not alike. .Shuffle(seed) above the split puts them in an order drawn from that seed so each part holds the same
mixture; without it they keep the order they were read in. A series in time is the exception that proves it: there you
write .OrderBy, and the split by time wants exactly that order.
Rows are divided here and never taken out below. Dropping a row below the split would quietly change how many each part holds, which is the one thing the split promised. That is why every verb that removes rows stands above it.
Next: Naming the answer · All three splits and the predict share: Pipeline