Skip to content

Splitting the rows

H.P. Gansevoort edited this page Oct 5, 2026 · 3 revisions

Step 5 — Splitting the rows

The split is the line the library is built around. Above it: everything that is the same for every row, whoever reads it. Below it: everything learned from data, and only the training rows may teach it.

using DeepSharp.Pipelines;

var passengers = Pdd.Create()
    .ReadCsv("titanic.csv")
    .Declare(schema => schema.Integer("survived").Optional("age", ColumnKind.Number).Number("fare"))
    .SplitStratified("survived", train: 0.70, validation: 0.15)   // test is what is left
    // ---- nothing above this line learns from the rows ----
    .Target("survived")                                           // the answer, named right below the line
    .FillMissing(fill => fill.Median("age"))
    .Normalise("age", "fare")
    .Build()
    .Run();

Console.WriteLine($"{passengers.CountIn(Part.Train)} train, {passengers.CountIn(Part.Validation)} validation, {passengers.CountIn(Part.Test)} test");

Seventy to learn from, fifteen to choose with, and what is measured on is never written down — it is what is left. Naming the test share as well is allowed and says the same thing twice; nothing can ask for more rows than there are.

Which split, and why it matters more than it looks. A stratified split keeps the shares of an answer the same in every part, which is what you want when one answer is rare. A split by time gives the earliest rows to learn from and the latest to be measured on, and takes a gap of days between the parts so that a row just before the boundary cannot leak into the part just after it through a feature that looks ahead. Using a random split on a series in time is the most common way to get a wonderful validation number and a worthless model.

What the split writes down. How many rows went to each part, and a digest of the rows it divided that does not depend on the order they arrived in — so what a fit saw travels with what it learned, and a replay can be held to it.

A file's own order is not a mixture. Rows written every survivor first, or every month in turn, give parts that are not alike. .Shuffle(seed) above the split puts them in an order drawn from that seed so each part holds the same mixture; without it they keep the order they were read in. A series in time is the exception that proves it: there you write .OrderBy, and the split by time wants exactly that order.

Rows are divided here and never taken out below. Dropping a row below the split would quietly change how many each part holds, which is the one thing the split promised. That is why every verb that removes rows stands above it.

Next: Naming the answer · All three splits and the predict share: Pipeline

Clone this wiki locally