Skip to content

Filling and scaling

H.P. Gansevoort edited this page Oct 5, 2026 · 3 revisions

Step 7 — Filling and scaling

Everything in this step learns a number from the training rows, replays it unchanged on validation, on test and on every row served a year later, and writes it into the file so it can be replayed at all. That is why it stands below the split, and it is the half of the course that a hand-written script almost always gets wrong.

using DeepSharp.Pipelines;

var prepared = Pdd.Create()
    .ReadCsv("titanic.csv")
    .Declare(schema => schema.Integer("survived", "sibsp").Category("pclass", "sex").Optional("age", ColumnKind.Number).Number("fare"))
    .SplitStratified("survived", train: 0.70, validation: 0.15)
    // ---- nothing above this line learns from the rows ----
    .Target("survived")                              // the answer, named below the line and before what learns
    .FillMissing(fill => fill.Median("age"))         // Mean, Median, Zero, Constant, Previous, Refuse
    .EncodeCategories()                              // every category column the schema named
    .Normalise(scale => scale                        // a kind a group of columns
        .Columns("age", "fare")                      // the default: between minus one and one
        .MaxAbs("sibsp"))
    .Build()
    .Run();

Console.WriteLine(prepared.Fitted[4].Number("value"));   // the median the training rows gave

One line, many columns, one kind each. Every per-column verb here takes a line of its own — .FillMissing, .Normalise, .Reshape, .ClipOutliers, .FillNaN, .Cyclical, .TimeParts — and what reaches the declaration is one step a column either way, so the file reads the same as if you had written the verb once per column.

Where features land is said once. A scaling that names no kind puts them between minus one and one, which is where a network takes its features. .DefaultFeatures(Form.Unit) says nothing to one instead, once, above the split, and a column that names its own kind keeps it.

What Refuse is for. Several of these verbs offer a refusal as a strategy, and .FillNaN refuses by default. A not-a-number is somebody's broken division, and carrying it into a model as though it were a measurement is the one thing a pipeline should never do quietly.

A scaling whose bounds you give stands above the split. Everything on this page reads its numbers from the training rows, which is why it is here. .ScaleGiven("fare", 0, 512) is told them instead — a fare runs from nothing to the most anybody paid, an hour from nought to twenty-three — and that is knowledge about the column rather than about the rows. So it stands where the features are worked out, a feature after it is already on the scale a model takes, it writes nothing into the fitted half, and the way back is exact without anything written down. One line does many columns: .ScaleGiven(scale => scale.Between("fare", 0, 512).Between("age", 0, 100)).

A learner that does without some of this still gets it declared. A tree takes numbers of any size, so a run made for one leaves the scalings out — and writes down which it left out, in the file, so what that learner was fed is never a guess. Declare them anyway: the same declaration then trains a network too, and the two are measured on the same rows. What is never left out is the way back, since an answer has to come back in its own units whoever predicted it.

The sharp edge, once more. A column settled above the split has no gap left to fill, so a fill of it here is refused at the line — see Settling the gaps.

Next: Measuring and drawing · Every strategy, every scale, every encoder: Pipeline

Clone this wiki locally