Skip to content

Starting from a course

H.P. Gansevoort edited this page Oct 5, 2026 · 2 revisions

Starting from a course

A pipeline written from nothing starts with an empty chain and a reader who has to know which verb comes where. A course says it: every step of a prepared pipeline, in the order the steps belong, each waiting for what only you know — the file, the columns, the bounds — and starting everything else as its verb starts it. You fill in rather than remember.

There are two, because a table and a series in time do not take the same steps: the rules refuse some of the choices, so no one list can teach both.

The course for a table

For rows that do not depend on one another: a passenger list, a customer file. PipelineCourse.Table.

step waits for starts as
1 read.csv path
2 declare columns remainder: drop
3 settle.gaps column with: zero
4 feature.add column, left, right arithmetic: minus
5 scale.given column, lowest, highest lands: signed
6 split.stratified column the shares 0.7, 0.15 and 0.15, and a seed
7 target column
8 drop.columns columns
9 fill.missing column with: median
10 normalise column scale: midrange, outOfRange: pass
11 evidence.report nothing the root-mean-square error on the validation and test rows, as numbers
12 learn.network nothing a dense layer of sixteen and a relu, Adam, mean squared error, a hundred epochs

The course for a series in time

For rows that follow one another in time: prices, readings, anything where yesterday is not independent of today. PipelineCourse.SeriesInTime is the table's course with three differences the rules make.

step waits for starts as
1 read.csv path
2 declare columns remainder: drop
3 order.by columns
4 settle.gaps column with: zero
5 feature.add column, left, right arithmetic: minus
6 scale.given column, lowest, highest lands: signed
7 split.byTime column the shares 0.7, 0.15 and 0.15, and a gap of 1
8 target.ahead column ahead: 1, as: value
9 drop.columns columns
10 fill.missing column with: median
11 normalise column scale: midrange, outOfRange: pass
12 evidence.report nothing the root-mean-square error on the validation and test rows, as numbers
13 learn.network nothing a dense layer of sixteen and a relu, Adam, mean squared error, a hundred epochs

The rows are put in order first, because a series is read in its order and shuffling it would leak. They are divided along the clock rather than at random, with a gap of one moment kept apart. And the answer is the one that reads ahead, which refuses a split without a gap at least as wide as how far it reads — which is why the course says gap where it would otherwise leave it out. Make it as wide as the answer reads.

What waits, and what does not

What names something of yours waits as nothing: a column, the columns, a file, the columns a schema declares, a bound. An example of such a value is not a default — a column called column reads as well as any other name, so a step that kept it would pass for one somebody chose. A name a step may leave out stays out, which is how the step decides it itself. Everything that only settles how — a seed, a share, a word from a set, a number that is not a bound of your column — starts as its verb starts it, because a step you are asked to fill in from memory is not one you were helped with.

Why the steps stand where they do

  • The gaps, then the features. A feature worked out from a column with a gap is itself a gap, and what fills a gap is learned from the training rows, so a fill cannot stand where the features are worked out. Settling puts in a value no row decided, so it can. See Settling the gaps.
  • The features, then a scale from bounds, both above the split. A feature is worked out from the columns as they were read, so its arithmetic still means what it says; a scale into minus one to one from bounds you know then learns nothing from the rows, so it stands with the features rather than below the split with what learns. See [[Working out the features]].
  • The split, then the answer. The answer stands right below the split because a return must be made from its column as it was read, which no step above it may have changed; there it is valid whatever the answer is, and the ten pages of Getting Started name it there too. See Naming the answer.
  • What is not needed is dropped before what learns. The steps that learn from the training rows are not asked to work on columns that will not be there. A drop after them is possible too, and no smarter.
  • A gap is filled, then the column is scaled. The scale is learned from numbers, and a gap is not one. See [[Filling and scaling]].
  • The report, then the network. What a model is held to is said before the model is named. See [[Measuring and drawing]] and Naming the network.

The order is written down once, by the course, and held by tests. Filled in for the passenger list and for the price series, each course breaks no rule. The rules themselves fix only a few neighbours — the source before the schema, the schema before the settling and, for a series, the order after the schema and the split before the answer that reads ahead — because among ten steps most orders break nothing; the rest is the flow above, and each of its pairs is said by a test beside its reason.

In a notebook

Two buttons in the toolbar, Course for a table and Course for a series, write every step of a course the notebook does not hold yet, each as a block, in one turn. A block starts as its verb's skeleton, with what waits as null:

{
  "step": "normalise",
  "column": null,
  "scale": "midrange",
  "outOfRange": "pass"
}

A block that waits does not read as a step, so it says which key it waits for and nothing else, and it takes no part in the pipeline the blocks above it make. The pipeline is the longest run of blocks from the top that makes a declaration, so the blocks below a block that waits are not judged: a block nobody has touched never shows a rule broken. Fill them in from the top; a block starts to read the moment the last key it waits for is said, and then its properties panel takes it over — until then it is edited as text.

A course goes on from where the blocks stand. It never writes over a block, never moves one, and puts its steps right after the last block that is a step; text between the blocks stays where it is. A step the notebook left out is not written back: a pipeline that went straight from the schema to its split has decided it needs no settling, no features and no scale — and when a later step needs one that is missing, the rule that needs it says so at that block, as every rule does. A block of another kind of the same thing — a Parquet file where the course reads a comma-separated one, a split along the clock where the course divides at random — stands for the step; a shuffle is not the order of a series, so it does not make a notebook a series'. A block that drops columns says the course's drop wherever it stands, before what learns or after it, and neither moves the course on nor takes it away. A notebook that puts its rows in order is a series', so only the series' course is offered to it. A notebook whose steps stand in another order than the course teaches, or that holds the whole of it already, is offered none, and pressing a button twice writes the course once.

In DeepSharp's own host — an application that opens notebooks through DeepSharp.Verso.Api, and deepsharp-serve — the add button does the same one block at a time: a pipeline block arrives as the step of the course that belongs where it is added, and starts empty when no step belongs in that place. A new notebook starts with the course's first step, waiting for its file. Verso's own editors never ask a cell type what a cell starts with, so a block added there starts empty, as every cell does.

In code

A course is a value. Say returns the course with more said of a step, Waiting says what still waits, and Pdd.From starts a pipeline from a course filled in all the way:

using System.Text.Json.Nodes;
using DeepSharp.Learners.Networks;
using DeepSharp.Pipelines;

var catalog = StepCatalog.BuiltIn().WithNetworks();

var course = PipelineCourse.Table
    .Say("read.csv", new JsonObject { ["path"] = "titanic.csv" });

foreach (var waiting in course.Waiting(catalog))
{
    Console.WriteLine(waiting);       // Step 2, 'declare': waits for what only you can say: 'columns'.
}

course = course
    .Say("declare", JsonNode.Parse("""
        {"columns": [{"name": "survived", "kind": "integer", "optional": false},
                     {"name": "sibsp", "kind": "integer", "optional": false},
                     {"name": "parch", "kind": "integer", "optional": false},
                     {"name": "fare", "kind": "number", "optional": false},
                     {"name": "age", "kind": "number", "optional": true}]}
        """)!.AsObject())
    .Say("settle.gaps", new JsonObject { ["column"] = "fare" })
    .Say("feature.add", new JsonObject { ["column"] = "family", ["left"] = "sibsp", ["arithmetic"] = "plus", ["right"] = "parch" })
    .Say("scale.given", new JsonObject { ["column"] = "fare", ["lowest"] = 0, ["highest"] = 512 })
    .Say("split.stratified", new JsonObject { ["column"] = "survived" })
    .Say("target", new JsonObject { ["column"] = "survived" })
    .Say("drop.columns", new JsonObject { ["columns"] = new JsonArray("sibsp", "parch") })
    .Say("fill.missing", new JsonObject { ["column"] = "age" })
    .Say("normalise", new JsonObject { ["column"] = "age" });

var pipeline = Pdd.From(course, catalog);      // refused, with every step that still waits, until nothing does

A course not filled in all the way is refused whole, naming every step that still waits and what it waits for, every verb the catalog does not know, every key a verb does not take, and — through the verb's own words — every value a step that waits for nothing any more refuses. What is said is read as a file's steps are read, so a verb's own rules hold here, and the rules every declaration keeps hold for the steps together.

A pipeline rarely needs every step, and some steps come once a column. Without leaves the steps of some verbs out, and Also adds one more step of a verb the course holds, right after the last one of it; PipelineCourse.Named lists both courses. A scale is written once a column, so a course says them one at a time:

using System.Text.Json.Nodes;
using DeepSharp.Pipelines;

var plain = PipelineCourse.Table
    .Without("settle.gaps", "feature.add", "scale.given", "drop.columns")      // a pipeline that needs none of these
    .Say("normalise", new JsonObject { ["column"] = "age" })
    .Also("normalise", new JsonObject { ["column"] = "fare" });                // one step a column

foreach (var step in plain.Steps)
{
    Console.WriteLine(step.Verb);       // read.csv, declare, split.stratified, target, fill.missing, normalise, normalise, …
}

As a file

A course is a file of its own, as the decisions about a pipeline's columns are: what has been said of each verb, in order, with the version of the pipeline file it was written against.

File.WriteAllText("titanic.course.json", course.ToJson());

var read = PipelineCourse.FromJson(File.ReadAllText("titanic.course.json"), catalog);

Console.WriteLine(read == course);             // True
{
  "version": 7,
  "course": [
    { "step": "read.csv", "path": "titanic.csv" },
    { "step": "declare" }
  ]
}

A step holds its verb and the keys that have been said; a key left out, or written as null, has not been said. The file is read through the same door as a pipeline file, with the catalog of the verbs it may name, every fault at its line and column, and a file written by a newer DeepSharp refused whole. It is not a pipeline file and it is not the saved columns of a notebook: it holds no declaration and no fitted, and a file of the wrong kind is refused by name.

A step that still waits is judged by its keys alone, whether the verb takes them: a value is judged beside the keys it is tied to — a lower bound beside an upper bound that is still to be said — so it is judged once nothing the step waits for is missing. A course written half way reads back as the course it was, and a step with everything said is refused by its verb in the verb's own words.

What it leaves alone

  • A pipeline's own file. The pipeline file stays at version 7 and says nothing of having begun as a course: what a person decided is what the declaration says.
  • The saved columns. A course does not replace what a pipeline decided about its columns, and does not fill itself from them: the saved columns hold the schema, the drops and the output, a course the whole order, and each stands beside the other.
  • Saving your own course from a notebook. A course is written once and spent, so nothing writes one for you; write the file by hand, or build the value with PipelineCourse.Of.

Next: back to Getting Started, or every verb and every rule in Pipeline

Clone this wiki locally