Repository navigation
Declaring the columns
The schema is the one place that says what takes part. A column you do not name does not come along, and a column you do name is read as the kind you gave it or the run stops. This is deliberate: a file that grows a column next month must not quietly grow a feature.
using DeepSharp.Pipelines;
var prepared = Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema
.Integer("survived", "sibsp", "parch") // whole numbers
.Category("pclass", "sex") // words from a fixed set
.Optional("age", ColumnKind.Number) // a number, and it may be absent
.Number("fare")) // a number, and it may not
.SplitStratified("survived", train: 0.70, validation: 0.15)
.Target("survived")
.Build()
.Run();
Console.WriteLine($"{prepared.Table.Columns.Count} columns of {prepared.Table.RowCount} rows");Optional is the word that matters here. It says a cell may be empty — and only an optional column may be. A gap in
a column you declared Number stops the run at the row it found, with the row's number, because a number that is
sometimes absent is a different thing from a number, and a model trained on one and served the other is the bug this
whole library is shaped around.
The columns the file holds and the schema does not are dropped by default. Saying so differently, keeping them, or refusing a file that brings an unknown column, is one word on the schema — on Pipeline.
What the row looks like after this step. The empty age is now an absence the pipeline expects: declared optional,
still empty, and every later step knows the difference between "nobody wrote it down" and "it happened to be nought".
Next: Settling the gaps · Every kind and what refuses what: Pipeline