Repository navigation
Reading a file
The first step of the course names where the rows come from. Nothing is decided here and nothing is learned: the reader hands every cell over as text, and the next step says what each column holds.
using DeepSharp.Pipelines;
var proposal = Pdd.Create()
.ReadCsv("titanic.csv")
.ProposedKinds(); // every cell read, a kind proposed for each column
foreach (var column in proposal.Columns)
{
Console.WriteLine($"{column.Name}: {column.Kind}");
}ProposedKinds() is there because a schema typed from nothing is a schema typed from memory. It reads the file and
proposes a kind for each column — alive true or false, embarked a category, age a number — and it is a proposal and
nothing more. You write the schema; it just saves you from opening the file in a spreadsheet first.
A comma-separated file is one reader of several. Parquet, Excel workbooks and JSON each have their own, in a package of their own, so a project that reads only CSV carries only CSV. Which readers exist, and what each one does with a column it cannot make sense of, is on Pipeline.
What the row looks like after this step. The passenger whose age the file leaves empty arrives as an empty cell —
not a nought, not a zero-length word, an absence. It stays an absence until something settles or fills it, and every step
between here and there can see that it is one.
Next: Declaring the columns · Every reader and its options: Pipeline