Skip to content

Reading a file

H.P. Gansevoort edited this page Oct 4, 2026 · 1 revision

Step 1 — Reading a file

The first step of the course names where the rows come from. Nothing is decided here and nothing is learned: the reader hands every cell over as text, and the next step says what each column holds.

using DeepSharp.Pipelines;

var proposal = Pdd.Create()
    .ReadCsv("titanic.csv")
    .ProposedKinds();                      // every cell read, a kind proposed for each column

foreach (var column in proposal.Columns)
{
    Console.WriteLine($"{column.Name}: {column.Kind}");
}

ProposedKinds() is there because a schema typed from nothing is a schema typed from memory. It reads the file and proposes a kind for each column — alive true or false, embarked a category, age a number — and it is a proposal and nothing more. You write the schema; it just saves you from opening the file in a spreadsheet first.

A comma-separated file is one reader of several. Parquet, Excel workbooks and JSON each have their own, in a package of their own, so a project that reads only CSV carries only CSV. Which readers exist, and what each one does with a column it cannot make sense of, is on Pipeline.

What the row looks like after this step. The passenger whose age the file leaves empty arrives as an empty cell — not a nought, not a zero-length word, an absence. It stays an absence until something settles or fills it, and every step between here and there can see that it is one.

Next: Declaring the columns · Every reader and its options: Pipeline

Clone this wiki locally