Repository navigation
Saving and replaying
A pipeline is saved as one file with two halves, because they have two different authors: the declaration, which is what you wrote, and the fitted half, which is what the fit learned. Keeping them apart is what lets the same declaration be fitted again on fresh data, two runs be compared by diffing the declaration alone, and a serving process load the file without ever knowing a builder existed.
using DeepSharp.Pipelines;
var prepared = Pdd.Create()
.ReadCsv("titanic.csv")
.Declare(schema => schema.Integer("survived").Optional("age", ColumnKind.Number).Number("fare"))
.SplitStratified("survived", train: 0.70, validation: 0.15)
.Target("survived")
.FillMissing(fill => fill.Median("age"))
.Normalise("age", "fare")
.Build()
.Run();
var text = prepared.ToJson(); // both halves
var again = PreparedData.FromJson(text, StepCatalog.BuiltIn()); // anywhere, no data needed
var declared = PipelineDeclaration.FromJson(text, StepCatalog.BuiltIn()); // the person's half alone
Console.WriteLine($"{declared.Steps.Count} steps, version {PipelineDeclaration.Version}");A fit is tied to the steps it was fitted behind. Every fitted entry carries a key made from its own step and every step above it, so a fit is never used under steps that changed after it was learned. Read back by position instead, a fit spliced under another declaration once served a price of 135.7 where 0.9048 was meant — which is why it is keyed by content and not by place.
Replaying is not re-fitting. Replay(rows) walks the same steps over rows nobody had seen, with the numbers the
training rows produced and nothing fitted again — it puts the rows in order and drops the warm-up rows exactly as the run
did. Served(rows) hands rows over without an answer, because a served row is the question, and says which handed-in row
each one is so a prediction finds its way back. BackToOriginal(…) puts the predictions back in your own units.
A file from an older DeepSharp still reads. The file names its version, and each verb the version from which it means what it says now. A file newer than the library is refused whole rather than half-understood. A model's own file carries its pipeline's text exactly as that text was written and is tied to it by a digest over that text, so the version the text names never changes and a model trained behind an older pipeline keeps loading and keeps serving. Two pipelines are compared fit for fit by a second digest, and that is the one that leaves the version out.
A model from another library says more. A trainer from ML.NET produces that library's own archive, which only it opens, so the file carrying it also names the version of ML.NET that wrote it and the processor it was written on — a reader that cannot open the archive can then say which package and which version would. Its bytes are never what ties the file to anything: they carry the clock they were saved at, so two saves of one model differ while the model does not.
Where to go next. The same pipeline written as a notebook, one block per step, with the data and a profile at any block: Notebook. Every verb in the order you write them: Pipeline. Why the course is shaped this way: PDD.