Skip to content

Importing a model

H.P. Gansevoort edited this page Oct 1, 2026 · 1 revision

Importing a model

A model does not have to begin here to run here. Three packages read what other frameworks saved into the same network DeepSharp writes — to serve, to measure, or to train further behind a pipeline and keep with it as the one file:

dotnet add package DeepSharp.Import.PyTorch
dotnet add package DeepSharp.Import.Keras
dotnet add package DeepSharp.Import.Onnx
DeepSharp.Import.PyTorch SafetensorsFile, TorchSaveFile A network's state PyTorch saved, as a safetensors file or as the .pt file torch.save writes, into a network written here.
DeepSharp.Import.Keras KerasFile A model Keras 3 saved, as its .keras archive or the .h5 file it saved one to before, built as the file describes it.
DeepSharp.Import.Onnx OnnxFile An ONNX graph — PyTorch's, Keras's or tf2onnx's — lowered onto the layers it is.

Each is an IImporter, and importer.Read(stream) hands back a SavedNetwork, as reading a network's own file does: the network, every slot holding the file's numbers, and the loss it answers through. A file of numbers alone is read into the network it is handed; a file that describes its network as well has that network built.

Numbers go in by path

Every importer ends in one load. network.Load(entries) puts a tensor into each slot its path names, as PyTorch's load_state_dict does with strict=True: every slot of the network once, none it does not have, each of its slot's shape and holding finite numbers. A batch normalisation's scale, shift and statistics go in as any slot's do:

using DeepSharp.Networks;
using DeepSharp.Tensors;

var stream = new RandomStream(42);
var network = new LayerStack(new Dense(2, 2, stream.Draw("initialise:0", 0, 0)), new BatchNorm(2));

network.Load(
[
    new SlotEntry("0.weight", Tensor.From(new Shape(2, 2), [1f, 0.5f, -0.5f, 2f]), "line 1"),
    new SlotEntry("0.bias", Tensor.From(new Shape(2), [0.25f, -0.25f]), "line 2"),
    new SlotEntry("1.weight", Tensor.From(new Shape(2), [1.5f, 0.5f]), "line 3"),
    new SlotEntry("1.bias", Tensor.From(new Shape(2), [0.125f, -0.125f]), "line 4"),
    new SlotEntry("1.running_mean", Tensor.From(new Shape(2), [0.5f, -1f]), "line 5"),
    new SlotEntry("1.running_var", Tensor.From(new Shape(2), [4f, 0.25f]), "line 6"),
]);

The third value of each is where its file holds it, in the file's own words — a tensor's name, a dataset's path, a line — because a file of numbers has no lines and columns to name a fault by. It puts in all of them or none: whatever is wrong is refused at once with a SlotLoadException, whose Faults name each where its file holds it, and the network keeps what it held:

line 2: '0.bias' is a 2 slot here, and is written as 3.
line 3: '1.weight' holds finite numbers, and its value at 1 is NaN.
line 5: '1.bias' is written twice here, and only one of the two would be read.
line 7: '2.weight' is no slot of this network.
'1.running_var' is missing: every slot of the network is written.

A network's own file is read through the same load, so it refuses a missing, extra or misshapen slot in the same words. Every name a file holds is shown in a refusal as Quoted() shows a file's text: a line break, a control character or a character that turns text round is written as its escape, and a name past 200 characters is cut short with its length, so a file cannot write a line of its own into a log.

PyTorch

A file PyTorch saved holds a network's numbers and not the network, so the reader is handed the network they belong to and the loss it answers through: a stack — in Keras's words or written by hand, its layers numbered as PyTorch's Sequential numbers its modules — or a network written as code, its layers named as the PyTorch module named them.

using DeepSharp.Import.PyTorch;
using DeepSharp.Networks;
using DeepSharp.Tensors;

var titanic = new Sequential().Dense(16).Relu().Dense(1).Lower(new Shape(14), new RandomStream(7));

using var safetensors = File.OpenRead("titanic.safetensors");
var saved = new SafetensorsFile(titanic, new BinaryCrossEntropy()).Read(safetensors);

var images = new Sequential().Conv2D(4, new Window(3, 3)).Relu().Flatten().Dense(1)
    .Lower(new Shape(28, 28, 1), new RandomStream(7));

using var state = File.OpenRead("images.pt");                       // torch.save(model.state_dict(), "images.pt")
var read = new TorchSaveFile(images, new MeanSquaredError()) { Example = new Shape(28, 28, 1) }.Read(state);

SafetensorsFile reads a safetensors file and TorchSaveFile the archive torch.save has written since PyTorch 1.6; both are a PyTorchFile, and put the same numbers into the same slots, bit for bit, whichever file they came in.

Layouts. How PyTorch lays out each number is read off the layer that holds its slot, never off the number's shape — a square matrix turned round is as square as one that is not. A linear layer's weights are kept outputs by inputs there and inputs by outputs here, so they are turned round; a convolution's kernel is kept channels out, channels in, rows, columns there, and rows, columns, channels in, channels out here; biases and what a normalisation keeps are alike; and num_batches_tracked, which PyTorch keeps beside a batch normalisation's statistics, is left out by its name.

A flatten. PyTorch lays an image out channel by channel and DeepSharp place by place, so the rows a flatten makes of images are in another order there, and every number the next linear layer — and any normalisation before it — keeps for them is put in its place here. That takes knowing what the flatten is handed: an Example the network takes, which the reader runs through it as zeros, or Flattened, stating it by the flatten's path — for the network of images above, Flattened = new Dictionary<string, Shape> { ["2"] = new Shape(26, 26, 4) }. Told neither, the reader refuses those numbers rather than guess. Read without turning, the flattened rows of a test network landed 0.043 away from PyTorch's answers; turned, within 3e−8.

A network written as code shows the layers it holds, each with its path — layer.HeldLayers() lists them, as Slots() lists the numbers — but keeps to itself the order its forward pass runs them in. Its weights and kernels are turned as a stack's are, while whether one of its linear layers or normalisations reads rows made of images is never guessed: wherever such rows could reach one, it is read only as Flattened states it by that layer's path, and refused otherwise. A layer of a kind the reader does not know has none of its numbers read.

Numbers. Held in 16 bits — half precision or bfloat16 — they are widened exactly; in 64 bits, rounded to the nearest 32-bit float as PyTorch rounds them; whole numbers and truth values are refused.

A safetensors file is read by Onnxify.Safetensors, a port of safetensors' own reader, so a file that reader refuses is refused here in its words — of twenty written wrongly on purpose, the eighteen it refuses. A file whose note names another framework's layout than "format": "pt" is refused too, and one with no note is read as PyTorch's, since safetensors.torch.save_file and save_model write none unless handed one.

A .pt file. The pickle torch.save writes the state into is a program, whose instructions name functions and call them, so no pickle library reads it: an interpreter written here carries out only the instructions PyTorch's own weights-only reader, torch.load(weights_only=True), carries out, and builds nothing a file names but what a state dictionary is made of — an OrderedDict, a tensor rebuilt by torch._utils._rebuild_tensor_v2 or _rebuild_parameter, and the storages of a kind of number. Every other name is refused where the file names it, before anything is looked up, built or run. Tensors viewing their storage from any offset and with any strides, several sharing one storage, a file a big-endian machine wrote, and forty layers in a row, whose pickle numbers its memo past 255, are all read. Refused, each by name:

  • a pickle that is no dictionary of tensors by name — a checkpoint that keeps the state under a key of its own is read once model.state_dict() is saved on its own;
  • a torch.Size, bytes or a Counter, which PyTorch's reader allows and a state dictionary never holds;
  • the format before PyTorch 1.6, and a TorchScript archive;
  • a tensor saved as a negated or conjugated view — save the tensor itself — and one that repeats its storage's numbers.

Of 49 files written wrongly or with hostile intent on purpose, each handed to PyTorch's own reader first, every one it refuses is refused here too; of those it reads, six are refused here — a torch.Size, bytes, a Counter, the format before 1.6, a checkpoint under a key of its own and a negated view. What a file can make the reader do is bounded by the file and the network — Security says how.

Measured. The Titanic network PyTorch trained on the pipeline's training rows answers a man of 22 in third class within one rounding of a single-precision number of PyTorch's chance, and all 135 test passengers within five, written in Keras's words or as code; with a batch normalisation within ten, since PyTorch folds the normalisation into one scale and one shift where here a feature is normalised and then scaled. The files and what PyTorch answered are kept with the tests, with the scripts that made them.

Keras

using DeepSharp.Import.Keras;

using var file = File.OpenRead("titanic.keras");
var saved = new KerasFile().Read(file);

A Keras file describes its network as well as holding its numbers, so the reader builds the network from the description: each layer written in Keras's words and lowered as any description in those words is — a batch normalisation keeping Keras's epsilon and the complement of its momentum, a window padded as TensorFlow's 'same', the activation a Keras layer carries a layer of its own after it. The loss is the one the model was compiled with — a binary or a categorical cross-entropy, or a mean squared error — and the network ends where that loss takes over: a last sigmoid before a binary cross-entropy, or a last softmax before a categorical one, trained on the chances it gives, is lifted into the loss, as Keras's habit is; a model trained on its logits is read as it stands, and one whose end and loss disagree is refused at that layer. The numbers need no turning — Keras lays them out as the slots here keep them — and each layer's numbers are found by the name the file gives them, never by where they stand: Keras files them under the layer's class, layers/dense, while the layer is dense_2.

What is read: a Sequential model working in single precision, of Dense, Conv2D, BatchNormalization, LayerNormalization, Dropout, Flatten, Reshape, Activation and ReLU layers, with the linear, relu, tanh and sigmoid activations. What no network here is built of is refused at the layer that says it, every such layer at once: pooling and any other kind of layer or activation; a stride of its own for each axis, a dilated window, channels in groups, channels first; a layer without a bias; a normalisation over another axis than the last; a model of another kind than a Sequential, in another precision, or saved without its loss.

Measured. The networks sample's Titanic network, trained in Keras 3 on the pipeline's training rows and saved both ways, answers the 135 test passengers within five roundings of Keras's chances — 68 of them to the bit — and the two passengers the sample serves within two, from either file. A network of every kind the reader builds answers within six.

ONNX

using DeepSharp.Import.Onnx;
using DeepSharp.Networks;

using var graph = File.OpenRead("titanic.onnx");
var saved = new OnnxFile(new BinaryCrossEntropy()).Read(graph);    // a graph names no loss, so it is handed one

Three writers were measured and are read: PyTorch's torch.onnx.export, by its default exporter and by the TorchScript one before it; Keras's own model.export(format="onnx"); and tf2onnx, which is how a TensorFlow SavedModel reaches ONNX — python -m tf2onnx.convert --saved-model. Read the graph from its file: PyTorch's default exporter keeps larger numbers in a file beside the graph, and they are read from beside the graph's own file.

The graph is read once and lowered onto the layers it is, never kept beside them: a Gemm, or a MatMul and the Add of its bias, is a dense layer; a Conv a convolution; a BatchNormalization a batch normalisation, with ONNX's momentum, which is Keras's; a LayerNormalization; a Flatten, or a Reshape that flattens each example; a Relu, a Tanh or a Sigmoid. A Cast into single-precision numbers is nothing here, and a last Sigmoid or Softmax the loss applies itself is lifted into it.

Layouts are what each node declares, never what the numbers look like: a weight matrix is turned round where its Gemm says it is written outputs by inputs, as PyTorch writes one; a kernel is written channels out, channels in, rows, columns, as ONNX defines it; and an image flattened channel by channel there is flattened place by place here, so the numbers along such a row are turned from the one order to the other. Images go into the network, and come out, with their channels last. A graph that takes them channels last, as tf2onnx's does, says so with the Transpose that moves them after the batch before its convolutions, and that and the one moving them back before its flatten are nothing here; a convolution's pads written out one side at a time are read as TensorFlow's 'same' when they are the border 'same' gives. Numbers of half precision or bfloat16 are widened exactly, numbers of double precision rounded to the nearest single one.

What is refused, at its node and every such node at once: a branch — a node taking another value than the one the node before it made, or two — pooling and any other operator, a stride of its own for each axis, a dilated window, channels in groups, a Cast into another type, and a Transpose other than those two. tf2onnx writes a batch normalisation as a Mul, and a flatten for batches of any length through a side graph of Shape, Gather, Slice and Concat; neither is a layer here, so a network with a batch normalisation does not reach here through tf2onnx, and one that flattens images does when its SavedModel is exported for a batch of a fixed length.

Measured. The Titanic network PyTorch trained, exported both ways, answers the README's passenger within one rounding of PyTorch's chance and all 135 test passengers within five; the one Keras trained, exported by Keras or converted by tf2onnx 1.17.0 from TensorFlow 2.21.0's SavedModel, answers the 135 within five of Keras's chances and puts the very bits its .keras archive does into the slots. A network of convolutions, batch normalisations and a flatten gives each image PyTorch's output within 4.5e−8, and one of 'same' and strided convolutions that tf2onnx converted gives each TensorFlow's within 2.4e−7.

Trained further, and kept as one file

An import is a network and its loss, and records no fit of a pipeline — its TrainedOn is nothing — so it stands behind a pipeline as any network does: compiled with an optimizer and its loss, and fitted here on the rows the pipeline prepares, for one epoch or for as many as training it further takes. That fit is what the one file records, and the file is read back beside that fit alone:

using DeepSharp.Import.Keras;
using DeepSharp.Learners.Networks;
using DeepSharp.Networks;
using DeepSharp.Pipelines;

var prepared = Pdd.Create()
    .ReadCsv("titanic.csv")
    .Declare(schema => schema.Integer("survived", "sibsp", "parch").Category("pclass", "sex").Optional("age", ColumnKind.Number).Number("fare"))
    .SplitStratified("survived", train: 0.70, validation: 0.15)
    .FillMissing("age", With.Median)
    .EncodeCategories()
    .Normalise("age", Scale.MidRange)
    .Normalise("fare", Scale.MidRange)
    .Normalise("sibsp", Scale.MidRange)
    .Normalise("parch", Scale.MidRange)
    .Target("survived")
    .Build()
    .Run();

using var file = File.OpenRead("titanic.keras");
var saved = new KerasFile().Read(file);

var trained = saved.Network.Compile(new Adam(0.001), saved.Loss).Fit(prepared, new FitOptions(seed: 7) { Epochs = 1 });
File.WriteAllText("titanic.network.json", trained.ToJson());

The networks sample's Titanic network Keras trained, fitted one epoch further behind this pipeline, was read back from its file holding the same numbers to the bit and giving the passengers the same chances. Served without a fit, an import answers rows handed to it as a tensor of features, saved.Network.Predict(features, saved.Loss, engine), on any engine — the TorchSharp backend among them.

An importer of your own

A reader of another framework's files implements IImporter: Read(stream) hands back a SavedNetwork, and the numbers go in through network.Load, each as a SlotEntry — the path of its slot, the tensor laid out as the slot keeps it, and where its file holds it. Turning a number into the layout its slot keeps is the reader's, before it hands it over, and so is building the network when its file describes one; refusing what does not fit is the load's, in the same words for every importer.

Clone this wiki locally