Skip to content

TorchSharp backend

H.P. Gansevoort edited this page Oct 1, 2026 · 1 revision

TorchSharp backend

DeepSharp.Backends.TorchSharp runs a network's arithmetic on libtorch, the engine under PyTorch, through TorchSharp: on the processor, or on an NVIDIA graphics card. The model does not change a line. Every layer, loss, optimizer and gradient is the one DeepSharp runs on its light engine, held to the same contract operation by operation, and the engine is handed in where the arithmetic is asked for — it is a choice, never a dependency of the model.

dotnet add package DeepSharp.Backends.TorchSharp
dotnet add package libtorch-cpu-win-x64 --version 2.10.0

Your application brings libtorch

The package brings TorchSharp 0.107.0 and nothing native. A package that brought libtorch would bring every platform's to every application, and a card's gigabytes to one that only ever runs on the processor, so the application brings the one it runs on:

libtorch-cpu-win-x64, libtorch-cpu-linux-x64 or libtorch-cpu-osx-arm64, 2.10.0 The processor's, for one platform.
TorchSharp-cpu 0.107.0 The processor's, for all three.
TorchSharp-cuda-windows or TorchSharp-cuda-linux 0.107.0 An NVIDIA card's, in place of the processor's.

The processor's libtorch for Windows is 79.9 MB to download and 267.9 MB installed, for Linux 128.2 MB and 498.0 MB, and for macOS on Apple's processors 56.9 MB and 245.1 MB; a build for a card on Windows came to 4.2 GB. Made where there is no libtorch, the engine is refused with the names of the packages that bring one, and on a card where libtorch finds none — the libtorch brought is the processor's, or the machine has no NVIDIA card with its driver — it says so, naming the packages that bring a card's.

Choosing it

TorchBackend.OnCpu() is an engine on the processor and TorchBackend.OnGpu(0) one on the first graphics card: the device is the engine's, given when it is made. It is handed in where the arithmetic is asked for, as any engine is:

  • To a run, new FitOptions(seed: 42) { Backend = TorchBackend.OnCpu() }, which trains the network on it, judges it by the validation rows on it and takes the pipeline's report's measures on it, since the report is part of the run. Options hold the light engine unless they name another.
  • To a trained network as it serves, trained.Predict(rows, TorchBackend.OnGpu(0)), for that call alone; trained.Predict(rows) serves on a light engine of its own.
  • To a network, network.Predict(features, loss, engine), the one evaluation every door answers rows through.

Nothing keeps the engine. The network holds none and its file names none, so a network trained on libtorch is read back and served on the light engine, and one trained on the light engine is served on libtorch; a trained network's file is the same whichever engine did its arithmetic. The report and a served network ask an engine for a chunk of rows at a time — the run's own batch size for the report, thirty-two, Keras's own default, for rows served — and no answer depends on which rows share its pass.

What it is held to

An engine is held to the light engine operation by operation, and the tests that hold it are written once and run on every engine: the light one, an engine of the tests' own whose every value lives in memory allocated outside .NET and whose totals are single precision, and this one, on the processor and on a card.

  • A total — a product's inner sum, a column's sum, a mean, what folding adds onto one place — lies within three roundings of the size of the terms it adds up: three times a float's unit roundoff, 2⁻²⁴, times the sum of the terms' sizes, since two engines add the same terms in their own order. On libtorch on the processor a product came to 2.24 roundings at most, a column's sum to 0.28, a mean to 0.094 and a fold to 1.21, or 0.97 padded as 'same'.
  • A value worked out value by value lies within PyTorch's own float tolerance, a relative 1.3e−6 and an absolute 1e−5, and what only moves values moves them exactly.
  • One Titanic training step as the loop takes it is held the same way: its outputs, its loss and its gradients within three roundings, its parameters after Adam within the float tolerance.

On an RTX 5070 Ti, with libtorch built for CUDA 12.8, every contract held as well.

The promise is per step, not per run. Two engines that add up in another order drift apart over many steps, as any two float engines do. The networks sample's Titanic run went the same way on libtorch's processor build as on the light engine — forty-six epochs, the thirty-sixth kept, 0.815 of the test passengers right — and ended with its weights up to 0.08 apart, giving the sample's first passenger, a man of twenty-two in third class, 0.1655 where the light engine gives 0.1638. On the card it went the same way again and ended 2.4e−7 from the light engine's weights; served on the card and on the light engine, the network it trained answered the 135 test passengers within 1.2e−7 of itself. On one engine a run is the same run again, bit for bit.

The engine never sets how many threads libtorch works with, which is the application's to say for the whole process, and never draws from libtorch's random generator: every draw a run makes is DeepSharp's own, counted from its seed. Two hundred Titanic steps came to the same bits at 1, 4, 16 and 32 threads.

Gradients

Gradients are worked out on libtorch as on every engine: by the recording DeepSharp writes as a pass runs, from one rulebook of how each operation sends a gradient back, each rule checked by nudging its inputs and watching the loss. libtorch keeps a history of its own, but only for tensors marked, before the pass, as wanting a gradient, and only while a mode is switched on for the thread that runs it — a mark on a tensor the network, its best epoch and its checkpoints share, and a switch that is a shared instance by another name — so it is never switched on. The way back is worked out only towards the tensors asked for, as libtorch's own records only what wants a gradient: a training step asks for no gradient of its own rows. A tensor libtorch made stays where libtorch keeps it between operations, and an engine that keeps its tensors so takes two in a Titanic step — the batch's rows and their answers — and copies one out, the loss, where an engine that copied every tensor in and out of every operation copied 256, as measured on the tests' own engine on memory outside .NET.

Memory

What the engine makes stays in libtorch's memory, on the processor or the card, and is let go of once nothing holds it. TorchSharp ends every tensor with whatever dispose scope was open on the thread that made it; the engine takes each tensor out of any such scope the moment it exists, so a scope of yours around a run ends without touching the network's slots, what its optimizer remembers, the best epoch or a checkpoint you keep. Nothing but the collector ends one, and since the collector cannot see libtorch's memory, each tells it how many bytes it holds: of five hundred tensors of four megabytes made and dropped one after another, at most eleven were alive at once, where told nothing all five hundred were. Nothing is let go of sooner, at the end of a step, because what a step made may be held by code of yours — a layer written as code can keep a tensor of its pass.

What a run keeps is what something holds: the network's slots, what the optimizer remembers of each, the best epoch's slots when the run is to end holding them, and each checkpoint you keep, which holds its own epoch's slots, the optimizer's memory and the best epoch's slots as they stood when it was taken — on a card, in the card's memory. Two thousand steps of a small convolution on the card left its memory where the first thirty had: libtorch's context and cache took 283 MiB of it, and the most the run held besides was 12 MiB more, as much in the first thousand steps as in the second.

Checkpoints

A checkpoint records the engine its run was on: torch, the libtorch it runs on as TorchSharp states it — 2.10.0.0 for TorchSharp 0.107.0 — and its device, cpu or cuda:0. A run goes on from it only on an engine made the same way, on the same device with the same libtorch, and handed another it is refused, naming both, before anything is put back: gone on from on another engine, after its first, fifth, tenth or twentieth epoch alike, the networks sample's Titanic run kept the same epoch and ended with weights up to 3.0e−7 from the run it was taken of — another run, under the same seed, with nothing to tell the two apart. An engine of your own names its version and its device by implementing INamesItsVersionAndDevice; one that does not is recorded by its name alone.

When it is the faster

It depends on the size of the work. Measured side by side on one machine, an AMD Ryzen 9 9950X3D, in a harness that hands every step fresh copies of its rows:

light engine libtorch, processor RTX 5070 Ti
a Titanic training step: 32 rows of 14 features, a dense layer of 16 38.8 µs 336 µs on one thread, 365 on sixteen 1.2 ms
a convolution's step: 32 images of 28 by 28, eight filters of 3 by 3 12.1 ms 3.5 ms on one thread, 2.2 on sixteen 1.8 ms
the same over 256 images with 32 filters 345 ms 27.5 ms 5.0 ms

At Titanic's size the light engine is the faster: libtorch spends longer handing each small operation over than the operation takes, and a card longer still, each operation handed across to it and the loss read back. At a convolution's size libtorch is, and a card pays at a larger size still. These are one machine's numbers, not a promise about another's.

How it is checked

The engine's tests are a suite of their own, Tst/DeepSharp.Backends.TorchSharp, because they carry libtorch, which the library's own suite never does. They run every engine contract on libtorch on Windows and on Linux, on .NET 8 and .NET 10; built with -p:Libtorch=cuda, they bring a card's libtorch instead, and the tests that need a card run where there is one. And before a package leaves a run of the workflows, an application that references the engine's package just made, and nothing of DeepSharp's source, is started twice: brought no libtorch, it is refused, naming every package that brings one; brought the processor's, it fits a step of the networks sample's Titanic pipeline.

dotnet pack DeepSharp.slnx -c Release -o nupkgs
bash tools/torch/check.sh nupkgs

What it is not

  • Hand-written GPU code. The card is reached through libtorch, which already reaches it; no CUDA is written here.
  • A word on the pipeline. The pipeline is the recipe for the data and knows no learner, so there is no .UseTorch() on it: the engine is chosen where the arithmetic is asked for.
  • Required. The core carries the light engine, which needs nothing installed and travels inside an application; this is a package of its own, and a project that does not want libtorch never carries it.

Clone this wiki locally