Skip to content

Release v0.104.1

Choose a tag to compare

@github-actions github-actions released this 26 Jun 14:36
3652dc7

Changes in v0.104.1

Performance Improvements

  • *Into resident-scratch engine kernels (SDPA/AdaLN/FusedLinear) for diffusion inference (#696)
  • managed FP32 GEMM — GotoBLAS rewrite + CCX-2D-NUMA + short-M campaign (DiT ~2x; sq4096 > throttled MKL) (#695)

Packages

  • AiDotNet.Native.CLBlast
  • AiDotNet.Native.CUDA
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.CUDA.Fragments
  • AiDotNet.Native.MoltenVK
  • AiDotNet.Native.OneDNN
  • AiDotNet.Native.OpenBLAS
  • AiDotNet.Native.ROCm
  • AiDotNet.Tensors

Note: AiDotNet.Native.CUDA bundles the NVIDIA CUDA 12 runtime (cudart / cuBLAS / cuBLASLt / nvRTC, ~770MB). Because that exceeds nuget.org's 250 MB per-package limit, it ships as a small meta package plus a set of AiDotNet.Native.CUDA.Fragments.* chunk packages that are reassembled into the full DLLs at your build (win-x64) — just dotnet add package AiDotNet.Native.CUDA alongside AiDotNet.Tensors on NVIDIA GPU machines; the fragments restore automatically.
Note: AiDotNet.Native.ROCm requires ROCm installed on the target machine for the kernel library (~496MB).

Installation

dotnet add package AiDotNet.Tensors