Skip to content

Releases: CanReader/FastNN

Release list

FastNN 0.3.0

Choose a tag to compare

@CanReader CanReader released this 08 Aug 10:37

FastNN 0.3.0

The big change: CUDA is now opt-in. The default build is CPU only and needs no
toolkit, so cargo add fastnn works on any machine. Enable the GPU path with
--features cuda. If nvcc rejects your system compiler, point FASTNN_NVCC_CCBIN
at one it accepts.

Generation and transformers:

  • KV cached incremental decoding. Generating a token now costs O(n) attention
    instead of O(n^2), and a parity test proves the cached path matches the full
    forward exactly.
  • Sampler with temperature, top-k, top-p and repetition penalty. Beam search
    with length penalty and eos handling, usable with any step function.
  • Full encoder-decoder Transformer: decoder blocks with cross attention, plus
    key padding masks in both the encoder and the decoder.

Convolution:

  • im2col and col2im run on the GPU, so conv is on-device end to end.
  • Dilation, grouped and depthwise conv, Conv1d with a causal mode, and
    ConvTranspose1d/2d built on col2im as a first-class op. The transposed conv
    is tested against the adjoint identity, not just shapes.

Losses and optimizers:

  • New losses: huber, kl_divergence, nll, focal, dice, gaussian_nll,
    poisson_nll, cosine_embedding, triplet_margin, margin_ranking, contrastive,
    info_nce. A CrossEntropyLoss builder adds label smoothing, class weights,
    ignore_index and reduction modes.
  • New optimizers: RMSprop, Adagrad, Adadelta, RAdam, Lion, Lookahead and an
    EMA helper. All carry savable state and each update rule is unit tested
    against hand computed values.

Interchange:

  • safetensors save and load. Reads F32, F16 and BF16 (widened to f32 with
    exact round-to-nearest-even conversions), supports key remapping for foreign
    checkpoints, and reports missing and unexpected tensors instead of guessing.

Fixes worth knowing about:

  • try_matmul no longer underflows on 1-D operands.
  • A plain matrix times a batched tensor keeps its batch dimension.
  • CPU matmul is parallel per output row instead of per batch, which makes
    CPU training several times faster on transformer shapes.

Tests grew from 90 to about 160 on the CPU plus 16 GPU parity cases, spread
over gradient checks, convergence runs and hand computed reference values.