Repository navigation
Releases: CanReader/FastNN
Releases · CanReader/FastNN
Release list
FastNN 0.3.0
FastNN 0.3.0
The big change: CUDA is now opt-in. The default build is CPU only and needs no
toolkit, so cargo add fastnn works on any machine. Enable the GPU path with
--features cuda. If nvcc rejects your system compiler, point FASTNN_NVCC_CCBIN
at one it accepts.
Generation and transformers:
- KV cached incremental decoding. Generating a token now costs O(n) attention
instead of O(n^2), and a parity test proves the cached path matches the full
forward exactly. - Sampler with temperature, top-k, top-p and repetition penalty. Beam search
with length penalty and eos handling, usable with any step function. - Full encoder-decoder Transformer: decoder blocks with cross attention, plus
key padding masks in both the encoder and the decoder.
Convolution:
- im2col and col2im run on the GPU, so conv is on-device end to end.
- Dilation, grouped and depthwise conv, Conv1d with a causal mode, and
ConvTranspose1d/2d built on col2im as a first-class op. The transposed conv
is tested against the adjoint identity, not just shapes.
Losses and optimizers:
- New losses: huber, kl_divergence, nll, focal, dice, gaussian_nll,
poisson_nll, cosine_embedding, triplet_margin, margin_ranking, contrastive,
info_nce. A CrossEntropyLoss builder adds label smoothing, class weights,
ignore_index and reduction modes. - New optimizers: RMSprop, Adagrad, Adadelta, RAdam, Lion, Lookahead and an
EMA helper. All carry savable state and each update rule is unit tested
against hand computed values.
Interchange:
- safetensors save and load. Reads F32, F16 and BF16 (widened to f32 with
exact round-to-nearest-even conversions), supports key remapping for foreign
checkpoints, and reports missing and unexpected tensors instead of guessing.
Fixes worth knowing about:
- try_matmul no longer underflows on 1-D operands.
- A plain matrix times a batched tensor keeps its batch dimension.
- CPU matmul is parallel per output row instead of per batch, which makes
CPU training several times faster on transformer shapes.
Tests grew from 90 to about 160 on the CPU plus 16 GPU parity cases, spread
over gradient checks, convergence runs and hand computed reference values.