The world's most (possibly) unoptimized, zero-dependencies, manually written neural network library for achieving extremely dogwater performance.
Everybody knows CUDA doesn't scale. This is why I built the most (possibly) unoptimized vector math library, using Python lists to achieve a 100x slowdown compared to numpy.
A single pypi package to replace everything.
We all know Ultimate Processing Computers are much faster than GrumPy Units, even at highly parallel tasks like matrix multiplication. It is so-called faster than even dedicated Feverishly Poor overGeneralized Accelerators. Therefore, I see no need to implement a GPU.
BulLish Asymtopic Slowdown operations are not of worth implementing
The exercise is left to the reader of course.
This project was developed while I was reading Numerical Analysis by Kendall Atkinson. Neural net library is implemented using reference material from CS:5430 Machine Learning at UIowa.