Repository navigation
EveryO Architecture — Thoughts on the Next Stage of the Framework #6
Replies: 1 comment
|
Thank you for taking the time to explore EveryO and understand the thinking behind its architecture. You have highlighted exactly the balance I want the project to maintain: it should remain understandable enough to learn from, while gradually becoming capable of handling more realistic workloads. For autograd, the next direction is to strengthen the existing computation graph rather than creating special cases for transformers. Attention will be built from smaller differentiable operations such as matrix multiplication, softmax, reshaping, transposition, masking, and normalization. Once those operations have correct forward and backward implementations, transformer blocks can be composed from them naturally. Gradient correctness tests will remain important as the graph becomes more complex. For CUDA, the long-term goal is definitely to support GPU-resident tensors. The current CUDA functionality is an early acceleration layer, but repeatedly transferring data between the CPU and GPU would limit performance. Eventually, a tensor should own device-specific storage and remain on the selected device until the user explicitly moves it. I also plan to introduce a clearer operator and backend abstraction. The public tensor API should remain consistent, while NumPy and CUDA provide separate implementations underneath it. Device selection, memory management, and operation dispatch should happen below the user-facing API without duplicating the higher-level neural-network code. Maintaining readability is one of the project’s main principles. New functionality will be added in small, testable layers rather than through large abstractions. Features will be clearly separated as stable, experimental, or planned, and the documentation should explain both how something works and why it was designed that way. The likely architectural direction is:
I do not want EveryO to become a large framework that is difficult to understand. The goal is to grow it carefully, with correctness and clarity coming before feature count or performance claims. Thank you again for the thoughtful questions. Feedback like this is genuinely useful while deciding what the next architectural milestone should be. |
Uh oh!
There was an error while loading. Please reload this page.
Hi Krishanth,
I went through the EveryO repository and spent some time looking at the architecture. I really like the idea of building the framework from the ground up instead of hiding everything behind existing deep learning libraries.
The separation between tensors, autograd, neural network layers, optimizers, training, and the CUDA backend makes the project quite easy to understand.
I was curious about how you see the architecture evolving as the project grows.
For example, when adding things like attention, transformer blocks, GPU-resident tensors, and mixed-precision training, I think the current design decisions will become even more important.
A few things I was wondering about:
I think that balance between keeping EveryO simple enough to study while making it capable enough for more realistic workloads is one of the most interesting parts of the project.
Would be interested to hear your thoughts on the direction you want to take the architecture next.
All reactions