my approach to learning is simple: it's top-down. i like tearing things apart, and building them back up from the ground.
here's the leaderboard i topped while optimizing a vector addition kernel in CUDA at: Tensara
worklog coming soon btw.
oh, and a blog (still a WIP) where i documented my journey writing my first ever benchmarked kernel:
mat-mul
more kernels and updates from my side soon! stay tuned :)
