Final Project for CSDS 451 Description Implementation of direct convolution for VGG16 model for performance boost. Run main.py to get batches (install pytorch first) python main.py Run makefile to compile c++ script make Use make run to run the code make run modify the block size to tune the direct convolution algorithm performance Performance Intel(R) UHD Graphics Parallel Output Channel Parallel Output Channel and Output Width 12th Gen Intel(R) Core(TM) i7-12700H Parallel Output Channel Parallel Output Channel and Output Width