VIT-from-scratch The Vision Transformer, or ViT, is a model for image classification that employs a Transformer-like architecture over patches of the image.