This release brings an implementation of the Vision Transformer (ViT) model tailored specifically for the MNIST dataset, providing a powerful tool for image classification tasks.
Key Features:
- Vision Transformer Model: Leverage the state-of-the-art ViT architecture for image recognition tasks, now applied to the MNIST dataset.
- Modular Codebase: Enjoy a clean and modular code structure, making it easy to understand, extend, and integrate with other projects.
- Efficient Training Script: Train the ViT model on MNIST with ease using our customizable training script, offering flexibility in experimentation.
- Seamless Integration: Built using PyTorch and torchvision, ensuring compatibility with existing deep learning workflows and tools.
Get Started:
- Clone the Repository:
git clone https://github.com/your-username/vision-transformer-mnist.git - Install Dependencies:
pip install -r requirements.txt - Run Training Script:
python train.py
Dataset:
MNIST, a classic benchmark dataset in the machine learning community, comprises 28x28 grayscale images of handwritten digits (0-9). This dataset is widely used for testing and benchmarking image classification algorithms.