# Convolutional Neural Network (CNN) with PyTorch (on MNIST)
By [Zahra Taheri](https://github.com/zata213), September 6, 2020

## Transition from Feedforward neural networks to Convolutional neural networks

### 1 Hidden Layer Feedforward Neural Network

![alt text](feedforward-nn.png)

### Basic Convolutional neural network

- Additional convolution and pooling layers before Feedforward neural network.
- Layer with a linear function and a non-linearity name: Fully connected layer.

![alt text](convolutional-nn.png)

## One Convolutional layer: High level view

### One Convolutional Layer, Input Depth of 1


- Input_depth=1 means the input is a 'one single color' image, i.e., gray scale image.
- It can be imagined that the touch light shining the image is called a 'filter' or a 'kernel', and where it is shining is called a 'receptive field' or a 'patch'.
- a kernel slides through the hole image to detect the receptive fields' shapes, and the kernel returns $0$ if it cannot detect the shape of a receptive field. Also, if the kernel across and detects the shape that the kernel knows how to detect, then you get a non-zero number like. 
- A convolution is just a mapping of the input to a bunch of numbers in the output. Such a bunch of numbers is called a 'feature map' or an 'activation map'. 
- The number of feature maps depend on the number of kernels (output depths=the number of kernels).
- More kernels results in more feature map, and so more kernels gives us more information about the input.
- The width of input and output may be different.
![alt text](one-conv-layer-depth1.png)
![alt text](one-conv-layer-ex.png)
![alt text](one-conv-layer-ex2.png)

### One Convolutional Layer, Input Depth of 3

- Input_depth=3 means the input is an 'RGB' colored image.
- The kernel must have the same depth, 3.

![alt text](one-conv-layer-depth3.png)
![alt text](one-conv-layer-depth3-ex.png)

### Summary
- As the kernel is sliding/convolving across the image, 2 operations done per match
    1. Element-wise multiplication
    2. Summation

- More kernels means more feature map channels, and so can capture more information about the input.

## Multiple Convolutional Layers

### Pooling Layers
The essence of the pooling is to downsampling the images, i.e., reducing the size of your input.

**2 common types of pooling layers:**
   - Max pooling
   - Average pooling
   
![alt text](pooling.png)
![alt text](pooling-ex.png)
![alt text](pooling-ex2.png)
![alt text](pooling-ex3.png)

#### Multiple pooling layers

![alt text](multiple-pooling.png)

### Padding
- **Valid padding (zero padding)**
    - output size < input size
- **Same padding (zero padding)**
    - output size = input size

![alt text](zero-padding.png)
![alt text](same-padding.png)

### Output size calculation

$O = \frac{W-K+2P}{S}+1$, where
    - O = Output height
    - W = Input height
    - K = 

![alt text](1.png)

In [None]:
import torch
import torch.nn as nn
import torchvision.transforms as transforms
import torchvision.datasets as dsets
from torch.autograd import Variable

'''
STEP 1: LOADING DATASET
'''

train_dataset = dsets.MNIST(root='./data', 
                            train=True, 
                            transform=transforms.ToTensor(),
                            download=True)

test_dataset = dsets.MNIST(root='./data', 
                           train=False, 
                           transform=transforms.ToTensor())

'''
STEP 2: MAKING DATASET ITERABLE
'''

batch_size = 100
n_iters = 3000
num_epochs = n_iters / (len(train_dataset) / batch_size)
num_epochs = int(num_epochs)

train_loader = torch.utils.data.DataLoader(dataset=train_dataset, 
                                           batch_size=batch_size, 
                                           shuffle=True)

test_loader = torch.utils.data.DataLoader(dataset=test_dataset, 
                                          batch_size=batch_size, 
                                          shuffle=False)

'''
STEP 3: CREATE MODEL CLASS
'''
class CNNModel(nn.Module):
    def __init__(self):
        super(CNNModel, self).__init__()
        
        # Convolution 1
        self.cnn1 = nn.Conv2d(in_channels=1, out_channels=16, kernel_size=5, stride=1, padding=0)
        self.relu1 = nn.ReLU()
        
        # Max pool 1
        self.maxpool1 = nn.MaxPool2d(kernel_size=2)
     
        # Convolution 2
        self.cnn2 = nn.Conv2d(in_channels=16, out_channels=32, kernel_size=5, stride=1, padding=0)
        self.relu2 = nn.ReLU()
        
        # Max pool 2
        self.maxpool2 = nn.MaxPool2d(kernel_size=2)
        
        # Fully connected 1 (readout)
        self.fc1 = nn.Linear(32 * 4 * 4, 10) 
    
    def forward(self, x):
        # Convolution 1
        out = self.cnn1(x)
        out = self.relu1(out)
        
        # Max pool 1
        out = self.maxpool1(out)
        
        # Convolution 2 
        out = self.cnn2(out)
        out = self.relu2(out)
        
        # Max pool 2 
        out = self.maxpool2(out)
        
        # Resize
        # Original size: (100, 32, 7, 7)
        # out.size(0): 100
        # New out size: (100, 32*7*7)
        out = out.view(out.size(0), -1)

        # Linear function (readout)
        out = self.fc1(out)
        
        return out

'''
STEP 4: INSTANTIATE MODEL CLASS
'''

model = CNNModel()

#######################
#  USE GPU FOR MODEL  #
#######################

if torch.cuda.is_available():
    model.cuda()

'''
STEP 5: INSTANTIATE LOSS CLASS
'''
criterion = nn.CrossEntropyLoss()


'''
STEP 6: INSTANTIATE OPTIMIZER CLASS
'''
learning_rate = 0.01

optimizer = torch.optim.SGD(model.parameters(), lr=learning_rate)

'''
STEP 7: TRAIN THE MODEL
'''
iter = 0
for epoch in range(num_epochs):
    for i, (images, labels) in enumerate(train_loader):
        
        #######################
        #  USE GPU FOR MODEL  #
        #######################
        if torch.cuda.is_available():
            images = Variable(images.cuda())
            labels = Variable(labels.cuda())
        else:
            images = Variable(images)
            labels = Variable(labels)
        
        # Clear gradients w.r.t. parameters
        optimizer.zero_grad()
        
        # Forward pass to get output/logits
        outputs = model(images)
        
        # Calculate Loss: softmax --> cross entropy loss
        loss = criterion(outputs, labels)
        
        # Getting gradients w.r.t. parameters
        loss.backward()
        
        # Updating parameters
        optimizer.step()
        
        iter += 1
        
        if iter % 500 == 0:
            # Calculate Accuracy         
            correct = 0
            total = 0
            # Iterate through test dataset
            for images, labels in test_loader:
                #######################
                #  USE GPU FOR MODEL  #
                #######################
                if torch.cuda.is_available():
                    images = Variable(images.cuda())
                else:
                    images = Variable(images)
                
                # Forward pass only to get logits/output
                outputs = model(images)
                
                # Get predictions from the maximum value
                _, predicted = torch.max(outputs.data, 1)
                
                # Total number of labels
                total += labels.size(0)
                
                #######################
                #  USE GPU FOR MODEL  #
                #######################
                # Total correct predictions
                if torch.cuda.is_available():
                    correct += (predicted.cpu() == labels.cpu()).sum()
                else:
                    correct += (predicted == labels).sum()
            
            accuracy = 100 * correct // total
            
            # Print Loss
            print('Iteration: {}. Loss: {}. Accuracy: {}'.format(iter, loss.data[0], accuracy))