# Naive Machine Translation and LSH


### 1. The Word Embeddings Data for English and French Words
        
- **The Data**
    - Look at the data
    - Load two dictionaries mapping the English to French words
    - Looking at the English French dictionary

- **1.1 Generate Embedding and Transform Matrices**

### 2. Translations

#### 2.1 - Translation as Linear Transformation of Embeddings

- **Describing translation as the minimization problem**

Find a matrix `R` that minimizes the following equation. 

$$\arg \min _{\mathbf{R}}\| \mathbf{X R} - \mathbf{Y}\|_{F}\tag{1} $$

- **Frobenius norm**

The Frobenius norm of a matrix $A$ (assuming it is of dimension $m,n$) is defined as the square root of the sum of the absolute squares of its elements:

$$\|\mathbf{A}\|_{F} \equiv \sqrt{\sum_{i=1}^{m} \sum_{j=1}^{n}\left|a_{i j}\right|^{2}}\tag{2}$$
    
- Actual loss function
In the real world applications, the Frobenius norm loss:

$$\| \mathbf{XR} - \mathbf{Y}\|_{F}$$

is often replaced by it's squared value divided by $m$:

$$ \frac{1}{m} \|  \mathbf{X R} - \mathbf{Y} \|_{F}^{2}$$

where $m$ is the number of examples (rows in $\mathbf{X}$).

* The same R is found when using this loss function versus the original Frobenius norm.
* The reason for taking the square is that it's easier to compute the gradient of the squared Frobenius.
* The reason for dividing by $m$ is that we're more interested in the average loss per embedding than the  loss for the entire training set.
    * The loss for all training set increases with more words (training examples),
    so taking the average helps us to track the average loss regardless of the size of the training set.
- compute_loss
    - Computing the loss
    - compute_gradient
        - Computing the gradient of loss with respect to transform matrix R
        - Finding the optimal R with Gradient Descent Algorithm
    - align_embeddings

#### 2.2 - Testing the Translation

- k-Nearest Neighbors Algorithm

[k-Nearest neighbors algorithm](https://en.wikipedia.org/wiki/K-nearest_neighbors_algorithm) 
* k-NN is a method which takes a vector as input and finds the other vectors in the dataset that are closest to it. 
* The 'k' is the number of "nearest neighbors" to find (e.g. k=2 finds the closest two neighbors).

- Searching for the Translation Embedding
Since we're approximating the translation function from English to French embeddings by a linear transformation matrix $\mathbf{R}$, most of the time we won't get the exact embedding of a French word when we transform embedding $\mathbf{e}$ of some particular English word into the French embedding space. 
* This is where $k$-NN becomes really useful! By using $1$-NN with $\mathbf{eR}$ as input, we can search for an embedding $\mathbf{f}$ (as a row) in the matrix $\mathbf{Y}$ which is the closest to the transformed vector $\mathbf{eR}$

- Cosine Similarity
Cosine similarity between vectors $u$ and $v$ calculated as the cosine of the angle between them.
The formula is 

$$\cos(u,v)=\frac{u\cdot v}{\left\|u\right\|\left\|v\right\|}$$

* $\cos(u,v)$ = $1$ when $u$ and $v$ lie on the same line and have the same direction.
* $\cos(u,v)$ is $-1$ when they have exactly opposite directions.
* $\cos(u,v)$ is $0$ when the vectors are orthogonal (perpendicular) to each other.

- nearest_neighbor
- test_vocabulary
    
### 3. LSH and Document Search

#### 3.1 - Getting the Document Embeddings**
- get_document_embedding
- get_document_vecs
#### 3.2 - Looking up the Tweets
#### 3.3 - Finding the most Similar Tweets with LSH
#### 3.4 - Getting the Hash Number for a Vector
- hash_value_of_vector
#### 3.5 - Creating a Hash Table
- make_hash_table
#### 3.6 - Creating all Hash Tables
- approximate_knn
        

In [2]:
import pdb
import pickle
import string

import time

import nltk
import numpy as np
from nltk.corpus import stopwords, twitter_samples

from os import getcwd

# add folder, tmp2, from our local workspace containing pre-downloaded corpora files to nltk's data path
filePath = f"{getcwd()}/tmp2/"
nltk.data.path.append(filePath)

## Function

### 1. English to french dictionary

In [3]:
def get_dict(file_name):
    """
    This function returns the english to french dictionary given a file where the each column corresponds to a word.
    Check out the files this function takes in your workspace.
    """
    my_file = pd.read_csv(file_name, delimiter=' ')
    etof = {}  # the english to french dictionary to be returned
    for i in range(len(my_file)):
        # indexing into the rows.
        en = my_file.loc[i][0]
        fr = my_file.loc[i][1]
        etof[en] = fr

    return etof

### 2. Process tweet

In [4]:
def process_tweet(tweet):
    '''
    Input:
        tweet: a string containing a tweet
    Output:
        tweets_clean: a list of words containing the processed tweet

    '''
    stemmer = PorterStemmer()
    stopwords_english = stopwords.words('english')
    # remove stock market tickers like $GE
    tweet = re.sub(r'\$\w*', '', tweet)
    # remove old style retweet text "RT"
    tweet = re.sub(r'^RT[\s]+', '', tweet)
    # remove hyperlinks
    tweet = re.sub(r'https?:\/\/.*[\r\n]*', '', tweet)
    # remove hashtags
    # only removing the hash # sign from the word
    tweet = re.sub(r'#', '', tweet)
    # tokenize tweets
    tokenizer = TweetTokenizer(preserve_case=False, strip_handles=True,
                               reduce_len=True)
    tweet_tokens = tokenizer.tokenize(tweet)

    tweets_clean = []
    for word in tweet_tokens:
        if (word not in stopwords_english and  # remove stopwords
            word not in string.punctuation):  # remove punctuation
            # tweets_clean.append(word)
            stem_word = stemmer.stem(word)  # stemming word
            tweets_clean.append(stem_word)

    return tweets_clean

### 3. Cosine Similarity

In [5]:
def cosine_similarity(A, B):
    '''
    Input:
        A: a numpy array which corresponds to a word vector
        B: A numpy array which corresponds to a word vector
    Output:
        cos: numerical number representing the cosine similarity between A and B.
    '''
    # you have to set this variable to the true label.
    cos = -10    
    dot = np.dot(A, B)
    normb = np.linalg.norm(B)
    
    if len(A.shape) == 1: # If A is just a vector, we get the norm
        norma = np.linalg.norm(A)
        cos = dot / (norma * normb)
    else: # If A is a matrix, then compute the norms of the word vectors of the matrix (norm of each row)
        norma = np.linalg.norm(A, axis=1)
        epsilon = 1.0e-9 # to avoid division by 0
        cos = dot / (norma * normb + epsilon)
        
    return cos


### 4. Generate Embedding and Transform Matrices

Function :

```python
def get_matrices(en_fr, 
                 french_vecs, 
                 english_vecs):
```

Inputs:
- `en_fr` : English to French dictionary
- `en_embeddings` : English to embeddings dictionary
- `fr_embeddings` : French to embeddings dictionary

Return :
- Matrix `X` : each row is the word embedding for an english word
- Matrix `Y` : each same row (as X) is the word embedding for the french version word 

1. Lists of the english and french word embeddings

```python 
    X_l = list()
    Y_l = list()
```

2. Get the english words and rench words (the keys in the dictionary) and store in a set()

```python 
    english_set = english_vecs.keys()
    french_set = french_vecs.keys()
```

3. Store the french words that are part of the english-french dictionary (these are the values of the dictionary)

```python 
    french_words = set(en_fr.values())
```

4. Loop through all english, french word pairs in the english french dictionary

```python 
    for en_word, fr_word in en_fr.items():
```

5. Check that the french word has an embedding and that the english word has an embedding

```python 
        if fr_word in french_set and en_word in english_set:
```

6. Get the english embedding and french embedding

```python 
            en_vec = english_vecs[en_word]
            fr_vec = french_vecs[fr_word]
```

7. Add the english embedding and french embedding to their respectives lists

```python
            X_l.append(en_vec)
            Y_l.append(fr_vec)
```

8. Stack the vectors of X_l into a matrix X and the vectors of Y_l into a matrix Y

```python 
    X = np.vstack(X_l)
    Y = np.vstack(Y_l)
```

![image.png](attachment:image.png)

In [6]:
def get_matrices(en_fr, french_vecs, english_vecs):
    """
    Input:
        en_fr: English to French dictionary
        french_vecs: French words to their corresponding word embeddings.
        english_vecs: English words to their corresponding word embeddings.
    Output: 
        X: a matrix where the columns are the English embeddings.
        Y: a matrix where the columns correspong to the French embeddings.
        R: the projection matrix that minimizes the F norm ||X R -Y||^2.
    """
    # X_l and Y_l are lists of the english and french word embeddings
    X_l = list()
    Y_l = list()

    # get the english words (the keys in the dictionary) and store in a set()
    english_set = english_vecs.keys()

    # get the french words (keys in the dictionary) and store in a set()
    french_set = french_vecs.keys()

    # store the french words that are part of the english-french dictionary (these are the values of the dictionary)
    french_words = set(en_fr.values())

    # loop through all english, french word pairs in the english french dictionary
    for en_word, fr_word in en_fr.items():

        # check that the french word has an embedding and that the english word has an embedding
        if fr_word in french_set and en_word in english_set:

            # Get the english embedding and french embedding
            en_vec = english_vecs[en_word]
            fr_vec = french_vecs[fr_word]

            # Add the english embedding and french embedding to their respectives lists
            X_l.append(en_vec)
            Y_l.append(fr_vec)

    # stack the vectors of X_l into a matrix X
    X = np.vstack(X_l)

    # stack the vectors of Y_l into a matrix Y
    Y = np.vstack(Y_l)
    
    return X, Y

### 4. Translation as Linear Transformation of Embeddings - Computing the loss

* The loss function will be squared Frobenius norm of the difference between
matrix and its approximation, divided by the number of training examples $m$.
* Its formula is:
$$ L(X, Y, R)=\frac{1}{m}\sum_{i=1}^{m} \sum_{j=1}^{n}\left( a_{i j} \right)^{2}$$

where $a_{i j}$ is value in $i$th row and $j$th column of the matrix $\mathbf{XR}-\mathbf{Y}$.

**Function :**

```python 
def compute_loss(X, 
                 Y, 
                 R):
```

Input :
- X: a matrix of dimension `(m,n)` where the columns are the English embeddings.
- Y: a matrix of dimension `(m,n)` where the columns correspong to the French embeddings.
- R: a matrix of dimension `(n,n)` - transformation matrix from English to French vector space embeddings.

Return :
- L: a matrix of dimension (m,n) - the value of the loss function for given X, Y and R.

1. The number of rows in X
```python 
    m = X.shape[0]
```

2. Compute : XR - Y

```python 
    diff = np.dot(X,R)-Y
```

3. Compute the squarred of XR - Y

```python 
    diff_squared = diff**2
```

4. Compute sum of elements of this squarred matrix

```python 
    sum_diff_squared = np.sum(diff_squared)
```

5. Compute loss

```python 
    loss = sum_diff_squared/m
```

In [None]:
def compute_loss(X, Y, R):
    '''
    Inputs: 
        X: a matrix of dimension (m,n) where the columns are the English embeddings.
        Y: a matrix of dimension (m,n) where the columns correspong to the French embeddings.
        R: a matrix of dimension (n,n) - transformation matrix from English to French vector space embeddings.
    Outputs:
        L: a matrix of dimension (m,n) - the value of the loss function for given X, Y and R.
    '''
    # m is the number of rows in X
    m = X.shape[0]
    
    # diff is XR - Y
    diff = np.dot(X,R)-Y

    # diff_squared is the element-wise square of the difference
    diff_squared = diff**2

    # sum_diff_squared is the sum of the squared elements
    sum_diff_squared = np.sum(diff_squared)

    # loss i the sum_diff_squard divided by the number of examples (m)
    loss = sum_diff_squared/m

    return loss

### 5. Translation as Linear Transformation of Embeddings - Compute Gradient

* The formula for the gradient of the loss function $𝐿(𝑋,𝑌,𝑅)$ is:

$$\frac{d}{dR}𝐿(𝑋,𝑌,𝑅)=\frac{d}{dR}\Big(\frac{1}{m}\| X R -Y\|_{F}^{2}\Big) = \frac{2}{m}X^{T} (X R - Y)$$

**Function :**
```python 
def compute_gradient(X, 
                     Y, 
                     R):
```

**input :**
- X: a matrix of dimension (m,n) where the columns are the English embeddings.
- Y: a matrix of dimension (m,n) where the columns correspong to the French embeddings.
- R: a matrix of dimension (n,n) - transformation matrix from English to French vector space embeddings.

**Return :**

1. Compute the number of rows in X

```python 
    m = X.shape[0]
```

2. Compute gradient : X^T(XR - Y) * 2/m

```python 
    gradient = np.dot(X.transpose(),np.dot(X,R)-Y)*(2/m)
```

In [None]:
def compute_gradient(X, Y, R):
    '''
    Inputs: 
        X: a matrix of dimension (m,n) where the columns are the English embeddings.
        Y: a matrix of dimension (m,n) where the columns correspong to the French embeddings.
        R: a matrix of dimension (n,n) - transformation matrix from English to French vector space embeddings.
    Outputs:
        g: a scalar value - gradient of the loss function L for given X, Y and R.
    '''
    # m is the number of rows in X
    m = X.shape[0]

    # gradient is X^T(XR - Y) * 2/m
    gradient = np.dot(X.transpose(),np.dot(X,R)-Y)*(2/m)
    
    return gradient

### 6. Translation as Linear Transformation of Embeddings - Finding the optimal R with Gradient Descent Algorithm

Update $R$ with the formula:
$$R_{\text{new}}= R_{\text{old}}-\alpha g$$

**Function :**

```python 
def align_embeddings(X, 
                     Y, 
                     train_steps=100, 
                     learning_rate=0.0003, *
                     verbose=True, 
                     compute_loss=compute_loss, 
                     compute_gradient=compute_gradient):
```

**Input :**

- X: a matrix of dimension (m,n) where the columns are the English embeddings.
- Y: a matrix of dimension (m,n) where the columns correspong to the French embeddings.
- train_steps: positive int - describes how many steps will gradient descent algorithm do.
- learning_rate: positive float - describes how big steps will  gradient descent algorithm do.

**Return :**
- R: a matrix of dimension (n,n) - the projection matrix that minimizes the F norm ||X R -Y||^2

1. Compute a seed for random process

```python 
    np.random.seed(129)
```

2. R is a square matrix with length equal to the number of dimensions in the word embedding, the number of columns in X is the number of dimensions for a word vector (e.g. 300)

```python 
    R = np.random.rand(X.shape[1], X.shape[1])
```

3. Loop to train the model

```python 
    for i in range(train_steps):
```

4. Use the function that you defined to compute the gradient

```python 
        gradient = compute_gradient(X,Y,R)
```

5. Update R by subtracting the learning rate times gradient

```python 
        R -= learning_rate * gradient
```

In [None]:
def align_embeddings(X, Y, train_steps=100, learning_rate=0.0003, verbose=True, compute_loss=compute_loss, compute_gradient=compute_gradient):
    '''
    Inputs:
        X: a matrix of dimension (m,n) where the columns are the English embeddings.
        Y: a matrix of dimension (m,n) where the columns correspong to the French embeddings.
        train_steps: positive int - describes how many steps will gradient descent algorithm do.
        learning_rate: positive float - describes how big steps will  gradient descent algorithm do.
    Outputs:
        R: a matrix of dimension (n,n) - the projection matrix that minimizes the F norm ||X R -Y||^2
    '''
    np.random.seed(129)

    # the number of columns in X is the number of dimensions for a word vector (e.g. 300)
    # R is a square matrix with length equal to the number of dimensions in th  word embedding
    R = np.random.rand(X.shape[1], X.shape[1])

    for i in range(train_steps):
        if verbose and i % 25 == 0:
            print(f"loss at iteration {i} is: {compute_loss(X, Y, R):.4f}")
        # use the function that you defined to compute the gradient
        gradient = compute_gradient(X,Y,R)

        # update R by subtracting the learning rate times gradient
        R -= learning_rate * gradient
    return R

### 7. Testing the Translation - nearest_neighbor

* k-NN is a method which takes a vector as input and finds the other vectors in the dataset that are closest to it. 
* The 'k' is the number of "nearest neighbors" to find (e.g. k=2 finds the closest two neighbors).

**Function :**

```python 
def nearest_neighbor(v, 
                     candidates, 
                     k=1, 
                     cosine_similarity=cosine_similarity):
```

**Input :**

- v, the vector you are going find the nearest neighbor for
- candidates: a set of vectors where we will find the neighbors
- k: top k nearest neighbors to find

**Return :**
      
- k_idx: the indices of the top k closest vectors in sorted form


1. Create an empty list of similarity

```python 
    similarity_l = []
```

2. For each candidate vector get the cosine similarityand append it to the similarity list

```python 
    for row in candidates:
        cos_similarity = cosine_similarity(v,row)
        similarity_l.append(cos_similarity)
```

3. Sort the similarity list and get the indices of the sorted list  

```python 
    sorted_ids = np.argsort(similarity_l)[::-1]
```

4. Get the indices of the k most similar candidate vectors

```python 
    k_idx = sorted_ids[:k]
```

**Understand the sort**

```python
similarity_l = np.array([0.2, 0.9, 0.75, 0.4, 0.65])

sorted_ids = np.argsort(similarity_l)[::-1]

``` 

`sorted_ids = [1, 2, 4, 3, 0]`

```python
k_idx = sorted_ids[:k]
```

`k_idx = [1, 2, 4]`

In [7]:
def nearest_neighbor(v, candidates, k=1, cosine_similarity=cosine_similarity):
    """
    Input:
      - v, the vector you are going find the nearest neighbor for
      - candidates: a set of vectors where we will find the neighbors
      - k: top k nearest neighbors to find
    Output:
      - k_idx: the indices of the top k closest vectors in sorted form
    """
    similarity_l = []

    # for each candidate vector...
    for row in candidates:
        # get the cosine similarity
        cos_similarity = cosine_similarity(v,row)

        # append the similarity to the list
        similarity_l.append(cos_similarity)

    # sort the similarity list and get the indices of the sorted list    
    sorted_ids = np.argsort(similarity_l)[::-1]
    
    # get the indices of the k most similar candidate vectors
    k_idx = sorted_ids[:k]
    return k_idx

### 8. Testing the Translation - test_vocabulary

* Calculate accuracy as $$\text{accuracy}=\frac{\#(\text{correct predictions})}{\#(\text{total predictions})}$$

**Function :**

```python
def test_vocabulary(X, Y, R, nearest_neighbor=nearest_neighbor):
```

**Input :**

- X: a matrix where the columns are the English embeddings.
- Y: a matrix where the columns correspong to the French embeddings.
- R: the transform matrix which translates word embeddings from English to French word vector space.

**Return :**

- accuracy: for the English to French capitals

1. Compute prediction

```python 
    pred = np.dot(X,R)
```

2. Initialize the number correct to zero

```python 
    num_correct = 0
```

3. Loop through each row in pred (each transformed embedding)

```python 
    for i in range(len(pred)):
```

4. Get the index of the nearest neighbor of pred at row 'i'; also pass in the candidates in Y

```python 
        pred_idx = nearest_neighbor(pred[i],Y)
```

5. If the index of the nearest neighbor equals the row of i increment the number correct by 1.

```python 
        if pred_idx == i:
            num_correct += 1
```

6. Accuracy is the number correct divided by the number of rows in 'pred' (also number of rows in X)

```python 
    accuracy = num_correct / len(pred)
```

In [8]:
def test_vocabulary(X, Y, R, nearest_neighbor=nearest_neighbor):
    '''
    Input:
        X: a matrix where the columns are the English embeddings.
        Y: a matrix where the columns correspong to the French embeddings.
        R: the transform matrix which translates word embeddings from
        English to French word vector space.
    Output:
        accuracy: for the English to French capitals
    '''
    # The prediction is X times R
    pred = np.dot(X,R)

    # initialize the number correct to zero
    num_correct = 0

    # loop through each row in pred (each transformed embedding)
    for i in range(len(pred)):
        # get the index of the nearest neighbor of pred at row 'i'; also pass in the candidates in Y
        pred_idx = nearest_neighbor(pred[i],Y)

        # if the index of the nearest neighbor equals the row of i... \
        if pred_idx == i:
            # increment the number correct by 1.
            num_correct += 1

    # accuracy is the number correct divided by the number of rows in 'pred' (also number of rows in X)
    accuracy = num_correct / len(pred)

    return accuracy

### 9. Getting the Document Embeddings

python 

```python 

```

