<a href="https://colab.research.google.com/github/pavitrarajitha/FMML-LAB-ASSIGNMENT/blob/main/Mod3_lab3.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

# **FOUNDATIONS OF MODERN MACHINE LEARNING, IIIT Hyderabad**
### MODULE: CLASSIFICATION-1
### LAB-3 : Using KNN for Text Classification
#### Module Coordinator: Jashn Arora


---

## **Section 1: Understanding NLP tools**

In this lab we will be using KNN on a real world NLP application i.e. is text classification. But first look at some NLP techniques for text classification and tools that we use when we want to use python for NLP.

## Section 1.2: Data Cleaning and Preprocessing step

Raw text must be processed and converted into a form so that it is suitable to use with various machine-learning algorithms.  
In case of text, there are lots of things that need to be taken into account.  


1.   Removing numbers from the text
2.   Handling capitalization and punctuation.
3.   Stemming and Lemmatizing text.  

And most importantly, one can't just use words or images directly in algorithms; they need to be converted into vectors- a form that algorithms can understand.



### **NLTK**
NLTK (or Natural Language Tool Kit) is a commonly used library for processing text. We will use this tool in this lab. Lets first install it.


In [1]:
import nltk
nltk.download('stopwords')
nltk.download('punkt')
nltk.download('averaged_perceptron_tagger')
nltk.download('wordnet')
nltk.download('omw-1.4')

[nltk_data] Downloading package stopwords to /root/nltk_data...
[nltk_data]   Unzipping corpora/stopwords.zip.
[nltk_data] Downloading package punkt to /root/nltk_data...
[nltk_data]   Unzipping tokenizers/punkt.zip.
[nltk_data] Downloading package averaged_perceptron_tagger to
[nltk_data]     /root/nltk_data...
[nltk_data]   Unzipping taggers/averaged_perceptron_tagger.zip.
[nltk_data] Downloading package wordnet to /root/nltk_data...
[nltk_data] Downloading package omw-1.4 to /root/nltk_data...


True

In [2]:
import re
import numpy
from nltk.corpus import stopwords
from nltk.stem import WordNetLemmatizer
from nltk.stem import SnowballStemmer
from nltk.tokenize import word_tokenize
from nltk import pos_tag
from nltk.corpus import wordnet
from sklearn.feature_extraction.text import CountVectorizer, TfidfVectorizer
from bs4 import BeautifulSoup

def cleanText(text, lemmatize, stemmer):
    """Method for cleaning text from train and test data. Removes numbers, punctuation, and capitalization. Stems or lemmatizes text."""

    if isinstance(text, float):
        text = str(text)
    if isinstance(text, numpy.int64):
        text = str(text)
    try:
        text = text.decode()
    except AttributeError:
        pass

    soup = BeautifulSoup(text, "lxml")
    text = soup.get_text()
    text = re.sub(r"[^A-Za-z]", " ", text)
    text = text.lower()


    if lemmatize:
        wordnet_lemmatizer = WordNetLemmatizer()

        def get_tag(tag):
            if tag.startswith('J'):
                return wordnet.ADJ
            elif tag.startswith('V'):
                return wordnet.VERB
            elif tag.startswith('N'):
                return wordnet.NOUN
            elif tag.startswith('R'):
                return wordnet.ADV
            else:
                return ''

        text_result = []
        tokens = word_tokenize(text)  # Generate list of tokens
        tagged = pos_tag(tokens)
        for t in tagged:
            try:
                text_result.append(wordnet_lemmatizer.lemmatize(t[0], get_tag(t[1][:2])))
            except:
                text_result.append(wordnet_lemmatizer.lemmatize(t[0]))
        return text_result

    if stemmer:
        text_result = []
        tokens = word_tokenize(text)
        snowball_stemmer = SnowballStemmer('english')
        for t in tokens:
            text_result.append(snowball_stemmer.stem(t))
        return text_result

In [3]:
sample_text = "Troubling"
sample_text_result = cleanText(sample_text, lemmatize=False, stemmer=True)
sample_text_result = " ".join(str(x) for x in sample_text_result)
print(sample_text)
print(sample_text_result)
sample_text_result = cleanText(sample_text, lemmatize=True, stemmer=False)
sample_text_result = " ".join(str(x) for x in sample_text_result)
print(sample_text_result)

Troubling
troubl
trouble


## Section 1.2: BAG OF WORDS

A bag-of-words model, or BoW for short, is a way of extracting features from text for use in modeling, such as with machine learning algorithms.

The approach is very simple and flexible, and can be used in many ways for extracting features from documents.

A bag-of-words is a representation of text that describes the occurrence of words within a document.
It is called a “bag” of words, because any information about the order or structure of words in the document is discarded. The model is only concerned with whether known words occur in the document, not where in the document.

In [4]:
# Functions to convert document(s) to a list of words, with the option of removing stopwords. Returns document-term matrix.

def createBagOfWords(train, test, remove_stopwords, lemmatize, stemmer):
    if remove_stopwords:
        vectorizer = CountVectorizer(analyzer='word', input='content', stop_words=stopwords.words('english'))
    else:
        vectorizer = CountVectorizer(analyzer='word', input='content')

    clean_train = []
    for paragraph in train:
        paragraph_result = cleanText(paragraph, lemmatize, stemmer)
        paragraph = " ".join(str(x) for x in paragraph_result)
        clean_train.append(paragraph)

    clean_test = []
    for paragraph in test:
        paragraph_result = cleanText(paragraph, lemmatize, stemmer)
        paragraph = " ".join(str(x) for x in paragraph_result)
        clean_test.append(paragraph)

    bag_of_words_train = vectorizer.fit_transform(clean_train).toarray()
    bag_of_words_test = vectorizer.transform(clean_test).toarray()
    return bag_of_words_train, bag_of_words_test


## Section 1.3: TF-IDF
TF-IDF technique is used to find meaning of sentences consisting of words and cancels out the incapabilities of Bag of Words technique which is good for text classification or for helping a machine read words in numbers.

The number of times a term occurs in a document is called its Term frequency (TF).

 Document frequency is the number of documents in which the word is present.  Inverse DF (IDF) is the inverse of the document frequency which measures the informativeness of term *t*.




In [5]:
def createTFIDF(train, test, remove_stopwords, lemmatize, stemmer):
    if remove_stopwords:
        vectorizer = TfidfVectorizer(analyzer='word', input='content', stop_words=stopwords.words('english'))
    else:
        vectorizer =  TfidfVectorizer(analyzer='word', input='content')

    clean_train = []
    for paragraph in train:
        paragraph_result = cleanText(paragraph, lemmatize, stemmer)
        paragraph = " ".join(str(x) for x in paragraph_result)
        clean_train.append(paragraph)

    clean_test = []
    for paragraph in test:
        paragraph_result = cleanText(paragraph, lemmatize, stemmer)
        paragraph = " ".join(str(x) for x in paragraph_result)
        clean_test.append(paragraph)

    tfidf_train = vectorizer.fit_transform(clean_train).toarray()
    tfidf_test = vectorizer.transform(clean_test).toarray()
    return tfidf_train, tfidf_test

# **Section 2: UNDERSTANDING THE DATA : A REVIEWS DATASET**

Sentiment analysis is the interpretation and classification of emotions (such as positive, negative and neutral) within text data using text analysis techniques.  
Given below is a dataset consisting of reviews along with sentiment class (positive or negative).

In [6]:
# Upload the Reviews CSV file that has been shared with you.
# Run this cell, click on the 'Choose files' button and upload the file.
from google.colab import files
uploaded = files.upload()

Saving reviews.csv to reviews.csv


In [7]:
import pandas as pd
df = pd.read_csv('reviews.csv')

In [8]:
df = df.dropna()

In [9]:
df.to_csv('reviews.csv', index=False)

# **Section 3: KNN MODEL**

Given below are two KNN models; in the first case we are using Bag-of-Words and in the second case we are using TF-IDF.
Note the different metrics and parameters used in each.

In [10]:
from sklearn import metrics, neighbors
from sklearn.model_selection import train_test_split, cross_val_score, cross_val_predict

## TASK - 1: Tweak the models below and see results with different parameters and distance metrics.

def bow_knn():
    """Method for determining nearest neighbors using bag-of-words and K-Nearest Neighbor algorithm"""

    training_data = pd.read_csv('reviews.csv')
    X_train, X_test, y_train, y_test = train_test_split(training_data["sentence"], training_data["sentiment"], test_size=0.2, random_state=5)
    X_train, X_test = createBagOfWords(X_train, X_test, remove_stopwords=True, lemmatize=True, stemmer=False)
    # print(X_train)
    knn = neighbors.KNeighborsClassifier(n_neighbors=5, weights='uniform', algorithm='auto', leaf_size=30, p=2, metric='euclidean', metric_params=None, n_jobs=1)

    knn.fit(X_train, y_train)
    predicted = knn.predict(X_test)
    acc = metrics.accuracy_score(y_test, predicted)
    print('KNN with BOW accuracy = ' + str(acc * 100) + '%')

    scores = cross_val_score(knn, X_train, y_train, cv=3)
    print("Cross Validation Accuracy: %0.2f" % (scores.mean()))
    print(scores)
    print('\n')
    return predicted, y_test


def tfidf_knn():
    """Method for determining nearest neighbors using tf-idf and K-Nearest Neighbor algorithm"""

    training_data = pd.read_csv('reviews.csv')
    X_train, X_test, y_train, y_test = train_test_split(training_data["sentence"], training_data["sentiment"],
                                                        test_size=0.2, random_state=5)
    X_train, X_test = createTFIDF(X_train, X_test, remove_stopwords=True, lemmatize=True, stemmer=False)
    # print(X_train)
    knn = neighbors.KNeighborsClassifier(n_neighbors=5, weights='distance', algorithm='brute', leaf_size=30, p=2,
                                         metric='cosine', metric_params=None, n_jobs=1)

    knn.fit(X_train, y_train)
    predicted = knn.predict(X_test)
    acc = metrics.accuracy_score(y_test, predicted)
    print('KNN with TFIDF accuracy = ' + str(acc * 100) + '%')

    scores = cross_val_score(knn, X_train, y_train, cv=3)
    print("Cross Validation Accuracy: %0.2f" % (scores.mean()))
    print(scores)
    return predicted, y_test

Note: Cross-validation will be discussed in detail in the upcoming lab session.

In [11]:
## KNN accuracy after using BoW
predicted, y_test = bow_knn()

  soup = BeautifulSoup(text, "lxml")


KNN with BOW accuracy = 62.30366492146597%




Cross Validation Accuracy: 0.62
[0.60784314 0.58431373 0.66141732]




In [12]:
## KNN accuracy after using TFIDF
predicted, y_test = tfidf_knn()

  soup = BeautifulSoup(text, "lxml")


KNN with TFIDF accuracy = 70.15706806282722%
Cross Validation Accuracy: 0.73
[0.7254902  0.74117647 0.72834646]




# Section 4: SPAM TEXT DATASET
Now let's use what we've learnt to classify texts as spam or not spam.

In [13]:
# Upload the spam text data CSV file that has been shared with you. You can also download the file from https://www.kaggle.com/datasets/uciml/sms-spam-collection-dataset
# Run this cell, click on the 'Choose files' button and upload the file.
from google.colab import files
uploaded = files.upload()

Saving spam.csv to spam.csv


In [14]:
import pandas as pd
df = pd.read_csv('spam.csv', error_bad_lines=False)
df



  df = pd.read_csv('spam.csv', error_bad_lines=False)


Unnamed: 0,Category,Message
0,ham,"Go until jurong point, crazy.. Available only ..."
1,ham,Ok lar... Joking wif u oni...
2,spam,Free entry in 2 a wkly comp to win FA Cup fina...
3,ham,U dun say so early hor... U c already then say...
4,ham,"Nah I don't think he goes to usf, he lives aro..."
...,...,...
5567,spam,This is the 2nd time we have tried 2 contact u...
5568,ham,Will Ã¼ b going to esplanade fr home?
5569,ham,"Pity, * was in mood for that. So...any other s..."
5570,ham,The guy did some bitching but I acted like i'd...


In [15]:
df['Category'] = df['Category'].map({'ham': 0, 'spam': 1})

In [16]:
df.head(5)

Unnamed: 0,Category,Message
0,0,"Go until jurong point, crazy.. Available only ..."
1,0,Ok lar... Joking wif u oni...
2,1,Free entry in 2 a wkly comp to win FA Cup fina...
3,0,U dun say so early hor... U c already then say...
4,0,"Nah I don't think he goes to usf, he lives aro..."


In [17]:
len(df)

5572

In [18]:
from sklearn import metrics, neighbors
from sklearn.model_selection import train_test_split, cross_val_score, cross_val_predict

## TASK - 2: Tweak the models below and see results with different parameters and distance metrics.

def bow_knn():
    """Method for determining nearest neighbors using bag-of-words and K-Nearest Neighbor algorithm"""

    training_data = pd.read_csv('spam.csv')
    training_data['Category'] = training_data['Category'].map({'ham': 0, 'spam': 1})
    X_train, X_test, y_train, y_test = train_test_split(training_data["Message"], training_data["Category"], test_size=0.2, random_state=5)
    X_train, X_test = createBagOfWords(X_train, X_test, remove_stopwords=True, lemmatize=True, stemmer=False)
    knn = neighbors.KNeighborsClassifier(n_neighbors=5, weights='uniform', algorithm='auto', leaf_size=30, p=2, metric='euclidean', metric_params=None, n_jobs=1)

    knn.fit(X_train, y_train)
    predicted = knn.predict(X_test)
    acc = metrics.accuracy_score(y_test, predicted)
    print('KNN with BOW accuracy = ' + str(acc * 100) + '%')

    scores = cross_val_score(knn, X_train, y_train, cv=3)
    print("Cross Validation Accuracy: %0.2f" % (scores.mean()))
    print(scores)
    print('\n')
    return predicted, y_test


def tfidf_knn():
    """Method for determining nearest neighbors using tf-idf and K-Nearest Neighbor algorithm"""

    training_data = pd.read_csv('spam.csv')
    training_data['Category'] = training_data['Category'].map({'ham': 0, 'spam': 1})
    X_train, X_test, y_train, y_test = train_test_split(training_data["Message"], training_data["Category"], test_size=0.2, random_state=5)
    X_train, X_test = createTFIDF(X_train, X_test, remove_stopwords=True, lemmatize=True, stemmer=False)
    knn = neighbors.KNeighborsClassifier(n_neighbors=5, weights='distance', algorithm='brute', leaf_size=30, p=2, metric='cosine', metric_params=None, n_jobs=1)

    knn.fit(X_train, y_train)
    predicted = knn.predict(X_test)
    acc = metrics.accuracy_score(y_test, predicted)
    print('KNN with TFIDF accuracy = ' + str(acc * 100) + '%')

    scores = cross_val_score(knn, X_train, y_train, cv=3)
    print("Cross Validation Accuracy: %0.2f" % (scores.mean()))
    print(scores)
    return predicted, y_test

In [19]:
# This cell may take some time to run
predicted, y_test = bow_knn()

  soup = BeautifulSoup(text, "lxml")


KNN with BOW accuracy = 92.19730941704036%
Cross Validation Accuracy: 0.91
[0.90713324 0.90040377 0.91245791]




In [20]:
# This cell may take some time to run
predicted, y_test = tfidf_knn()

  soup = BeautifulSoup(text, "lxml")


KNN with TFIDF accuracy = 98.56502242152466%
Cross Validation Accuracy: 0.97
[0.96837147 0.96769852 0.96363636]


### Questions to Think About and Answer
1. Why does the TF-IDF approach generally result in a better accuracy than Bag-of-Words ?
2. Can you think of techniques that are better than both BoW and TF-IDF ?
3. Read about Stemming and Lemmatization from the resources given below. Think about the pros/cons of each.

### Useful Resources for further reading
1. Stemming and Lemmatization: https://nlp.stanford.edu/IR-book/html/htmledition/stemming-and-lemmatization-1.html
2. TF-IDF and BoW : https://www.analyticsvidhya.com/blog/2020/02/quick-introduction-bag-of-words-bow-tf-idf/
3. TF-IDF: https://scikit-learn.org/stable/modules/generated/sklearn.feature_extraction.text.TfidfTransformer.html


1

The TF-IDF (Term Frequency-Inverse Document Frequency) approach often performs better than the Bag-of-Words (BoW) model in certain natural language processing (NLP) tasks because it addresses some limitations of the BoW model. Here are some reasons why TF-IDF can lead to better accuracy:

1. **Term Importance Weighting:**
   - In the Bag-of-Words model, each word in a document is represented by a count, and the order of words is ignored. This representation does not consider the importance of words in distinguishing between documents.
   - TF-IDF assigns weights to words based on their frequency (TF) in a document and their rarity (IDF) across the entire corpus. This means that common words that appear in many documents receive lower weights, while rare words that are specific to certain documents receive higher weights. This allows the model to focus on the discriminative power of terms.

2. **Handling Stop Words:**
   - Bag-of-Words does not distinguish between common words (stop words) and rare words. Common words may not contribute much to the meaning of a document but can dominate the representation due to their high frequency.
   - TF-IDF mitigates the impact of common words by assigning lower weights to them based on their high document frequency (low IDF).

3. **Normalization:**
   - TF-IDF normalizes the term frequencies to account for variations in document lengths. This normalization helps in making the representations more robust and less dependent on the absolute frequency of terms.

4. **Addressing Word Frequency Imbalances:**
   - In Bag-of-Words, frequent words have higher counts and can dominate the representation. TF-IDF reduces the impact of high-frequency terms, giving more weight to terms that are discriminative and unique to specific documents.

5. **Handling Synonyms:**
   - TF-IDF can capture synonymy to some extent. Synonymous terms that are used interchangeably across documents may have similar TF-IDF scores, helping the model understand their semantic similarity.

6. **Global Context Consideration:**
   - TF-IDF considers the global context of the entire corpus when assigning weights. This helps in capturing the overall importance of a term in the entire dataset, making the representation more informative.

While TF-IDF can provide better accuracy in certain scenarios, it's important to note that the effectiveness of either approach depends on the specific task and dataset. In some cases, Bag-of-Words might be sufficient, especially when the task involves capturing local word patterns rather than considering the semantic meaning of words. Additionally, more advanced techniques, such as word embeddings or transformer-based models, have gained popularity and often outperform both Bag-of-Words and TF-IDF in various NLP tasks.

2

Yes, there are several advanced techniques in natural language processing (NLP) that have shown superior performance compared to Bag-of-Words (BoW) and TF-IDF in various tasks. Some of these techniques include:

1. **Word Embeddings:**
   - Word embeddings, such as Word2Vec, GloVe, and FastText, represent words as dense vectors in a continuous vector space. These embeddings capture semantic relationships between words and are capable of capturing more nuanced meanings than BoW or TF-IDF.

2. **Doc2Vec (Paragraph Embeddings):**
   - Doc2Vec extends the idea of Word2Vec to represent entire documents as continuous vectors. This approach captures not only word semantics but also the overall context of a document.

3. **BERT (Bidirectional Encoder Representations from Transformers):**
   - BERT, and transformer-based models in general, have revolutionized NLP. BERT considers bidirectional context during pre-training and has shown state-of-the-art performance in a wide range of tasks, including text classification, named entity recognition, and question answering.

4. **TF-IDF with Word Embeddings:**
   - Combining TF-IDF with word embeddings can leverage the benefits of both approaches. TF-IDF can capture document-specific term importance, while word embeddings provide a more sophisticated representation of word semantics.

5. **ULMFiT (Universal Language Model Fine-tuning):**
   - ULMFiT is a transfer learning approach that pre-trains a language model on a large corpus and fine-tunes it on a smaller task-specific dataset. This has been effective in improving performance, especially in scenarios with limited labeled data.

6. **ELMo (Embeddings from Language Models):**
   - ELMo generates context-dependent word embeddings by considering the entire sentence context. It captures word meanings that can change based on their context within a sentence.

7. **Attention Mechanisms:**
   - Attention mechanisms, as used in transformers, allow models to focus on different parts of the input sequence when making predictions. This enables the model to capture long-range dependencies and relationships between words.

8. **Subword Embeddings:**
   - Models like Byte Pair Encoding (BPE) and SentencePiece can be used to generate subword embeddings. These embeddings are capable of capturing morphological variations and can handle out-of-vocabulary words more effectively.

These advanced techniques often outperform traditional Bag-of-Words and TF-IDF representations, especially in tasks requiring a deep understanding of language semantics, context, and relationships between words. The choice of the most suitable technique depends on the specific task, dataset characteristics, and available computational resources.

3

I don't have the ability to access external resources or read specific articles. However, I can provide you with information about stemming and lemmatization based on my training data.

**Stemming:**

**Pros:**
1. **Simplicity:** Stemming is a simpler and computationally less expensive process compared to lemmatization. It involves removing or replacing suffixes to obtain the root form of a word.
  
2. **Reduced Vocabulary:** Stemming can help in reducing the vocabulary size by grouping words with similar stems together. This can be beneficial in tasks where a smaller feature space is desired.

**Cons:**
1. **Over-Stemming:** Stemming may result in over-stemming, where the algorithm removes too many characters and reduces words to the point where the meaning is altered. For example, "universe" and "university" might be stemmed to the same root.

2. **Lack of Semantic Understanding:** Stemming doesn't consider the context or meaning of words. It may produce stems that are not actual words or may have different meanings in different contexts.

**Lemmatization:**

**Pros:**
1. **Accurate Word Representation:** Lemmatization provides a more accurate representation of the base or dictionary form of a word (lemma). This can be important in applications where the semantic meaning of words needs to be preserved.

2. **Context Awareness:** Lemmatization considers the context of the word in a sentence, providing better results in tasks where the meaning of words depends on their context.

**Cons:**
1. **Computational Complexity:** Lemmatization is computationally more complex compared to stemming. It involves dictionary lookups and morphological analysis, making it slower and potentially resource-intensive.

2. **Increased Vocabulary Size:** Since lemmatization aims to retain the actual base form of words, it may result in a larger vocabulary size compared to stemming. This can be a consideration in tasks where a smaller feature space is preferred.

**Choosing Between Stemming and Lemmatization:**

- **Application Requirements:** The choice between stemming and lemmatization often depends on the specific requirements of the NLP task. If computational efficiency is crucial and a less sophisticated approach is acceptable, stemming might be preferred. If preserving the semantic meaning and context of words is more important, lemmatization may be a better choice.

- **Task Type:** In information retrieval or text mining tasks where word matching is crucial, stemming may be sufficient. In tasks like sentiment analysis or question answering where understanding the semantics is essential, lemmatization might be more appropriate.

- **Language Characteristics:** The effectiveness of stemming and lemmatization can vary based on the characteristics of the language being processed. Some languages may benefit more from one technique over the other.

Ultimately, the decision to use stemming or lemmatization depends on the trade-offs between simplicity, computational cost, and the linguistic accuracy required for a specific NLP task.