## Tokenization
Before we can classify any posts, we'll need to clean and tokenize the text data. Use what you remember from the last lesson on NLP to implement the function `tokenize`. This function should perform the following steps on the string, `text`, using nltk:

1. Identify any urls in `text`, and replace each one with the word, `"urlplaceholder"`.
2. Split `text` into tokens.
3. For each token: lemmatize, normalize case, and strip leading and trailing white space.
4. Return the tokens in a list!

For example, this:
```python
text = 'Barclays CEO stresses the importance of regulatory and cultural reform in financial services at Brussels conference  http://t.co/Ge9Lp7hpyG'

tokenize(text)
```
should return this:
```txt
['barclays', 'ceo', 'stress', 'the', 'importance', 'of', 'regulatory', 'and', 'cultural', 'reform', 'in', 'financial', 'service', 'at', 'brussels', 'conference', 'urlplaceholder']
```

Hint: You'll have to add an import statement to use the `re` package (which supports regular expressions) and two import statements to use the appropriate functions from `nltk`! Add them to this first code cell.

In [24]:
# download necessary NLTK data

import nltk
nltk.download(['punkt', 'wordnet'])

# import statements
import pandas as pd
import re
from nltk.tokenize import word_tokenize,sent_tokenize
from nltk.stem.wordnet import WordNetLemmatizer

[nltk_data] Downloading package punkt to /root/nltk_data...
[nltk_data]   Package punkt is already up-to-date!
[nltk_data] Downloading package wordnet to /root/nltk_data...
[nltk_data]   Package wordnet is already up-to-date!


In [25]:
def load_data():
    df = pd.read_csv('corporate_messaging.csv', encoding='latin-1')
    df = df[(df["category:confidence"] == 1) & (df['category'] != 'Exclude')]
    X = df.text.values
    y = df.category.values
    return X, y

#### For step 1, the regular expression to detect a url is given below

In [26]:
url_regex = 'http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\(\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+'


In [27]:
def tokenize(text):
    # get list of all urls using regex
    detected_urls = re.findall(url_regex,text) #all cases of pattern save in string for each case
    print('***',detected_urls)
    pattern = re.compile(r'[^a-zA-Z0-9]')
    
    # replace each url in text string with placeholder
    for url in detected_urls:
        
        text = re.sub(url,'urlplaceholder',text)   #text.replace(url,'urlplaceholder')
        
    text = re.sub(pattern,' ',text)
    # tokenize text
    tokens = word_tokenize(text.lower())
    
    # initiate lemmatizer
    lemmatizer = WordNetLemmatizer()

    # iterate through each token
    clean_tokens = []
    for tok in tokens:
        
        # lemmatize, normalize case, and remove leading/trailing white space
        clean_tok = lemmatizer.lemmatize(lemmatizer.lemmatize(tok.strip()),pos='v')
        clean_tokens.append(clean_tok)

    return clean_tokens

In [28]:
# test out function
X, y = load_data()
for message in X[:10]:
    tokens = tokenize(message)
    print(message)
    print(tokens, '\n')

*** ['http://t.co/Ge9Lp7hpyG']
Barclays CEO stresses the importance of regulatory and cultural reform in financial services at Brussels conference  http://t.co/Ge9Lp7hpyG
['barclays', 'ceo', 'stress', 'the', 'importance', 'of', 'regulatory', 'and', 'cultural', 'reform', 'in', 'financial', 'service', 'at', 'brussels', 'conference', 'urlplaceholder'] 

*** ['http://t.co/LbIqqh3wwG']
Barclays announces result of Rights Issue http://t.co/LbIqqh3wwG
['barclays', 'announce', 'result', 'of', 'right', 'issue', 'urlplaceholder'] 

*** ['http://t.co/YZk24iE8G6']
Barclays publishes its prospectus for its å£5.8bn Rights Issue: http://t.co/YZk24iE8G6
['barclays', 'publish', 'it', 'prospectus', 'for', 'it', '5', '8bn', 'right', 'issue', 'urlplaceholder'] 

*** ['http://t.co/nkuHoAfnSD']
Barclays Group Finance Director Chris Lucas is to step down at the end of the week due to ill health http://t.co/nkuHoAfnSD
['barclays', 'group', 'finance', 'director', 'chris', 'lucas', 'be', 'to', 'step', 'down', '