# word2vec

how to get feature vectors:
1. take fake problem
2. solve it using neural network
3. we get word embedding(feature vectors) as side effect

CBOW continuous bag of words:  given context words , predict target word

skip gram: given target , predict context words

# code

The dataset we are using here is a subset of Amazon reviews from the Cell Phones & Accessories category. The data is stored as a JSON file and can be read using pandas.

Link to the Dataset: http://snap.stanford.edu/data/amazon/productGraph/categoryFiles/reviews_Cell_Phones_and_Accessories_5.json.gz

In [2]:
import gensim # NLP library for python
import pandas as pd

In [3]:
df = pd.read_json(".\\datasets\\word2vec\\Cell_Phones_and_Accessories_5.json", lines=True)

In [13]:
df

Unnamed: 0,reviewerID,asin,reviewerName,helpful,reviewText,overall,summary,unixReviewTime,reviewTime
0,A30TL5EWN6DFXT,120401325X,christina,"[0, 0]",They look good and stick good! I just don't li...,4,Looks Good,1400630400,"05 21, 2014"
1,ASY55RVNIL0UD,120401325X,emily l.,"[0, 0]",These stickers work like the review says they ...,5,Really great product.,1389657600,"01 14, 2014"
2,A2TMXE2AFO7ONB,120401325X,Erica,"[0, 0]",These are awesome and make my phone look so st...,5,LOVE LOVE LOVE,1403740800,"06 26, 2014"
3,AWJ0WZQYMYFQ4,120401325X,JM,"[4, 4]",Item arrived in great time and was in perfect ...,4,Cute!,1382313600,"10 21, 2013"
4,ATX7CZYFXI1KW,120401325X,patrice m rogoza,"[2, 3]","awesome! stays on, and looks great. can be use...",5,leopard home button sticker for iphone 4s,1359849600,"02 3, 2013"
...,...,...,...,...,...,...,...,...,...
194434,A1YMNTFLNDYQ1F,B00LORXVUE,eyeused2loveher,"[0, 0]",Works great just like my original one. I reall...,5,This works just perfect!,1405900800,"07 21, 2014"
194435,A15TX8B2L8B20S,B00LORXVUE,Jon Davidson,"[0, 0]",Great product. Great packaging. High quality a...,5,Great replacement cable. Apple certified,1405900800,"07 21, 2014"
194436,A3JI7QRZO1QG8X,B00LORXVUE,Joyce M. Davidson,"[0, 0]","This is a great cable, just as good as the mor...",5,Real quality,1405900800,"07 21, 2014"
194437,A1NHB2VC68YQNM,B00LORXVUE,Nurse Farrugia,"[0, 0]",I really like it becasue it works well with my...,5,I really like it becasue it works well with my...,1405814400,"07 20, 2014"


we are creating word2vec model using only review text

## Simple Preprocessing and Tokenization

The first thing to do for any data science task is to clean the data. For NLP, we apply various processing like converting all the words to lower case, trimming spaces, removing punctuations. This is something we will do over here too.


Additionally, we can also remove stop words like 'and', 'or', 'is', 'the', 'a', 'an' and convert words to their root forms like 'running' to 'run'.

In [7]:
review_text = df.reviewText.apply(gensim.utils.simple_preprocess)

In [9]:
review_text[0]

['they',
 'look',
 'good',
 'and',
 'stick',
 'good',
 'just',
 'don',
 'like',
 'the',
 'rounded',
 'shape',
 'because',
 'was',
 'always',
 'bumping',
 'it',
 'and',
 'siri',
 'kept',
 'popping',
 'up',
 'and',
 'it',
 'was',
 'irritating',
 'just',
 'won',
 'buy',
 'product',
 'like',
 'this',
 'again']

## Training the Word2Vec Model
Train the model for reviews. Use a window of size 10 i.e. 10 words before the present word and 10 words ahead. A sentence with at least 2 words should only be considered, configure this using min_count parameter.

Workers define how many CPU threads to be used.

In [10]:
model = gensim.models.Word2Vec(
    window=10,
    min_count=2,
    workers=4
)

In [11]:
model.build_vocab(review_text, progress_per=1000)

In [12]:
model.epochs

5

In [14]:
model.train(review_text, total_examples=model.corpus_count, epochs=model.epochs) # corpus_count is our dataset size

(61508038, 83868975)

In [15]:
model.save('.//word2vec_amazon_reviews.model')

### experiment the model:

In [27]:
model.wv.most_similar('lame') # find words that have similar weights

[('buggy', 0.553232729434967),
 ('irritating', 0.5526553392410278),
 ('bad', 0.5370842218399048),
 ('gimmick', 0.5243535041809082),
 ('obvious', 0.5088638663291931),
 ('annoying', 0.5060213208198547),
 ('frustrating', 0.49379169940948486),
 ('laggy', 0.4848453104496002),
 ('confusing', 0.48418086767196655),
 ('sucks', 0.47617700695991516)]

In [24]:
model.wv.similarity(w1='great', w2='good')

0.7755563