# Document clustering

This is an example showing how the scikit-learn can be used to cluster documents by topics using a bag-of-words approach.
TfidfVectorizer uses a in-memory vocabulary (a python dict) to map the most frequent words to features indices and hence compute a word occurrence frequency (sparse) matrix. The word frequencies are then reweighted using the Inverse Document Frequency (IDF) vector collected feature-wise over the corpus.


In [7]:
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.cluster import KMeans
from sklearn.metrics import adjusted_rand_score

documents = ["Human machine interface for lab abc computer applications",
             "A survey of user opinion of computer system response time",
             "The EPS user interface management system",
             "System and human system engineering testing of EPS",
             "Relation of user perceived response time to error measurement",
             "The generation of random binary unordered trees",
             "The intersection graph of paths in trees",
             "Graph minors IV Widths of trees and well quasi ordering",
             "Graph minors A survey"]

vectorize the text i.e. convert the strings to numeric features

In [8]:
vectorizer = TfidfVectorizer(stop_words='english')
X = vectorizer.fit_transform(documents)

cluster documents

In [16]:
true_k = 2
model = KMeans(n_clusters=true_k, init='k-means++', max_iter=100, n_init=1)
model.fit(X)

KMeans(algorithm='auto', copy_x=True, init='k-means++', max_iter=100,
    n_clusters=2, n_init=1, n_jobs=1, precompute_distances='auto',
    random_state=None, tol=0.0001, verbose=0)

print top terms per cluster clusters

In [17]:
print("Top terms per cluster:")
order_centroids = model.cluster_centers_.argsort()[:, ::-1]
terms = vectorizer.get_feature_names()
for i in range(true_k):
    print ("Cluster :",  i)
    for ind in order_centroids[i, :10]:
        print (terms[ind])
    

Top terms per cluster:
Cluster : 0
graph
trees
minors
human
survey
intersection
paths
testing
engineering
generation
Cluster : 1
user
response
time
management
eps
interface
opinion
relation
perceived
error


### Reference:
 - http://scikit-learn.org/stable/auto_examples/text/document_clustering.html
 - https://datasciencelab.wordpress.com/2013/12/12/clustering-with-k-means-in-python/