<a href="https://colab.research.google.com/github/cbadenes/notebooks/blob/main/probabilistic_topic_models/LabeledLDA_Topics.ipynb" target="_parent"><img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab"/></a>

Use a Python implementation of Labeled-LDA based on Gibbs sampling
 from [JoeZJH](https://github.com/JoeZJH/Labeled-LDA-Python)

In [1]:
import os

!git clone https://github.com/JoeZJH/Labeled-LDA-Python
os.chdir('Labeled-LDA-Python/')
!ls -la

Cloning into 'Labeled-LDA-Python'...
remote: Enumerating objects: 175, done.[K
remote: Counting objects: 100% (24/24), done.[K
remote: Compressing objects: 100% (14/14), done.[K
remote: Total 175 (delta 8), reused 24 (delta 8), pack-reused 151[K
Receiving objects: 100% (175/175), 322.44 KiB | 17.91 MiB/s, done.
Resolving deltas: 100% (69/69), done.
total 44
drwxr-xr-x 6 root root 4096 Feb 28 15:25 .
drwxr-xr-x 1 root root 4096 Feb 28 15:25 ..
drwxr-xr-x 2 root root 4096 Feb 28 15:25 assets
drwxr-xr-x 2 root root 4096 Feb 28 15:25 example
drwxr-xr-x 8 root root 4096 Feb 28 15:25 .git
-rw-r--r-- 1 root root   47 Feb 28 15:25 .gitignore
-rw-r--r-- 1 root root 1069 Feb 28 15:25 LICENSE
drwxr-xr-x 2 root root 4096 Feb 28 15:25 model
-rw-r--r-- 1 root root 5684 Feb 28 15:25 README.md
-rw-r--r-- 1 root root   14 Feb 28 15:25 requirements.txt


## Initialize a Labeled Topic Model
* `K`: the number of topics
* `M`: the number of documents
* `W`: the corpus, a list of terms list,
              W[m] is the document vector, W[m][n] is the id of the term
* `T`: the number of terms
* `WN`: the number of all words in W
* `LN`: the number of all original labels
* `alpha`: doc-topics ratio
* `eta`: topic-terms ratio

In [11]:
import model.labeled_lda as llda

# initialize data
labeled_documents = [("example example example example example"*10, ["example"]),
                     ("test llda model test llda model test llda model"*10, ["test", "llda_model"]),
                     ("example test example test example test example test"*10, ["example", "test"]),
                     ("good perfect good good perfect good good perfect good "*10, ["positive"]),
                     ("bad bad down down bad bad down"*10, ["negative"])]

# new a Labeled LDA model
# llda_model = llda.LldaModel(labeled_documents=labeled_documents, alpha_vector="50_div_K", eta_vector=0.001)
# llda_model = llda.LldaModel(labeled_documents=labeled_documents, alpha_vector=0.02, eta_vector=0.002)
llda_model = llda.LldaModel(labeled_documents=labeled_documents, alpha_vector=0.01, eta_vector=0.001)
print(llda_model)


Labeled-LDA Model:
	K = 6
	M = 5
	T = 12
	WN = 344
	LN = 7
	alpha = 0.01
	eta = 0.001
	perplexity = 3.9683718068695697
	


## Train 

In [12]:
# training
# llda_model.training(iteration=10, log=True)
while True:
    print("iteration %s sampling..." % (llda_model.iteration + 1))
    llda_model.training(1)
    print("after iteration: %s, perplexity: %s" % (llda_model.iteration, llda_model.perplexity()))
    print("delta beta: %s" % llda_model.delta_beta)
    if llda_model.is_convergent(method="beta", delta=0.01):
        break

iteration 1 sampling...
gibbs sample count:  344
after iteration: 1, perplexity: 2.948204518307306
delta beta: 1.829836168167582
iteration 2 sampling...
gibbs sample count:  344
after iteration: 2, perplexity: 2.6239852397791985
delta beta: 1.6409206682875814
iteration 3 sampling...
gibbs sample count:  344
after iteration: 3, perplexity: 2.573645211454954
delta beta: 1.2799773845474132
iteration 4 sampling...
gibbs sample count:  344
after iteration: 4, perplexity: 2.576271045875313
delta beta: 0.11655324642799364
iteration 5 sampling...
gibbs sample count:  344
after iteration: 5, perplexity: 2.576271045875313
delta beta: 0.0


## Update an existing model with new documents

In [13]:
# update
print("before updating: ", llda_model)
update_labeled_documents = [("new example test example test example test example test", ["example", "test"])]
llda_model.update(labeled_documents=update_labeled_documents)
print("after updating: ", llda_model)

before updating:  
Labeled-LDA Model:
	K = 6
	M = 5
	T = 12
	WN = 344
	LN = 7
	alpha = 0.01
	eta = 0.001
	perplexity = 2.576271045875313
	
after updating:  
Labeled-LDA Model:
	K = 6
	M = 6
	T = 13
	WN = 353
	LN = 11
	alpha = 0.01
	eta = 0.001
	perplexity = 2.683861643881013
	


In [14]:
# train again
# llda_model.training(iteration=10, log=True)
while True:
    print("iteration %s sampling..." % (llda_model.iteration + 1))
    llda_model.training(1)
    print("after iteration: %s, perplexity: %s" % (llda_model.iteration, llda_model.perplexity()))
    print("delta beta: %s" % llda_model.delta_beta)
    if llda_model.is_convergent(method="beta", delta=0.01):
        break

iteration 6 sampling...
gibbs sample count:  353
after iteration: 6, perplexity: 2.619759333193657
delta beta: 0.953266522426173
iteration 7 sampling...
gibbs sample count:  353
after iteration: 7, perplexity: 2.6014164700095175
delta beta: 1.166284331758653
iteration 8 sampling...
gibbs sample count:  353
after iteration: 8, perplexity: 2.5969196211475443
delta beta: 0.2049851733233044
iteration 9 sampling...
gibbs sample count:  353
after iteration: 9, perplexity: 2.595314188042728
delta beta: 0.028529499507375995
iteration 10 sampling...
gibbs sample count:  353
after iteration: 10, perplexity: 2.5969196211475443
delta beta: 0.028529499507375995
iteration 11 sampling...
gibbs sample count:  353
after iteration: 11, perplexity: 2.5961458918908167
delta beta: 0.10507898616959299
iteration 12 sampling...
gibbs sample count:  353
after iteration: 12, perplexity: 2.598830102713859
delta beta: 0.0553038760279235
iteration 13 sampling...
gibbs sample count:  353
after iteration: 13, perple

## Explore Topics

In [21]:
for topic in llda_model.topic_vocabulary.keys():
  print("Topic '",topic,"' :")
  for term in llda_model.top_terms_of_topic(topic,5):
    print("\t-",term)

Topic ' common_topic ' :
	- ('new', 0.9881539980256663)
	- ('exampleexample', 0.000987166831194472)
	- ('perfect', 0.000987166831194472)
	- ('down', 0.000987166831194472)
	- ('testexample', 0.000987166831194472)
Topic ' positive ' :
	- ('good', 0.6665814937842309)
	- ('perfect', 0.33329630164531776)
	- ('exampleexample', 1.1109506404630438e-05)
	- ('down', 1.1109506404630438e-05)
	- ('testexample', 1.1109506404630438e-05)
Topic ' llda_model ' :
	- ('llda', 0.49990835319014215)
	- ('model', 0.3499408461500009)
	- ('modeltest', 0.1499841700964791)
	- ('exampleexample', 1.666305633779348e-05)
	- ('perfect', 1.666305633779348e-05)
Topic ' negative ' :
	- ('bad', 0.5081048301181715)
	- ('down', 0.3442053332896268)
	- ('downbad', 0.14752593709537312)
	- ('exampleexample', 1.6389949682854473e-05)
	- ('perfect', 1.6389949682854473e-05)
Topic ' example ' :
	- ('example', 0.8674959198976664)
	- ('exampleexample', 0.1323423463161454)
	- ('perfect', 1.4703071471630419e-05)
	- ('down', 1.4703071471

## Inference

* `document`: some sentence like "this is a method for inference"
* `times`: the number of samples of the target distribution (one whole iteration(sample for all words) generates a sample)
* `iteration`: the times of iteration until Markov chain converges

In [15]:
# inference
# note: the result topics may be different for difference training, because gibbs sampling is a random algorithm
document = "example llda model example example good perfect good perfect good perfect" * 100

topics = llda_model.inference(document=document, iteration=100, times=10)
print(topics)

[('positive', 0.5554065139791144), ('common_topic', 0.4445491430725229), ('llda_model', 1.1085737090659159e-05), ('negative', 1.1085737090659159e-05), ('example', 1.1085737090659159e-05), ('test', 1.1085737090659159e-05)]


## Evaluation

In [16]:
# perplexity
# calculate perplexity on test data
perplexity = llda_model.perplexity(documents=["example example example example example",
                                              "test llda model test llda model test llda model",
                                              "example test example test example test example test",
                                              "good perfect good good perfect good good perfect good",
                                              "bad bad down down bad bad down"],
                                   iteration=30,
                                   times=10)
print("perplexity on test data: %s" % perplexity)
# calculate perplexity on training data
print("perplexity on training data: %s" % llda_model.perplexity())

perplexity on test data: 5.2135218369149285
perplexity on training data: 2.5990739998167807


## Save Model

In [9]:
# save to disk
save_model_dir = "data/model"
# llda_model.save_model_to_dir(save_model_dir, save_derivative_properties=True)
llda_model.save_model_to_dir(save_model_dir)

## Load Model

In [10]:
# load from disk
llda_model_new = llda.LldaModel()
llda_model_new.load_model_from_dir(save_model_dir, load_derivative_properties=False)
print("llda_model_new", llda_model_new)
print("llda_model", llda_model)
print("Top-5 terms of topic 'negative': ", llda_model.top_terms_of_topic("negative", 5, False))
print("Doc-Topic Matrix: \n", llda_model.theta)
print("Topic-Term Matrix: \n", llda_model.beta)

llda_model_new 
Labeled-LDA Model:
	K = 6
	M = 6
	T = 13
	WN = 353
	LN = 11
	alpha = 0.01
	eta = 0.001
	perplexity = 2.5742566681649377
	
llda_model 
Labeled-LDA Model:
	K = 6
	M = 6
	T = 13
	WN = 353
	LN = 11
	alpha = 0.01
	eta = 0.001
	perplexity = 2.5742566681649377
	
Top-5 terms of topic 'negative':  ['bad', 'down', 'downbad', 'exampleexample', 'perfect']
Doc-Topic Matrix: 
 [[0.         0.         0.         0.         1.         0.        ]
 [0.         0.         1.         0.         0.         0.        ]
 [0.         0.         0.         0.         0.43661972 0.56338028]
 [0.87777778 0.12222222 0.         0.         0.         0.        ]
 [0.         0.         0.         1.         0.         0.        ]
 [0.00110742 0.         0.         0.         0.4440753  0.55481728]]
Topic-Term Matrix: 
 [[1.26561452e-05 3.79697012e-01 1.26561452e-05 1.26561452e-05
  1.26561452e-05 6.20163771e-01 1.26561452e-05 1.26561452e-05
  1.26561452e-05 1.26561452e-05 1.26561452e-05 1.26561452e