A list of data driven lexica developed by the World Well-Being Project.
Read the full publication here.
@inproceedings{sap2014developing,
author={Sap, Maarten and Park, Greg and Eichstaedt, Johannes C and Kern, Margaret L and Stillwell, David J and Kosinski, Michal and Ungar, Lyle H and Schwartz, H Andrew},
title={Developing age and gender predictive lexica over social media},
booktitle={Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP)},
year={2014},
}
Read the full publication here.
@{h. andrew schwartz2016predicting,
author={H. Andrew Schwartz, Maarten Sap, Margaret L. Kern, Johannes C. Eichstaedt, Adam Kapelner, Megha Agrawal, Eduardo Blanco, Lukasz Dziurzynski, Gregory Park, David Stillwell, Michal Kosinski, Martin E.P. Seligman, Lyle H. Ungar.},
title={Predicting Individual Well-Being Through the Language of Social Media},
year={2016},
pages={516-527}
}
Read the full publication here.
@inproceedings{smith2016does,
title={Does ‘well-being’translate on Twitter?},
author={Smith, Laura and Giorgi, Salvatore and Solanki, Rishi and Eichstaedt, Johannes and Schwartz, H Andrew and Abdul-Mageed, Muhammad and Buffone, Anneke and Ungar, Lyle},
booktitle={Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing},
pages={2042--2047},
year={2016}
}
Read the full publication here.
@inproceedings{preoctiuc2016modelling,
title={Modelling valence and arousal in facebook posts},
author={Preo{\c{t}}iuc-Pietro, Daniel and Schwartz, H Andrew and Park, Gregory and Eichstaedt, Johannes and Kern, Margaret and Ungar, Lyle and Shulman, Elisabeth},
booktitle={Proceedings of the 7th Workshop on Computational Approaches to Subjectivity, Sentiment and Social Media Analysis},
pages={9--15},
year={2016}
}
Read the full publication here.
@inproceedings{schwartz2015extracting,
title={Extracting human temporal orientation from Facebook language},
author={Schwartz, H Andrew and Park, Gregory and Sap, Maarten and Weingarten, Evan and Eichstaedt, Johannes and Kern, Margaret and Stillwell, David and Kosinski, Michal and Berger, Jonah and Seligman, Martin and others},
booktitle={Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies},
pages={409--419},
year={2015}
}
A weighted lexicon is often applied as the sum of all weighted word relative frequencies over a document:
where is the lexicon
weight for the word,
is frequency of the word in the document (or for a given user), and
is the total word count for that document (or user).
For example, let's say a lexicon has the following weights for words a, b, and c:
and two documents with the following frequencies of words:
therefore the total word uses in the documents are:
The documents' lexicon usage are given by summing the weighted relative frequencies:
Once the usages have been computed, the intercept of the lexicon needs to be added to the usages:
If the lexicon used represents age, and
are the predicted ages for both documents. If it represents gender, simply take the sign of the result and if it's positive, the document is female, else it's male.
Unless specified in the lexica's subdirectory all lexica are licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 3.0 Unported License.
Developed by the World Well-Being Project based out of the University of Pennsylvania and Stony Brook University.