Rada Mihalcea & Hugo Liu (2006). A corpus-based approach to finding happiness. Proceedings of the AAAI Spring Symposium on Computational Approaches to Analyzing Weblogs, AAAI Press.
Mihalcea, R., & Liu, H. (2006). A corpus-based approach to finding happiness. In Proceedings of the AAAI Spring Symposium on Computational Approaches to Analyzing Weblogs. AAAI Press.
@inproceedings{mihalcea2006a,
author = {Rada Mihalcea and Hugo Liu},
title = {A corpus-based approach to finding happiness},
booktitle = {Proceedings of the AAAI Spring Symposium on Computational Approaches to Analyzing Weblogs},
publisher = {AAAI Press},
url = {https://starheartsong.com/papers/pdf/CAAW2006-Happiness.pdf},
year = {2006}
}This paper treats mood-tagged LiveJournal posts as a form of linguistic ethnography. From 5,000 happy and 5,000 sad entries, a unigram Naive Bayes model reached 79.13 percent accuracy and yielded a 446-word happiness lexicon. The derived words correlated moderately with ANEW pleasure and dominance scores but not arousal. Further analyses found happy words more social and sad words more human-centered, and explored grammatical mode, time, and repeated phrases. Several conclusions are deliberately playful rather than causal. The hourly curve, for example, combines web-search counts for a word plus an hour; it does not use the posts' timestamps. Voluntary mood labels and LiveJournal users are self-selected, first WordNet senses can be wrong, and public posts cannot establish “private” happiness.