Caveat lector — built quickly with a guild of magic elves. If something is wrong, please tell me. — Aguas Infinitas, Star
the record · 2006 · sentiment analysis

A corpus-based approach to finding happiness

Rada Mihalcea & Hugo Liu (2006). Proceedings of the AAAI Spring Symposium on Computational Approaches to Analyzing Weblogs, AAAI Press.

272 citations
PDF

This paper treats mood-tagged LiveJournal posts as a form of linguistic ethnography. From 5,000 happy and 5,000 sad entries, a unigram Naive Bayes model reached 79.13 percent accuracy and yielded a 446-word happiness lexicon. The derived words correlated moderately with ANEW pleasure and dominance scores but not arousal. Further analyses found happy words more social and sad words more human-centered, and explored grammatical mode, time, and repeated phrases. Several conclusions are deliberately playful rather than causal. The hourly curve, for example, combines web-search counts for a word plus an hour; it does not use the posts' timestamps. Voluntary mood labels and LiveJournal users are self-selected, first WordNet senses can be wrong, and public posts cannot establish “private” happiness.

← the record, in the papers room