Caveat lector — built quickly with a guild of magic elves. If something is wrong, please tell me. — Aguas Infinitas, Star
the record · 2007 · selected writing

Of men, women, and computers: data-driven gender modeling for improved user interfaces

Hugo Liu & Rada Mihalcea (2007). Proceedings of the International Conference on Weblogs and Social Media (ICWSM), Boulder, CO.

PDF

From roughly 300,000 Blogspot entries gathered on two days in 2006, this paper formed a balanced corpus of 75,000 male-labeled and 75,000 female-labeled posts. A unigram classifier reached 71 percent author-gender accuracy; further analyses explored time, food, color, size, pronouns, social reference, and affect. GenderLens then reranked Google News with the 14,000 most discriminating words. Thirty readers compared hidden male- and female-ranked columns; reported preferences aligned in four of five categories, while entertainment did not. The limits matter as much as the result: binary self-report, one platform, a two-day window, thirty participants, correlational and stereotype-adjacent interpretation, and one category accepted at a weaker significance threshold. These are patterns in particular corpora and participants, not fixed properties of women and men.

← the record, in the papers room