Caveat lector — this site is being built at unreasonable speed with a guild of magic elves. If something reads not quite right, someone might have gotten carried away. Please tell me so I can fix it. — Aguas Infinitas, Star
the academic record

Papers

From 2002 to 2017, under the name Hugo Liu, I published work on how machines might engage ordinary meaning: common sense, feeling, stories, taste, point of view, and cultural value. The early papers grew from the MIT Media Lab; the later work followed those questions into Hunch, eBay, the Guggenheim, ArtAdvisor, and Artsy. They are quiet ancestors of systems that now speak, recommend, and reason with us. The record is gathered here, and it is free: every paper is one click to read, every citation one click to copy.

6,629 citations 24 h-index 31 i10-index
View on Google Scholar updated July 2026

2002 → 2017

selected writingcomputational aestheticssentiment analysissemantics of natural languagecommon sense reasoninglater work
20022003200420052006200720142017selected writing · 6 papers · 2006–2007 · 780 citationscomputational aesthetics · 9 papers · 2004–2006 · 436 citationssentiment analysis · 4 papers · 2003–2006 · 1,220 citationssemantics of natural language · 8 papers · 2003–2006 · 450 citationscommon sense reasoning · 10 papers · 2002–2004 · 3,445 citationslater work · 2 papers · 2014–2017 · citations not yet reconciled

selected writing

Social network profiles as taste performances journal article

Hugo Liu (2007). Journal of Computer-Mediated Communication 13(1), 252–275, Blackwell Publishing.

780 citations

Most early social-network research studied the link graph. This paper reads what people placed inside the nodes: books, music, films, and interests used to compose a public self. A grounded-theory reading of 100 profiles identified four kinds of taste statement—prestige, differentiation, authenticity, and theatrical persona—then natural-language processing and principal components analysis examined 127,477 MySpace profiles. The latent maps exposed interpretable cultural oppositions; profiles at both popular and obscure extremes were especially coherent. Users' listed interests were also less like those of their Top 8 friends than chance suggested, although several causal explanations fit that result. The study covers one platform and one period; demographic correlations were inconclusive, and the paper calls for longitudinal, cross-platform, cross-cultural comparison.

PDF

Small happiness: aesthetic strategies for witting consumers

Hugo Liu, trans. Wu Gang (2007). Cultural Review, November 2007: 32–39, Shanghai.

This essay asks whether ordinary pleasures can support a durable happiness rather than merely distract from the epic promises of career, romance, and achievement. Small pleasures are inexpensive, numerous, and granular, making them material for repeated experiments in taste and identity. The essay proposes rotating pleasures before familiarity exhausts them, saving “dark horse” pleasures for moments of need, and treating taste as an ecosystem whose diversity makes happiness more resilient. Reappropriation and multipurposing let one object participate in several aesthetics, allowing imagination rather than accumulation to enlarge the field of pleasure. The argument draws on philosophy and examples from soap, food, clothing, and collecting. Its “aesthetic regeneration” and ecosystem analogies are normative proposals, not experimentally measured findings.

PDF

From programming the unconscious to aesthetic technologies Portuguese

Hugo Liu & Paulo Urbano (2007). Interview, Nada 9, Portugal.

This Portuguese-language interview explains the human purpose behind Aesthetiscope, Taste Fabric, and the broader idea of computational aesthetics. Aesthetiscope combines several modes of reading—thought, feeling, sensation, intuition, and cultural association—to render words or poems as abstract fields of color. Taste Fabric learns relationships among interests from large collections of social-network profiles. The interview distinguishes factual common sense from experiential common sensibility: knowledge of how life feels from different perspectives. It also marks firm boundaries. Aesthetiscope must be given a personality setting rather than discovering one, and superficial emotional display should not be confused with genuine machine feeling. The lasting proposal is for technologies that help people encounter more perspectives and cultivate happiness and beauty as capacities.

PDF

Introduction to the semantics of people and culture (editorial preface)

Hugo Liu & Pattie Maes (2007). International Journal on Semantic Web and Information Systems, Special Issue on Semantics of People and Culture (Eds. H. Liu & P. Maes) 3(1), Idea Group.

This editorial preface contrasts an over-architected formal Semantic Web with meanings people already grow through tags, ratings, profiles, and communities. Its principle of semantical ergonomics asks systems to make annotation immediately useful, learn from off-label behavior, and interpret a tag within its social setting. Three themes organize the issue: the politics of tagging as a strategic public act; subjectivity, where collaborative filtering and trust localize truth rather than forcing one global answer; and cultural inheritance, where shared defaults help translate knowledge across communities. Examples and findings come from the four papers introduced in the issue, not new experiments by the editors. “Cultural modules” and programmatic inheritance remain explanatory proposals, but the agenda is concrete: human meaning is situated, participatory, and inseparable from people.

PDF

Of men, women, and computers: data-driven gender modeling for improved user interfaces

Hugo Liu & Rada Mihalcea (2007). Proceedings of the International Conference on Weblogs and Social Media (ICWSM), Boulder, CO.

From roughly 300,000 Blogspot entries gathered on two days in 2006, this paper formed a balanced corpus of 75,000 male-labeled and 75,000 female-labeled posts. A unigram classifier reached 71 percent author-gender accuracy; further analyses explored time, food, color, size, pronouns, social reference, and affect. GenderLens then reranked Google News with the 14,000 most discriminating words. Thirty readers compared hidden male- and female-ranked columns; reported preferences aligned in four of five categories, while entertainment did not. The limits matter as much as the result: binary self-report, one platform, a two-day window, thirty participants, correlational and stereotype-adjacent interpretation, and one category accepted at a weaker significance threshold. These are patterns in particular corpora and participants, not fixed properties of women and men.

PDF

Superconsumer: a postmodern romance

Hugo Liu (2006). Nada 8, Portugal.

The “superconsumer” begins inside cultural systems whose authority feels natural. Contradictions among those systems reveal that their truths are constructed, making culture available as material for deliberate self-formation. The essay describes three experiences in that movement: multicultural encounter exposes disagreement; cosmopolitan experience distinguishes deep cultural languages from flattened commodities; and perspectival experience lets several viewpoints enrich one another. The mature superconsumer still submits to culture's formative power but chooses which languages to inhabit—controls the controller. This is a theoretical romance assembled through Lévi-Strauss, Derrida, Jameson, Bhabha, Simmel, Lacan, Lyotard, and others. It contains no corpus, experiment, or user study, and does not demonstrate that consumers actually pass through the proposed stages.

PDF

computational aesthetics

Unraveling the taste fabric of social networks

Hugo Liu, Pattie Maes & Glorianna Davenport (2006). International Journal on Semantic Web and Information Systems 2(1), 42–71. Reprinted 2008, Idea Group.

250 citations

Taste Fabric turns co-occurrence among favorite things into a reusable cultural map. A six-month crawl gathered 100,000 profiles from two social networks. Heuristic segmentation and an ontology of 22,000 interest and identity descriptors recognized 68 percent of tokens; pointwise mutual information produced a dense relation matrix later reduced to 12,000 descriptors. Spreading activation formed “taste ethoi” for recommendation, comparison, Ambient Semantics, and Identity Mirror. In five-fold held-out-profile tests, the full model scored 0.86 on graded percentile rank, compared with 0.73 for direct pairwise association. The paper also records the costs: overlapping populations, a static snapshot, unrecognized and falsely mapped terms, unevaluated neighborhood morphology, and no re-identification audit. Its metric measures recovery of withheld profile items, not lived recommendation quality.

PDF

Computing Point-of-View: Modeling and Simulating Judgments of Taste PhD dissertation

Hugo Liu (2006). Ph.D. Dissertation, Program in Media Arts & Sciences, MIT, 163pp.

This dissertation asks whether weblogs, profiles, and email contain enough evidence to model how a particular person judges culture, events, food, humor, and perception. Its sequence is acquire, generalize, and apply: read textual traces for affective themes; extend sparse personal evidence with semantic resources and cultural patterns; then use the model in systems for art, introductions, self-reflection, virtual mentors, family recipes, and humor. Results varied by realm. Taste recommendation scored 0.86 against a 0.73 pairwise baseline; attitude models predicted arousal better than baselines but pleasure and dominance less reliably; one perceptual classification reached 0.62 while another was not significant; and a thirty-six-person mentor study improved learning of explicit attitudes and personality. The dissertation calls these evaluations “modestly successful.” Its five realms overlap, several studies are small, and associative reading can assign surrounding affect to the wrong concept.

PDF

Rendering aesthetic impressions of text in color space

Hugo Liu & Pattie Maes (2006). International Journal on Artificial Intelligence Tools 15(4), 515–550, World Scientific.

Aesthetiscope renders a word, poem, or lyric as a 16-by-9 field of color. Five readers—Think, Culturalize, See, Intuit, and Feel—draw on common sense, popular culture, stock imagery, free association, affective vocabulary, and color theory. Their palettes are blended and arranged into one animated response. Four graduate judges found Intuit and Feel the most consistent, while Think and Culturalize were weaker. In a separate study, fifty-one MIT and Harvard students chose between the correct “Golden Setting” image and one generated from another text; correct pairings were selected 75.2 percent of the time for poems and songs and 80.7 percent for words. The setting was manually chosen, samples were narrow, associations culturally bounded, and reader weights were not yet personalized.

PDF

Taste fabrics and the beauty of homogeneity

Hugo Liu, Glorianna Davenport & Pattie Maes (2006). AIS SIGSEMIS Bulletin, vol. 3.

This paper asks whether a homogeneous relation can sometimes reveal more than an elaborate ontology. Favorite things from 100,000 social-network profiles become “taste acts”; domain ontologies normalize them, and co-occurrence produces a 12,000-by-12,000 affinity fabric including 600 identity descriptors. Spreading activation converts a person's interests into a taste ethos, intersects two ethoi for comparison, and surfaces neighbors for recommendation. One relation type, one propagation mechanism, and continuous similarity offer a compact complement to long chains of symbolic categories. The paper is a six-page conceptual and technical condensation of the larger Taste Fabric work, not a separate experiment. It assumes that selected interests form a coherent statement and gains uniformity by discarding distinctions among different kinds of cultural relation.

PDF

Self-reflexive performance: dancing with the computed audience of culture

Hugo Liu & Glorianna Davenport (2005). International Journal of Performance Arts and Digital Media 1(3), 237–247, Intellect Ltd.

Identity Mirror overlays a tracked silhouette with a moving field of cultural descriptors drawn from a person's favorite books, music, films, foods, sports, television, and subcultures. Distance changes descriptive granularity; deliberate and abrupt motion change which tenuous associations remain visible. The mirror draws on Taste Fabric's network of 12,000 cultural symbols, while daily news can bias which facets become salient. The paper frames reflection as a negotiation among performer, culture-as-audience, and the self as meta-audience, with “facets,” “shadows,” and co-performance as ways to explore contextual identity. It is a design and performance-theory paper, not a controlled evaluation. Several off-stage histories and connected performances remain proposals, and every reflection inherits Taste Fabric's coverage and inference limits.

PDF

Synesthetic Recipes: foraging for food with the family, in taste-space

Hugo Liu, Matthew Hockenberry & Ted Selker (2005). ACM SIGGRAPH 2005 Posters, Los Angeles.

Synesthetic Recipes begins with the desired experience of a meal rather than a known ingredient or dish. Its interface searches 60,000 recipes through 5,000 ingredients, 1,000 sensory terms, 400 nutrients, and negations. A cooking model adds 21,000 sensory facts about 4,200 ingredients and 1,300 facts about procedures. As someone types “hearty,” “mushy,” “moist,” “aromatic,” or “no beef,” recipe scraps fade in and out; opening one highlights the ingredients that explain its match. Family avatars carry manually programmed tastes, react to candidates, and explain why. Earlier versions were tried by hundreds of Media Lab visitors, but the poster reports no controlled retrieval or usability evaluation. Its reach is bounded by the parser, keyword profiles, and finite food knowledge.

PDF

The aesthetiscope: visualizing aesthetic readings of text in color space

Hugo Liu & Pattie Maes (2005). Proceedings of the 18th International FLAIRS Conference (AI in Music & Art track), 74–79, AAAI Press.

Aesthetiscope reads a word, poem, or lyric through five modules—Thinking, Feeling, Sensation, Intuition, and Culturalizing—then maps their associations into a glimmering 16-by-9 color field. ConceptNet supplies rational links; popular-culture text cultural ones; stock images remembered colors; free-association norms intuition; and affective vocabularies feeling. Sliders set each reader's contribution. Informal feedback from psychologists, designers, colorists, and hundreds of visitors shaped the installation: grid borders disappeared, source associations became briefly visible, and the image learned to breathe. Visitors seemed to prefer either intuitive/feeling or thinking/sensation mixtures, but personality type remained a hypothesis. This paper contains no controlled experiment; text is mostly a bag of local features, cultural sources are American, sensation is visual, and weights are manual.

PDF

InterestMap: harvesting social network profiles for recommendations

Hugo Liu & Pattie Maes (2005). Proceedings of IUI Beyond Personalization 2005, San Diego, CA, 54–59.

186 citations

InterestMap tries to recommend for the whole person rather than for a single shopping session. It mined 100,000 public social-network profiles, segmented their free-form interest lists, normalized them against 21,000 interest and 1,000 identity descriptors, and learned a weighted network from co-occurrence. Spreading activation moved from a person's stated tastes through identity hubs and cultural cliques to candidate recommendations. In five-fold tests that hid half of each profile, the full network scored 0.86 on the paper's graded percentile measure; removing identity nodes reduced it to 0.81, and direct pairwise tallies scored 0.73. This was held-out profile reconstruction, not a user trial. The data also inherited duplicate accounts, heuristic normalization, public self-presentation, and the blind spots of a fixed ontology.

PDF

Articulation, the letter, and the spirit in the aesthetics of narrative

Hugo Liu (2004). Proceedings of the 2004 ACM Workshop on Story Representation, Mechanism, and Context (SRMC’04), New York.

This essay distinguishes the spirit of a story—personal, affective, mythical meaning—from its letter, the explicit form through which a culture makes that meaning shareable. Articulation moves between them, but perfect articulation can flatten the ambiguity that lets a reader participate. Drawing on literary and cultural theory, the paper describes partial articulation, defamiliarization, re-articulation through myth, and hyperarticulation in a culture saturated with familiar forms. It proposes four ways to preserve aesthetic life: intertextuality, unusual representation, a recognizable aesthetic signature, and personalization to a reader's background or interests. This is a position paper, not an implemented narrative system or experiment. It also admits that intertextuality can exclude readers and that personalization can harden into another cliché.

PDF

sentiment analysis

A model of textual affect sensing using real-world knowledge

Hugo Liu, Henry Lieberman & Ted Selker (2003). Proceedings of the ACM International Conference on Intelligent User Interfaces (IUI 2003), 125–132, ACM Press.

790 citationsACM IUI Impact Award (2018)

A sentence such as “I got fired” carries feeling without naming one. This paper uses everyday knowledge to infer that implicit affect. It filters affective statements from Open Mind Common Sense, grounds them in six Ekman emotions, and combines four models of event structure, concepts, valence, and modifiers. Smoothing represents emotional decay, local interpolation, overall mood, and transitions such as relief. EmpathyBuddy made the result visible as a reacting face beside an email composer. In a twenty-person interface study, the real affect model was judged more intelligent and drew the strongest interest in continued use; randomized faces were slightly more entertaining. The study did not measure classifier precision or recall, and the paper leaves sarcasm, humor, personal history, and long narrative context unresolved.

PDF

A corpus-based approach to finding happiness

Rada Mihalcea & Hugo Liu (2006). Proceedings of the AAAI Spring Symposium on Computational Approaches to Analyzing Weblogs, AAAI Press.

273 citations

This paper treats mood-tagged LiveJournal posts as a form of linguistic ethnography. From 5,000 happy and 5,000 sad entries, a unigram Naive Bayes model reached 79.13 percent accuracy and yielded a 446-word happiness lexicon. The derived words correlated moderately with ANEW pleasure and dominance scores but not arousal. Further analyses found happy words more social and sad words more human-centered, and explored grammatical mode, time, and repeated phrases. Several conclusions are deliberately playful rather than causal. The hourly curve, for example, combines web-search counts for a word plus an hour; it does not use the posts' timestamps. Voluntary mood labels and LiveJournal users are self-selected, first WordNet senses can be wrong, and public posts cannot establish “private” happiness.

PDF

What would they think? a computational model of attitudes

Hugo Liu & Pattie Maes (2004). Proceedings of the ACM International Conference on Intelligent User Interfaces (IUI 2004), 38–45, ACM Press.

60 citations

What Would They Think? converts first-person writing into an affective memory of someone else's attitudes. Salient events become episodic memories; repeated concept-and-feeling pairings become reflexive ones. When a reader opens new material, the system estimates pleasure, arousal, and dominance, renders several modeled people as expressive faces, and lets a click reveal the source quotations behind a reaction. Four long-running webloggers supplied model checks; a separate study asked thirty-six students to learn about unfamiliar writers using raw blogs, a textual model, or the full interface. The full version helped most with explicit attitudes and speed, but implicit attitudes remained difficult. Pleasure and dominance predictions were imprecise, sarcasm and belief revision were unresolved, and the authors reject autonomous fail-hard use.

PDF

Visualizing the affective structure of a text document

Hugo Liu, Ted Selker & Henry Lieberman (2003). Proceedings of CHI 2003, 740–741, ACM Press.

97 citations

Poseidon's Eye asks whether the emotional progression of a story can become a navigable document structure. A commonsense model classifies sentences into six emotions plus neutral; Bayesian smoothing, layout boundaries, and discourse cues gather them into larger regions. A colored bar then links each affective region to its place in the text. Four people used four stories each. Event-location tasks were reported 27 percent faster for familiar stories and 36 percent faster for unfamiliar ones than with a uniformly yellow control. That is an intriguing demonstration, but an extremely small one. The affect condition came first, leaving a practice effect possible; color meanings required learning and could vary by culture; and readers sometimes returned to thematic browsing when emotion alone did not identify the passage.

PDF

semantics of natural language

NLP (natural language processing) for NLP (natural language programming)

Rada Mihalcea, Hugo Liu & Henry Lieberman (2006). Computational Linguistics and Intelligent Text Processing (Ed. A. Gelbukh), LNCS 3878, 319–330, Springer.

161 citations

This paper asks how much program structure can be recovered from ordinary English while preserving language's useful incompleteness. It combines Metafor's mappings with detectors for steps, loops, examples, conditions, passive statements, and assumptions, then emits Perl-like scaffolding. The evaluation began with 120 introductory programming assignments gathered from the web; twenty-five were manually annotated. Step detection reached 86.0 percent precision and 75.4 percent recall, while loop detection reached 80.6 percent precision and 71.4 percent recall. The output is deliberately an outline, not complete runnable software. Its rules depend on surface cues and the conventions of classroom assignments, some potentially executable statements become comments, and the study does not ask whether learners finish programs more successfully with the generated skeletons.

PDF

Metafor: visualizing stories as code

Hugo Liu & Henry Lieberman (2005). Proceedings of the ACM International Conference on Intelligent User Interfaces (IUI 2005), 305–307, ACM Press.

151 citations

Metafor treats a story as an early program model. MontyLingua parses typed English into subject-verb-object structures; ConceptNet helps constrain plausible readings; semantic recognizers update a code model; and a Python renderer continuously shows the emerging scaffolding. Nouns can become objects or classes, verbs functions, adjectives properties, and “when” clauses conditions. Later sentences may refactor earlier interpretations as context accumulates. Thirteen beginner and intermediate programmers estimated the effort required to build Pacman, brainstormed with Metafor, and revised their estimates. Beginners reported a 22 percent reduction and intermediates 11 percent, with greater willingness to use the system than paper. These are confidence judgments, not completed programming tasks; the examiner sometimes rephrased input, coverage was partial, and generated code might not run.

PDF

Programmatic semantics for natural language interfaces

Hugo Liu & Henry Lieberman (2005). Proceedings of CHI 2005, 1597–1600, ACM Press.

76 citations

Rather than reduce instructions to keywords, this paper asks what procedural meaning ordinary grammar already carries. It develops four families of mappings: syntactic, procedural, relational and set-theoretic, and representational. Nouns suggest structures, verbs functions, conditions establish control, plural descriptions imply iteration, and comparative phrases select dynamic sets. Metafor demonstrates how later context can refactor a simple representation into a richer one. Evidence comes from prior novice-programming descriptions, three system iterations, and a thirteen-person Pacman brainstorming study in which beginners especially preferred Metafor. The study measured confidence and estimated time, not completed code or mapping accuracy. An examiner helped enter and sometimes rephrase stories, generated code could remain non-executable, and ambiguous conditions still depended on world and program context.

PDF

Feasibility studies for programming in natural language

Henry Lieberman & Hugo Liu (2006). End-User Development (Eds. Lieberman, Paternò, Wulf), HCI Series Vol. 9, 459–474, Springer.

62 citations

This chapter asks whether a non-programmer can communicate intent without first translating it into formal syntax. Its proposed answer combines broad but partial parsing, mixed-initiative dialogue, and programming by demonstration. The paper closely reads children's plain-English descriptions of Pacman and identifies the real interaction problems a system must survive: top-down and detail-first styles, assumed context, out-of-order instructions, advice rather than actions, missing cases, conflicts, and unannounced changes of viewpoint. A future system would maintain goals and a semantic code model, confirm interpretations, ask about unknown actions, and connect language to direct manipulation. This is a feasibility study and design proposal, not an implemented end-to-end product or quantitative trial. The authors explicitly concede that the imagined dialogue may exceed available language understanding.

PDF

Toward a programmatic semantics of natural language

Hugo Liu & Henry Lieberman (2004). Proceedings of VL/HCC’04: 20th IEEE Symposium on Visual Languages and Human-Centric Computing, 281–282, IEEE.

This two-page position paper catalogs recurrent correspondences between ordinary English and program structure. Nouns can suggest types and objects; possessives act like namespace paths; descriptive phrases select sets; “when” and “if” establish conditions; plurals invite iteration; ordered narration implies steps; and incomplete descriptions can preserve useful delayed commitment. Examples come from children's explanations of Pacman and early Metafor work. The important claim is modest: natural language carries programmatic structure that interfaces can exploit, even when it does not specify a whole program. The paper presents no new corpus, parser-coverage score, user study, or runnable-code evaluation. Many mappings still require common sense, contextual disambiguation, and structural inference, so the taxonomy is a research agenda rather than proof that arbitrary prose compiles.

PDF

Saurus: an emotionally-weighted thesaurus

Jim Gouldstone, Hugo Liu, Henry Lieberman & Hiroshi Ishii (2006). Proceedings of the AAAI-06 Workshop on Computational Aesthetics, 107–110, AAAI Press.

Saurus asks why a thesaurus should treat synonyms as emotionally interchangeable. It augments roughly 35,000 Roget terms with pleasure, arousal, and dominance values. Starting from more than 1,000 rated ANEW words, it follows co-occurrences in Open Mind Common Sense to estimate values for about 70 percent of the thesaurus. A writer supplies text and a short guide phrase; after lemmatization, tagging, and attempted sense selection, Saurus chooses synonyms nearer that tone. Single-word examples can be apt, but full passages often become nonsensical when the wrong sense is chosen. The paper contains no formal accuracy study or user trial. Long passages average toward emotional neutrality, changed word lengths complicate comparison, and synonymy alone cannot guarantee local grammar, discourse fit, or intended meaning.

PDF

Langutils: a natural language toolkit for common lisp

Ian Eslick & Hugo Liu (2005). Proceedings of the International Lisp Conference (ILC 2005), Stanford, CA.

Langutils carries a MontyLingua-derived pipeline into Common Lisp while preserving the language's high-level, inspectable style. It combines contextual tokenization, a rule-based Brill tagger, and shallow phrase chunking with compact integer-token arrays and macros that compile recognizers into specialized code. On a 1.67 GHz PowerBook, the tokenizer processed about 30,000 words per second, tagging reached 15,000–21,000, and chunking 250,000 tokens per second. Example applications include command mapping, ConceptNet extraction, and web mining; the last produced only about 30 percent high-quality phrases. The paper measures throughput, not independent linguistic accuracy. Brill training was absent, token representation was unbounded, multilingual handling incomplete, genre sensitivity substantial, and shallow chunks were a deliberate trade for speed.

PDF

Unpacking meaning from words

Hugo Liu (2003). Modeling and Using Context (Eds. Blackburn et al.), LNCS 2680, 218–232, Springer.

Ordinary dictionaries divide a word into fixed senses; this paper asks how a machine might instead assemble the meaning needed for a particular situation. The Bubble Lexicon places words and concepts in a weighted network of lexical and everyday knowledge, then uses spreading activation to select a context-sensitive neighborhood around a term. In trials over an adapted OMCSNet of roughly 140,000 items, contexts such as money, culture, and transportation changed the readings of phrases including “fast horse,” “cheap apartment,” and “talk music.” The examples show both the promise and the fragility of the method: sparse links and exaggerated context weights could produce arbitrary associations. This was an implemented exploratory model, not a statistical evaluation.

PDF

common sense reasoning

ConceptNet: a practical commonsense reasoning toolkit

Hugo Liu & Push Singh (2004). BT Technology Journal 22(4), 211–226, Kluwer Academic.

2,458 citations

ConceptNet begins with a practical question: can ordinary knowledge gathered from ordinary people become usable infrastructure for software? Roughly fifty extraction rules transformed Open Mind Common Sense sentences into relations; MontyLingua parsed the language, and later normalization merged grammatical variants, reconciled vocabulary, and added thematic links. ConceptNet 2.0 contained about 1.6 million assertions over 300,000 nodes, drawn from some 700,000 sentences contributed by more than 14,000 people. Its toolkit supported contextual neighborhoods, analogy, projection, topic gisting, classification, novel-concept detection, and affect sensing. The paper is equally clear about the unfinished work: coverage was uneven, paraphrases fragmented the evidence, and most assertions appeared only once. Its human evaluation described the earlier 1.2 network, not the new thematic links in 2.0.

PDF

Commonsense reasoning in and over natural language

Hugo Liu & Push Singh (2004). Knowledge-Based Intelligent Information and Engineering Systems (Eds. Negoita, Howlett, Jain), LNCS 3215, 293–306, Springer.

218 citations

Formal logic offers precision, but everyday knowledge arrives in elastic phrases. This paper keeps ConceptNet's common sense close to natural language, then supplies the “glue” needed to reason across different wordings. Concepts are decomposed and located in WordNet, dictionaries, verb classes, and FrameNet; heuristic semantic distance then allows spreading activation, fuzzy inference, and analogy to bridge phrases such as “buy food” and “purchase groceries.” The source corpus held nearly 700,000 Open Mind statements, yielding more than 250,000 knowledge elements across nineteen relation types. The reported similarity percentages are worked examples rather than scores against a gold standard. Ambiguity, duplication, surface parsing, and task-dependent weights remain explicit costs of choosing linguistic flexibility over a complete formal semantics.

PDF

Beating common sense into interactive applications

Henry Lieberman, Hugo Liu, Push Singh & Barbara Barry (2004). AI Magazine 25(4), 63–76, AAAI Press.

188 citations

This article rejects the premise that a machine must possess complete, perfectly reliable common sense before applications can benefit from it. Its alternative is the fail-soft interface agent: inferred suggestions remain optional inside an interaction the user can complete conventionally. Bad suggestions can be ignored; missing knowledge merely makes the agent quiet. The paper surveys ARIA, affective text interfaces, documentary-video assistance, search reformulation, predictive text, and story systems, reporting small prototype studies where available. None was a large deployment, and teams sometimes added knowledge by hand. The contribution is therefore a product principle more than a claim of solved intelligence: design the interaction so partial understanding can help without gaining enough authority to break the task.

PDF

Teaching machines about everyday life

Push Singh, Barbara Barry & Hugo Liu (2004). BT Technology Journal 22(4), 227–240, Kluwer Academic.

101 citations

No single representation comfortably holds all of everyday life. This architectural paper compares three complementary systems: ConceptNet for broad semantic relations, LifeNet for probabilistic transitions between nearby moments, and StoryNet for actors, events, settings, goals, and themes across longer sequences. ConceptNet grew from nearly 700,000 public contributions; LifeNet contained about 80,000 propositions and 415,000 pairwise probability tables. StoryNet, and a browser joining the three, remained under development. The paper's value lies in the division of labor it makes visible: broad graphs are associative but local, probabilistic models tolerate uncertainty but lose expressiveness, and scripts preserve narrative structure but are harder to acquire. It proposes an integration rather than presenting a controlled comparison.

PDF

Goose: a goal-oriented search engine with commonsense

Hugo Liu, Henry Lieberman & Ted Selker (2002). Adaptive Hypermedia and Adaptive Web-Based Systems (Eds. De Bra, Brusilovsky, Conejo), LNCS 2347, 253–263, Springer.

116 citationsBest AI Paper Award · Asociación Española para la Inteligencia Artificial

GOOSE starts from the mismatch between how a novice states a goal and how a search engine expects a query. It parses a natural-language request into a semantic frame, classifies the goal, combines everyday knowledge with search expertise, and emits a reformulated query; raw keywords remain the fallback. Four novice users tried constrained tasks. The system inferred seven of eight household-problem queries, where average relevance rose from Google's 3.5 to 6.1, but only one of eight product-research queries, where it did slightly worse. Those results reveal the system honestly: useful inside familiar domains, brittle outside them. The sample was tiny, categories were chosen manually, trademarked products were largely missing, and the paper explicitly says the prototype was not yet generally robust.

PDF

Makebelieve: using commonsense to generate stories

Hugo Liu & Push Singh (2002). Proceedings of the 18th National Conference on Artificial Intelligence (AAAI 2002), 957–958, AAAI Press.

121 citations

MAKEBELIEVE asks whether fragments of everyday causal knowledge can sustain an interactive story. From 9,000 cause-and-effect statements selected out of Open Mind Common Sense, it builds crude event frames and links one event's effect to another's cause using WordNet and verb classes. A user supplies the opening sentence; the system extends a causal chain, rejects cycles and contradictions, and invites another line when inference stalls. Generated stories ran five to twenty lines. In a preliminary study, eighteen readers gave five-line stories an average combined score of 10 out of 15 for creativity, quality, and coherence. There was no control condition, narrative structure remained mostly local, and ambiguous bindings made stories with several characters unreliable.

PDF

Adaptive linking between text and photos using common sense reasoning

Henry Lieberman & Hugo Liu (2002). Adaptive Hypermedia and Adaptive Web-Based Systems (Eds. De Bra, Brusilovsky, Conejo), LNCS 2347, 2–11, Springer.

108 citations

ARIA watches an email or webpage take shape, continuously reorders a photo collection, and learns when the writer drags an image into the story. Parsing extracts people, places, things, and events; a weighted graph built from Open Mind Common Sense expands annotations across ordinary associations. That lets “bride” retrieve wedding, groom, veil, tuxedo, and wedding dress even when those words never labeled the image. Reinforcing convergent paths and penalizing high fan-out restrains some noise. The paper reports qualitative mutual adaptation—the suggestions changed what users remembered and wrote—but does not isolate commonsense expansion in a controlled trial. Its knowledge was culturally narrow, incomplete, ambiguous, and prone to drift, so the interface deliberately leaves irrelevant suggestions harmless and ignorable.

PDF

Semantic Understanding and Commonsense Reasoning in an Adaptive Photo Agent M.Eng. thesis

Hugo Liu (2002). M.Eng. Thesis, School of EECS, MIT, 160pp.

This master's thesis develops ARIA as an integrated photo-storytelling agent rather than a single retrieval algorithm. WALI turns prose into event structures spanning people, places, times, things, events, and emotions. CRIS compiles about 80,000 Open Mind facts into a weighted network for real-time semantic expansion. PCSL learns personal relations—such as a family connection—from explicit patterns and repeated co-occurrence. Together they let ARIA learn annotations from prose, cross gaps such as bride to wedding, and adapt later searches to its user. The thesis carefully separates demonstration from proof: these new components were not formally user-tested. WALI needed multiple rules for linguistic alternations; the commonsense corpus remained noisy; cold start, collection scale, and controlled comparison were left open.

PDF

Robust photo retrieval using world semantics

Hugo Liu & Henry Lieberman (2002). Proceedings of the LREC 2002 Workshop on Creating and Using Semantics for Information Retrieval and Filtering, Las Palmas, 15–20, LREC Press.

48 citations

This paper gives the commonsense retrieval engine inside ARIA a clear technical account. Pattern rules convert more than 400,000 Open Mind sentences into 50,000 predicate structures; normalization produces a graph of 30,000 concepts and 160,000 weighted edges. Spreading activation expands a query, rewards associations supported by several paths, and discounts overly general nodes. “Bride,” for example, reaches wedding, groom, church, flower girl, veil, and wedding dress—but also lake, mountain, and other noise. That mixture defines the boundary. The mechanism had no formal evaluation, reflected a predominantly middle-class American corpus, and did not distinguish word senses. It was intended for ordinary consumer photographs and remained useful because ARIA allowed the writer to ignore bad suggestions.

PDF

MontyLingua: an end-to-end natural language processor with common sense software / tech report

Hugo Liu (2004). MIT Media Laboratory, Cambridge, MA.

87 citations

MontyLingua packages the everyday operations of English processing into one inspectable path: sentence splitting, tokenization, part-of-speech tagging, chunking, lemmatization, semantic extraction, and later surface generation. Its jist interface emits Lisp-shaped verb-subject-object structures that applications can use without assembling a research stack of their own. The original 2.1 distribution survives with Python and Java source, documentation, data files, release archives, and licenses. Its notes report 97 percent word-level tagging accuracy after adding common sense, plus substantial speedups, but do not include the protocol needed to reproduce those figures. Subordinate-clause objects were excluded, generation was beta, and the Python 2-era code depends on companion data files. It is best understood as working software with candid edges, not a benchmark paper.

Project

later work

What social media sentiment tells us about the ebb and flow of a city's moods book chapter

Nancy Etcoff & Hugo Liu (2017). In A. Karandinou (Ed.): Data and Senses — Architecture, Neuroscience and the Digital Worlds, University of East London.

The Positivity Pulse asks whether time-stamped, geolocated public language can complement surveys of urban mood. It collected 10,000 Santa Monica tweets and 590,000 Greater Los Angeles tweets over six days. LIWC positive and negative word counts formed a mood ratio; first-person plural and singular pronouns formed a proposed measure of social engagement. Against 1,000 tweets labeled by one judge, the classifier reported 76 percent accuracy, 91 percent precision, and 61 percent recall. The paper finds a strong relationship between its We/I and positive/negative ratios and describes different hourly patterns for residents, visitors, and Greater LA. Its boundaries are material: six days, one annotator, proxy location labels, speculative explanations for the peaks, and unresolved privacy questions.

PDF

Brand choice as gender identity workshop paper

Elizabeth F. Churchill & Hugo Liu (2014). Proceedings of the CHI 2014 Workshop: Perspectives on Gender and Product Design, ACM Press.

This workshop paper treats brand choice as a possible form of gender expression rather than assuming product design should simply mirror a profile's male/female field. Using 2013 eBay purchases, it assigns each of 1,500 brands a femininity percentile from its reported buyer mix, then averages the distinct brands bought by each of 608,000 accounts across shoes, watches, and small kitchen appliances. Female-labeled accounts formed a near-normal distribution around 0.54; male-labeled accounts had a more masculine mode of 0.34 and a long feminine tail. The interpretation is deliberately tentative. Reported gender was unverified and menu-order biased; accounts could be shared, used for gifts, or represent households. Purchases cannot establish one person's identity, much less causation.

PDF
← back home