Caveat lector — built quickly with a guild of magic elves. If something is wrong, please tell me. — Aguas Infinitas, Star
the record · 2005 · semantics of natural language

Langutils: a natural language toolkit for common lisp

Ian Eslick & Hugo Liu (2005). Proceedings of the International Lisp Conference (ILC 2005), Stanford, CA.

PDF

Langutils carries a MontyLingua-derived pipeline into Common Lisp while preserving the language's high-level, inspectable style. It combines contextual tokenization, a rule-based Brill tagger, and shallow phrase chunking with compact integer-token arrays and macros that compile recognizers into specialized code. On a 1.67 GHz PowerBook, the tokenizer processed about 30,000 words per second, tagging reached 15,000–21,000, and chunking 250,000 tokens per second. Example applications include command mapping, ConceptNet extraction, and web mining; the last produced only about 30 percent high-quality phrases. The paper measures throughput, not independent linguistic accuracy. Brill training was absent, token representation was unbounded, multilingual handling incomplete, genre sensitivity substantial, and shallow chunks were a deliberate trade for speed.

← the record, in the papers room