Linguabase is a structured map of English—two million ranked words and phrases wired together by 250 million+ weighted semantic relationships, with senses, antonyms, word families, definitions, and clues—built over fifteen years and released free, in the public domain. It powers word games especially well (the association and categorization genre that took off after NYT Connections), and it’s open to any use. Linguabase is a product of IDEA.org.
Founder & AI Systems Architect
Michael led the project from its origins as internal infrastructure for two word games through to the current system—the LLM pipeline, the production data, and the release of the entire dataset into the public domain. When you email linguabase@idea.org, you’re talking to the person who designed every layer of the data.
The data accumulated over fifteen years across three phases. Professional lexicography and hand-built vocabulary came first—thousands of definitions, sense-grouped associations, and thematic word lists authored by people who understand English at a level automation can’t reach. Computational linguistics came next: 70+ structured linguistic sources integrated algorithmically, 648,000 Library of Congress subject classifications used as idea seeds, 2.3 million supercomputer hours of topic modeling and word embeddings via an NSF grant, and the algorithms that map and weight relationships across the full vocabulary. The current system layers 130 million LLM inferences on top of that foundation—generating, validating, ranking, and auditing at a scale the earlier phases couldn’t. Read the full story →
The deployed vocabulary is larger than the Scrabble dictionary, Merriam-Webster Collegiate, or Collins, because it includes familiar multi-word expressions other dictionaries don’t. Roughly the size of Webster’s Third Unabridged, excluding the obscure and technical words that aren’t fun. (There’s no true count of English — here’s why.)
The data layers built before LLMs existed—algorithms, lexicography, and source integration that the current system inherits and builds on.
Language Data Architect
Designed the algorithms that map and weight word relationships across the full vocabulary. Mathematics background (Shandong University), decades of software engineering.
Lexicographer
Established the lexicographic framework for how word relationships should be structured—which senses matter, how to group associations, where human judgment is non-negotiable. Wrote 2,000+ custom definitions and 4,400+ sense-grouped associations. Contributor to Oxford, Macmillan, and other major dictionaries.
Sally Smith manually curated 100 sets of mutually exclusive topics for OtherWordly—a manual process that became the precursor to the current automated pipeline.
The following contributed as content writers and reviewers many as linguistics grad students or post-docs, typically working on a block of topics from the Dewey Decimal or Library of Congress classification system—sports played with balls, cathedral architectural elements, newspaper brand names (Globe, Post, Tribune), soft candies.
Linguabase was supported by a $295,000 NSF SBIR Phase I grant and $300,000 in Microsoft for Startups compute credits.
Share your work (we’d love to see what you’re doing), let us know about errors in the site or the data, or engage Michael for implementation consulting.