Free data downloads

Start building with Linguabase’s comprehensive map of English vocabulary and its connections.

Relative distribution of words by letter count
135791113151719
Single words vs. words with spaces, by familiarity
50K100K150K200K250K300K350K400K
Single words Words with spaces

This language atlas has thirteen tables of data. Before you even get into this, you need to know that there’s no exact count of how many words are in English. If you want to understand why, see how many words are in English? — but the heart of the matter is that words range from the most common and foundational words (the, and, of), to the meat of English (sister, forgot, amazing), to obscure English for wordsmiths (incongruences, extemporization), to technical terms (vanishing fraction), and then it extends out, ad infinitum, with proper nouns, uncommon spelling variants, and eventually blurs into endless specialized word combinations (lighthouse keeper), noise and typos that appear in online documents.

In our experience, an arbitrary threshold of the top 400,000 words is a really solid pool of legitimate words, but we provide larger headword sets, often expanding out to ~2 million headwords, to power various use cases or provide post-processing signals. This is all wired together by 250 million+ weighted relationships and annotated with senses, antonyms, word families, definitions, clues, and per-word scores and labels.

The data is available for download from three mirrors, all carrying the same files:

Hugging Face →

Load straight into Python with datasets or pandas. Parquet, with a dataset viewer for browsing before you download.

Parquet + dataset viewer

Zenodo →

A versioned, citable archive with a permanent DOI — for research and long-term reference.

Permanent DOI

Kaggle →

Explore and prototype in hosted notebooks, or pull the files into your own pipeline.

Hosted notebooks

13 tables · 22 core/full subsets · ~29.4M rows · 6.22 GB · CC0 public domain.

For citation, use the Zenodo DOI: 10.5281/zenodo.20788428.


What kind of data is here

The data comes in three parts — words, links, and meanings — thirteen tables in all. The words are the ranked vocabulary; the links wire them together; the meanings annotate them. It’s built for anyone working with English at scale: game designers, app builders, researchers, lexicographers, artists.

Words

400K terms ranked by familiarity—from everyday vocabulary to crossword-worthy rarities. Includes 200K multi-word expressions. Filterable by letter count or difficulty for your audience.

Content filters at two severity levels: a hard-block list of offensive words, and a soft-block list of words carrying unwanted innuendo.

Links

Over 250 million weighted relationships, including both categorical associations (soft → gentle, fuzzy, velvet) and associative relations (bunny → soft, fuzzy, rodent).

Every word broken down into senses and facets, with ~100 related words per entry on average. This is the layer that powers mutual exclusivity. Word families connect related forms (morphology): run → runs, running, runner, runway.

Meanings

400K definitions as ~55-word readable paragraphs covering all senses. Short clues — 1 to 5 words, multiple angles per term.

1.5 million usage examples from literature and journalism. Metadata that makes each word usable, not just present in a list.

How it’s better than anything else

Free word lists are easy to find. What sets this apart is that the relationships are engineered rather than scraped — balanced across senses, cleaned of the artifacts that wreck a naïve association dump, and graded so you can trust the strong links.

Half of it is words with spaces — 200,000 multi-word expressions like “night sky” and “old wives’ tale” that name a concept of their own. Ordinary dictionaries cover about 3% of them; here they’re first-class terms, doubling the pool of ideas to work with.

And every word is split by sense, so elephant keeps its anatomy, circus, heraldry, and symbolism meanings apart instead of collapsing to one blurred average. That’s the gap a language model can’t close: asked for associations, it leans on the two or three dominant senses and misses the rest.

Under the hood, the cleanup that makes those links trustworthy:

291K false cognates removed dig → digress, pan → pandemic
Capitalization intelligence “polish” vs. “Polish” get separate lists
Superconnector demotion Common hubs penalized so rare links surface
Experiential associations (gestalt) crisis → siren, wedding → white
Function word coverage “and,” “while,” “but”—not skipped

Side by side: vs. the free alternatives / vs. an LLM / vs. a traditional dictionary.

How the data is set up

Concretely, the atlas is thirteen tables — and they nearly all line up on the headword as the join key.

Table Key What it holds Scale
Words, rankings & labels
vocabulary
live ↗
word English terms ranked by familiarity (0–1), with a single-word / multi-word flag and morphology type. Core is a curated set of 400,000 words — not simply the top 400K by rank: about a quarter are included or excluded by other criteria. Past the full ranking’s 400K mark, words simply get rarer, and well past ~1.5M the mix shifts toward compounds and technical terms of little use outside research or error detection. ~2M ranked
~400K core
scores & labels
live ↗
word Per-word signals: iconability (how well it draws as an icon), game-suitability (1–7), and a 13-dimension accessibility profile — americana, pop-culture, sensitivity, religiosity, kid-appropriate, register, concreteness … — plus a G–X content rating. 1.3M
364K core
How words relate
associations
live ↗
word ~100 weighted, quality-graded (A–E) associated words per term, balanced across senses, with frequency-inflated “superconnectors” demoted so distinctive links surface. 1.92M words
~127M edges
senses
live ↗
word The sense inventory: each word mapped to its distinct meanings (bridge → anatomy, cards, computing, connection) — the label layer that says which senses exist. 1.56M
201K core
sense associations
live ↗
word · sense A separate association cloud per sense (bridge_cards, bridge_computing…) — the layer that lets a clue aim at one meaning instead of blurring across them. 6.13M
3.08M core
word families
live ↗
word Morphological and etymological relatives, split into substitutable variants (run / running) vs. semantic kin (run / runway). 291K false cognates audited out. 1.46M
359K core
Opposites
antonyms
live ↗
word · axis Multi-axis opposites per word — a word can be opposed along several dimensions (light → dark / heavy / serious) — each opposite rated 1–7. 4.14M
1.6M core
category antonyms
live ↗
category · axis Opposition at the category level — a category set against its opposite, with a member-word pool for each side. 197K
opposite pairs
live ↗
pair Curated symmetric opposite sets — two facing word pools (miniature ↔ giant), scored for how cleanly they separate. 163K
How words read
definitions
live ↗
word Readable 45–65-word paragraphs covering all senses of a word in plain prose — written for display, not dictionary lookup. Independently authored. 1.88M avail.
364K core
clues
live ↗
word Short 1–8-word hints from 26 different angles — synonym, function, category, contrast, figurative, example, plus learner and expert tiers — none sharing morphology with the answer. 1.5M words
× 26 types
Ready-made topics
categories
live ↗
category Curated topic→member-word sets (martial arts → karate, judo, aikido…), each with its own classification — scope, difficulty, sensitivity. Each set begins with a theme — a free-form label like “70s Rock” or “accountability” — and gathers the words that belong to it. 72.8K
topics
live ↗
headword · sense Iconic-headword idea sets built for grouping games — a short headword + sense + member words + a rich audience profile. Each set begins with one word in one sense and gathers what it brings to mind (courthouse → marble, flag, rotunda) — where categories begins with the theme, topics begins with the word. 373K

License

Everything in the release is public domain under CC0 1.0 — use it for anything, commercial work included. Ship it, remix it, sell what you build on it.

The word relationships are facts about language that no one can copyright anyway; the definitions and clues are our own, released the same way. The only thing we hold back is the literary quotes — that’s other people’s writing.

Using it

The files ship as plain TSV and Parquet — UTF-8, one record per line, joinable on word: readable by eye, loadable into a dataframe, or queried directly with DuckDB, no server. A companion guide covers loading, joining, filtering by familiarity or content rating, and feeding the data to your own code or an LLM agent. (Coming with the upload.)

Why free

Linguabase comes from IDEA.org, a charity, and giving the data away is the mission: we want to promote new projects built on semantic data, and help a new generation of data-thirsty word games come onto the market. If it powers your game, your research, your classroom tool, or something we haven’t imagined yet, it’s doing its job. How it was built →  ·  About us →