No one knows.
Count a dictionary’s entries and you get a number: around 600,000 word forms in the Oxford English Dictionary, 470,000 in Webster’s Third. Start counting what dictionaries leave out — technical vocabulary (every chemical compound, every species), proper nouns, slang, brand names, fresh coinages — and the total runs past two million and keeps climbing. There is no answer because there is no fixed thing to count.
A fluent adult knows roughly 20,000 to 40,000 word families.The trouble starts with “word” itself. Is run, runs, running, runner one word or four? A fluent adult knows roughly 20,000 to 40,000 word families — “liberates” and “liberator” counted as one. Depending on what you fold, the same Webster’s Third is 470,000 entries or 54,000 families.
Then there’s the fuzzy borderland between a word and a phrase. “Ice cream,” “hot dog,” “comfort food,” “hold it together” — each is a single unit of meaning that behaves like a word, yet wears a space. We call these words with spaces, and there are hundreds of thousands of them; dictionaries cover a small fraction.
Linguabase doesn’t decide what counts as a word. It takes everything — single words and words with spaces, the everyday and the obscure — and scores each for familiarity.
By construction, the bottom is invalid English — typos, OCR errors, noise. Familiarity fades through a gray zone where “is this even a word?” has no answer. Even dictionaries get caught in it: “dord” slipped into Webster’s in 1934 as a misreading of “D or d” — an abbreviation for density — and went five years before anyone noticed. A ranking can live in that gray; a fixed list can’t. Far enough down, the data flips purpose: terms to recognize and screen out.
A familiarity ranking works because the core of English barely moves. Shakespeare, a children’s picture book, a sitcom, a sci-fi novel — wildly different writing, almost the same handful of high-frequency words doing most of the work. Linguists call the pattern Zipf’s law: “the” alone is about 7% of nearly any English text, and the hundred most common words cover roughly half. Where they part ways is the long tail, and especially the words each one uses exactly once. Switch the view:
Counts are tallied from the full text of each source. The cell marked sample uses stand-in data until that text is added.
The once-only words are hapax legomena: in almost any sizable text, nearly half the distinct words appear exactly once. Those found in no reference book hint at what one Harvard study called “lexical dark matter” — by its estimate, half the real lexicon is in no dictionary.
The singletons are the tell: as long as half a text’s vocabulary appears just once, the well is nowhere near empty. Statisticians use that singleton rate to estimate the vocabulary a text never shows, and the discovery curve never flattens. The vocabulary of English isn’t a number awaiting a careful count — it’s an extrapolation. You can’t finish counting a moving target. You can rank it.
We lean on one threshold ourselves: about 400,000. We set it by reading down the ranked list: around 350,000 still left out words we wanted; around 450,000 began letting in strings that aren’t really English; 400,000 is where the trade-off settled. We drew that line by hand — English itself has none.
Any robust project can draw everything it needs from there: broad enough never to feel thin, shallow enough to stay clean. It’s a default, not a limit — a casual mobile game might live in the top 10,000; a crossword or a research corpus reaches much further down. Pick the slice your project needs. Get the data → · Why we keep two million words internally →
One more sign the word was never a clean unit: large language models — the systems that have read more English than anything in history — don’t use words at all. They split text into statistical fragments: “unhappiness” becomes un + happi + ness. The field that most needed to count words gave up on the word.
← More interesting questions