What is a confound?

When a word fits more than one match in your game.

Arcimboldo’s Vertumnus: a human portrait composed entirely of fruit, vegetables, and flowers

A face and a heap of produce at once — neither reading is wrong. That doubleness is the whole problem.

Giuseppe Arcimboldo, Vertumnus, c. 1590. Public domain.

Most word games are exercises in resolving ambiguity — narrowing the field of matches by meaning or by letters. And most words carry several kinds of valid association: bunnies are fuzzy and long-eared; clouds are pillowy, white vapor. Words with more than one meaning sharpen the problem — a key opens a lock, is a coral island, and is a song’s signature; a ball is round, but also a party.

Whether it’s different facets of one meaning or outright polysemy, a confound lives in two unrelated worlds at once. Whether that helps or hurts comes down to control:

The term comes from experimental design, where a confound is a variable you failed to isolate, so an effect can’t be pinned to one cause. You can use Linguabase either way — to plant more confounds on purpose, or to root them out of your puzzles.

Detecting confounds using Linguabase

WordRankTopic ATopic B
ball105round things balloon, bubble, discSoccer dribble, field, foul
window147IT vocabulary click, copy, displayhouse attic, basement, ceiling
light123physics atom, energy, gravitysenses see, hear, touch
bridge198structures arch, cable, spancard games deal, trump, bid
rain112weather cloud, hail, lightningnatural resources coal, gas, fish

How Linguabase finds them. Across 640 curated categories we build a word → categories index, keep common single words (rank ~90–20K), and flag any that land in two categories whose labels share no words and whose member sets barely overlap — truly different domains, not near-synonyms. Ball in both “round things” and “Soccer” is the textbook confound: drop it on a board and a solver can’t isolate which group it’s for. The data sees it because it knows every sense of a word, including the one no player is thinking of.

What about red herrings, distractors, and naticks?

Puzzlemakers have more jargon for variations of disambiguation. We can differentiate four related types of obstacle:

ObstacleIs it in the group?What it doesFair?
ConfoundIn two at onceYou can’t tell which group it’s forThe umbrella
Red herringIn it — but misleadingPoints you toward the wrong groupFair
DistractorNot in itStands in as a believable wrong pickFair
NatickYou can’t tellLeaves you guessing blindUnfair
1

Red herring

it’s in the group, but everything about it points the wrong way

Active misdirection: a clue or word that aims you at a specific wrong answer. It’s fair — a careful solver can resolve it, because the truth was visible all along. In puzzle hunts and escape rooms it’s the standard word for a deliberately planted false trail, usually a whole path rather than one wrong option. (They can be accidental, too, so “always deliberate” overstates it.)

Popularized by the English polemicist William Cobbett in 1807: writing about a press misled by false reports of Napoleon’s defeat, he described dragging a pungent smoked herring to throw hounds off a hare’s scent. Cured herring turns reddish, and pungent.

In the data — a word whose top links bait you toward a different group

desktop∈ computersbaits devicesvia mouse, monitor, keyboard, laptop
bathroom∈ homebaits fixturesvia toilet, sink, faucet, soap, shower
NASA∈ aerospacebaits spaceflightvia astronaut, satellite, orbit, shuttle

How Linguabase finds it. For a word in a category, pull its top-weighted associations and ask whether four or more of its top twelve links belong to some other, label-disjoint category. If so, its strongest pulls point away from home. Desktop sits in “computers,” but mouse, monitor, and keyboard all describe the devices around the machine. This is measurable only because the edges are weighted — the data knows which way a solver gets pulled, not just that a link exists.

2

Distractor

a believable wrong answer standing beside the right one

The plain word for a tempting wrong option — but it’s borrowed. In a multiple-choice item the stem is the question, the key is the right answer, and the distractors are the plausible wrong choices around it. A good one is built on a common misconception, kept parallel in form to the key, and is fair: you beat it by knowing the content.

A genuine term of art — but in test design, not puzzles. It first surfaces in 1951, in Robert L. Ebel’s chapter “Writing the Test Item” (the multiple-choice format itself is some 35 years older). The first rule of writing them is blunt: “make all distractors plausible.” One chosen by under ~5% of test-takers is “non-functional”; in practice most questions carry fewer than two that actually work.

In crosswords the native craft isn’t called “distractor” — it’s misdirection, and by Will Shortz’s own account it was nearly absent from the Times until he took over in 1993. The devices are precise:

Misdirection in the wild — the clue reads one way, the answer is another

“Leaves from a club”LETTUCE a noun, not the verb “leaves”
“Santa Fe and Tucson, e.g.”SUV the car models, not the cities
“Ticker tape?”ECG the “?” flags wordplay: ticker = heart

The deep cut is Jeremiah Farrell’s 1996 election-day puzzle, where “Lead story in tomorrow’s newspaper” solved as either CLINTON or BOBDOLE — every crossing engineered to work both ways (BAT or CAT, LUI or OUI). Shortz has called it his favorite crossword ever. Linguabase’s part: a fair distractor isn’t a flaw in the data, it’s one of the genuine associative neighbors a theme summons — soar for rising, gaze for perception — handed to a constructor ranked by strength; what makes it a distractor is only where it’s placed.

3

Natick

you can’t tell if it belongs — a pure guess

A crossword crossing where both entries are obscure, so the shared letter can’t be deduced, only guessed. It’s the only one here that’s intrinsically a complaint — a natick is by definition unresolvable by reasoning — and it’s solver-relative: if you happen to know one of the two words, it isn’t a natick for you. It’s now both noun and verb (“I got naticked”).

Coined by Michael Sharp — the blogger “Rex Parker” — on July 6, 2008, after a Brendan Emmett Quigley Sunday puzzle crossed NATICK (“town at the eighth mile of the Boston Marathon”) with N. C. WYETH (“Treasure Island illustrator”) at a shared “N” nothing in the grid let you recover. His Natick Principle: a proper noun more than three-quarters of solvers won’t know must cross only common words.

In the data — genuinely unknowable deep-tail words

WordRankWhy it’s unfair
hypernym378,036linguistics jargon
cardioverter364,545medical-device term
paramountcy356,877rare legal / archaic noun

How Linguabase finds it. The crossing is the grid’s business, not the data’s — but Linguabase owns the ingredient the Natick principle rests on: a familiarity rank across the full 400K vocabulary. Take the deep tail (rank beyond ~340K) and almost no one can infer the words. Cross two and the shared square is a pure guess: ioniser (352,260) over deveiner (349,436), sharing an I no one can recover.

And the vocabulary runs deeper. Puzzle-makers also reach for misdirection (the umbrella craft term, borrowed from stage magic), a trap (a tempting wrong entry), a decoy or foil (a near-miss planted to mislead), crosswordese (obscure little words that survive in grids only for their convenient letters), and the Schrödinger clue (engineered so two answers are correct at once). Each marks a different way a clue or a word pushes the solver — or bends the reasoning they use to push back.

How they relate

Confound is the parent. Under it, a red herring pushes you the wrong way, a distractor crowds you with a credible neighbor, and a natick gives you nothing to push against at all. The first two are fair — deliberate craft, the way a puzzle is supposed to be hard, and a careful solver wins. The natick is the one that names unfairness. Direction and fairness: that’s the whole taxonomy.

They’re not mutually exclusive; they often describe the same element from different seats. A single ambiguous word in a grouping game can be a red herring (it misdirects), a confound (it blocks isolation), and have the flavor of a distractor (it’s a parallel option) all at once — which word you reach for depends on what you want to say about it.

Why Linguabase can see these

The reason we can find these and not just define them is that the whole language sits in one place as a single weighted, sense-split semantic graph. A confound is a word in two distant groups — visible because category membership and member overlap are both in the data. A red herring is a word whose strongest links point at a foreign group — visible because every association carries a weight, so “top twelve” is a real, sortable thing. A distractor is one of a theme’s genuine associative neighbors — visible because the links run both ways and carry weight, so a set’s strongest neighbors can be ranked. A natick is a word deep in the ranking with almost no one able to recognize it — visible because every word carries a familiarity rank.

That’s also why you can build a board that’s hard on purpose and clean by construction. Our board builder uses exactly these signals to plant fair red herrings and distractors while screening out the confounds that would make a grid unsolvable — the unfair, natick-shaped ones. The measurements come straight from the data.

Where this lives in the data

These signals ride on the core tables: the weighted association graph, the curated category sets, and the familiarity rank that runs the full 400K-word vocabulary. None of it is a special “confounds” column — it’s what falls out when the graph is one consistent object. See real rows from every table, try the board builder, or just take the whole thing — it’s public domain.