One word in fourteen is "the"

In the 1960s, researchers at Brown University assembled a million words of ordinary American English: newspapers, novels, science writing, all published in 1961. When they counted every word, the winner was not close. The appeared 69,971 times, about 7% of everything. In second place, of appeared 36,411 times, almost exactly half as often. Third was and, with 28,852.

Keep going down the list and the pattern holds with eerie regularity: the word in position n turns up roughly 1/n as often as the top word. Just 135 different words make up half of that entire million-word collection. The other half is shared among tens of thousands of words, most of which appear only a handful of times.

Zipf's first law

A French stenographer, Jean-Baptiste Estoup, noticed this in 1916. The American linguist George Kingsley Zipf made it famous in the 1930s and 40s, and it now carries his name. Zipf's law says that if you rank words by how often they appear, frequency times rank stays roughly constant. Word 1 appears F times, word 2 about F/2, word 10 about F/10, word 1,000 about F/1,000.

The trick for seeing it is to plot both rank and frequency on logarithmic scales, where every step is ten times bigger than the last. A 1/rank curve then becomes a straight line sloping down at 45°. Real text wobbles around that line, but the resemblance is hard to miss. Try it below: the default text is this very article.

a word (drag or tap to inspect)Zipf's 1/rank line
0words in total
0distinct words make up half the text
0 / 0avg letters: top 10 words vs words used once
Each dot is one distinct word. Left means common, right means rare. The red line is what a perfect 1/rank law would predict from the top word. Longer texts hug the line more closely; a short paragraph is too small to show much.

Zipf's second law: busy words get short

Look at the top of the table above. The, of, and, a, to, in, is: tiny words, two or three letters each. Now look at the words that appear only once. They are much longer. That is Zipf's other, less famous law, the law of abbreviation: the more often a word is used, the shorter it tends to be.

This one is not a quirk of English. It has been checked in close to a thousand languages from 80 different language families, and it keeps turning up. There is evidence for it in the calls of some other primates, too.

You can watch it happen in real time. When a long word gets used a lot, speakers wear it down: omnibus became bus, telephone became phone, examination became exam. Nobody voted on those changes. Saying the long form thousands of times a year is simply more effort than anyone wants to spend.

Laziness, or something smarter?

Zipf's own explanation was a "principle of least effort". Speakers want to say as little as possible; listeners want words that are distinct enough to understand. A language that settles between those two pressures ends up with a few short, overworked words and a long tail of precise, rare ones.

But frequency may not be the whole story. In 2011, Steven Piantadosi, Harry Tily and Edward Gibson at MIT measured word lengths in 11 languages. They found that a word's length is better predicted by how predictable it is from the words just before it than by how often it appears overall. Words that are easy to guess in context get shorter, because the listener needs less signal to catch them. That is the same logic behind a well-designed code: spend the fewest symbols on what's easy to guess.

The monkey problem

Here is the twist that keeps linguists humble. In 1957 the psychologist George A. Miller pointed out that you get a Zipf-like curve even from completely random typing. Press the Monkey at a keyboard button above: it hits letters and the space bar at random. Short "words" come out more often than long ones simply because there are fewer ways to make them, and the rank-frequency plot still falls away as a staircase of steps, one step per word length, whose overall trend is a straight-ish line on the log-log plot. Notice, too, that the monkey's common "words" are its shortest ones.

So a straight line on a log-log plot doesn't, by itself, prove that language is cleverly optimised. What the monkey can't do is make its common words mean anything, or wear examination down into exam. The debate over why Zipf's law appears is still open, and similar power laws turn up in city sizes, incomes and website traffic.

The takeaway

Every language is lopsided in the same way. A handful of short, hard-working words carry half the load, and a vast tail of long, specific words appears only now and then. Next time you write an email, you will be using those same few words over and over, and they will be pulling their weight exactly as Zipf predicted.