Take any long book in English, count every word, and sort the words from most to least common. The list always looks the same. The top word is far ahead of the rest, a few more words are common, and then a very long tail of words appears once or twice. It is true for novels, news, and school essays. Almost nobody notices, because nobody counts.

The pattern: rank times frequency stays roughly constant

In the 1930s and 1940s the Harvard linguist George Kingsley Zipf worked this out in detail, and his 1949 book Human Behavior and the Principle of Least Effort made it famous. The rough rule now carries his name: the most common word appears about twice as often as the second, three times as often as the third, and so on. The word at rank n appears about 1/n as often as the top word.

The numbers in a real corpus are close to that. In the Brown Corpus, a million-word sample of American English, the top word “the” makes up nearly 7% of all words (69,971 of slightly over a million). The runner-up “of” is slightly over 3.5% (36,411), and “and” follows with 28,852. Not exactly 1, 1/2, 1/3, but not far off either.

Try it on your own text

The box below holds a short passage. Press Count and the chart plots each word's rank against its count, both on logarithmic axes. If Zipf's rule holds, the dots fall along a straight line sloping down at roughly 45°. The dashed line is the ideal 1/n shape. Paste in anything you like. A short text is noisy, but the slope shows up quickly.

A monkey gets there too

Here is the odd part. You do not need a language, a vocabulary, or any meaning to get this shape. In 1957 the psychologist George Miller pointed out that random sequences of letters and spaces already give a Zipf-like word list. Imagine a monkey hitting a keyboard of 26 letters and a space bar, each key chosen at random. A “word” is whatever sits between two spaces. Short words such as single letters are few in kind but each one is likely to come up often. Long words are numerous but each one is rare. Sorted by frequency, the result slopes down much like real text. The mathematician Wentian Li showed in 1992 that this is a by-product of ranking words by length and frequency.

Use the monkey below. It presses 25,000 random keys, and you control how often it hits the space bar. Compare the dots with the line from real text above.

So what is going on?

The monkey result does not mean human language is random. It means that a straight line on this chart is a weak test of how language works. Real text has some features a monkey cannot reproduce. In the monkey's text, all words of the same length are equally likely, so the chart shows steps, with a block of tied words at each length. Real words do not tie like that, and real texts have grammar, topics, and repeated names that a random typist has no way to produce.

Zipf himself explained the pattern with what he called the principle of least effort. Speakers want short, easy words that they can reuse; listeners want words that are distinct and unambiguous. The compromise gives a few cheap, heavily used words and a huge tail of specialised ones. Other researchers have since argued over which explanation is right. Both sides can reproduce the pattern, which is exactly why it is so hard to settle.

One consequence is easy to see in the table above: a small handful of short words (“the”, “and”, “a”) account for a large share of every page, while most different words appear only once. However much you read, new rare words keep turning up.

Zipf-like curves also show up outside words. People have found rough versions in city sizes, company sizes, and incomes, though more recent studies have challenged the neat rule for cities. A straight line on a log-log chart is a good reason to look closer, but it does not tell you why.