Almost everyone has had the experience: you overhear a language you don't speak and it seems to tumble out at breakneck speed. Spanish and Japanese have a particular reputation for this. Others, like Vietnamese or Thai, can sound slower and more measured.
The speed is real. Some languages really do push out more syllables every second. But a syllable is not a fixed unit of meaning. In 2019 a team including Christophe Coupé and François Pellegrino asked a sharper question: how much information is actually getting through per second? Their answer, published in Science Advances, was surprisingly tidy.
The researchers recorded 170 native speakers of 17 languages, from 9 language families, each reading the same 15 short texts about everyday situations. That gave them roughly 240,000 syllables to time. For each language they measured two things.
Speech rate is easy: how many syllables come out per second. Across all speakers it averaged 6.63 syllables per second.
Information density is subtler. It asks how surprising the next syllable is, given the one before it, measured in bits. One bit is a coin flip between two options; 5 bits is like picking one of 32 equally likely syllables; 8 bits, one of 256. A language with only a few hundred possible syllables, such as Japanese, carries fewer bits in each one. English, with close to 7,000 distinct syllables, has far more room to surprise you. Across the sample, density ran from 4.8 bits per syllable in Basque to 8.0 in Vietnamese.
Multiply the two and you get an information rate:
Try it yourself. Drag the dot anywhere on the plane below.
When the team did that multiplication, the spread of the fast-and-slow languages mostly collapsed. The information rate centred on 39.15 bits per second, with a standard deviation of about 5 bits per second. Languages that spoke quickly tended to use low-density syllables; languages that spoke slowly tended to pack more into each one. Statistically, the higher a language's density, the slower its speakers talked.
That is exactly the trade-off you can see on the chart. Slide the dot to the right, into rapid speech, and you have to drop down to lighter syllables to stay near the curve. Move it up to dense syllables and it has to slide left. The corners are empty for a reason: fast and dense would swamp a listener, and slow and light would waste their time.
Nobody designed this. No language committee decided to aim for 39 bits a second. The favoured explanation is that the bottleneck is not the tongue but the brain. Speakers must plan and produce speech, and listeners must decode it in real time; both have limits. Over generations, a language that asks too much of its listeners, or too little, drifts back towards a comfortable middle.
Think of it like a pipe carrying water. You can push a thin stream very fast or a thick stream slowly, but what comes out the far end is limited by the narrowest point. For spoken language, the narrow point seems to be the human processing on either side of the conversation.
It is worth being careful about the word. The bits in this study measure how predictable the sound stream is, syllable by syllable, not how much meaning or nuance a sentence holds. A syllable that could have been one of many different ones counts as more informative, whether or not it changes what the speaker means. As the cognitive scientist Sean Trott points out in a review of the work, uncertainty over signals is not the same as uncertainty over meanings.
The recordings were also of people reading prepared texts aloud, not chatting freely, and the 17 languages are mostly from Europe and Asia. So the 39 bits per second is best treated as a striking pattern in a well-controlled sample rather than a universal constant. It builds on an earlier, smaller study by Pellegrino, Coupé and Egidio Marsico in 2011 that spotted the same speed-for-density swap.
Still, the next time a language seems to be racing past you, it probably isn't getting ahead. It is just sending its message in smaller packets.