Translate the same passage into two languages and the resulting texts may use very different numbers of words. One language might place tense, number, direction or social information inside a word where another uses several separate words. Meaning has not disappeared, but its packaging has changed.

Pedro Aceves of Johns Hopkins University and James A. Evans of the University of Chicago asked whether that packaging is associated with the way communication unfolds. Their study, published in Nature Human Behaviour in February 2024, estimated word-level information density across 998 languages from 101 language families.

The denser languages tended to deliver matched spoken material faster. In a separate dataset of natural conversations across 14 languages, their speakers also tended to travel through a narrower region of computationally mapped meaning, staying closer to the opening subject and making smaller conceptual moves between turns.

This is one study, not settled consensus. The phrase “roughly 1,000 languages” applies to the first density analysis, not to every later test. Communication speed was assessed in 265 languages, real conversations in 14 and Wikipedia articles in 140. The design found associations and cannot show that language structure caused any group to discuss topics differently.

The researchers needed translations of the same material

Comparing raw word counts from unrelated texts would make little sense. A legal judgment in one language and a children’s story in another differ in content before language structure enters the picture. Aceves and Evans instead assembled parallel corpora, collections in which the same underlying material had been translated into multiple languages.

The sources included complete Bibles and New Testaments, news commentary, movie subtitles, TED talks, United Nations and European Union documents, medical and banking material, software localization files and example sentences for language learners. Coverage was uneven. The New Testament corpus supplied 828 languages, while many institutional collections contained only a few dozen.

Across all sources, the peer-reviewed paper included 998 languages in 101 families. The researchers reported that density estimates were correlated across very different document types, suggesting the measure was not solely a feature of one subject. Still, Bible translation carried much of the study’s extraordinary language reach.

“Information per word” was a compression measure

The team counted how often every word appeared in each translation, then used Huffman coding to build an efficient binary code for that frequency distribution. Common words receive shorter codes and rare words longer ones. Adding those codes across a document produces its compressed size in bits.

Because translations within a corpus were intended to convey the same source material, a language whose encoded document required fewer bits was treated as more informationally dense. Each translation was expressed relative to the English version of the same text.

This is more specific than the everyday phrase “packs more meaning into each word.” The researchers did not ask readers how much meaning they consciously extracted from individual words. Their measure depends on translated written corpora, word boundaries, tokenization and frequency patterns. A “word” is not equally easy to define across all language structures.

The authors compared their density measure with other descriptions of morphological complexity, fusion and grammatical informativity. Correlations were modest, suggesting it captured something related to, but not identical with, how many prefixes, suffixes or obligatory distinctions a language uses.

Dense word systems also formed denser maps of meaning

The next step used word embeddings. These models represent each word as a point in a many-dimensional space based on the company it keeps. Words that occur in similar contexts are placed closer together. The method does not understand concepts as a person does, but it can measure recurring semantic associations across a corpus.

Aceves and Evans trained separate embedding spaces for the parallel texts. Languages with higher information density also tended to have words positioned more closely together on average. The authors called this semantic density.

That relationship supplies the proposed link to conversation. If words sit amid more nearby associations, speakers may be able to approach the current subject from several related directions without making a large jump to a new topic. This is an interpretation of the model, not a direct recording of what speakers were thinking.

Faster communication was tested with 265 Audio Bibles

To study speed, the researchers needed spoken versions of comparable content. They collected the total duration of Audio Bible recordings in 265 languages from 51 language families. Denser languages tended to take less time to deliver the matched material.

The finding fits earlier evidence that languages can balance the density of their units against the rate at which those units are spoken. ScienceBlog previously covered a 17-language study reporting similar information transmission rates despite different syllable speeds. The new study works at the level of words and extends the comparison to far more recordings.

Its scale brings a trade-off. The speed analysis had one duration value per language. An Audio Bible is a long, carefully read performance, not spontaneous conversation. Narrators differ in cadence, dialect and production choices, and a single recording cannot describe the distribution of speaking styles within a language. The result is best read as shorter delivery of one matched genre, not proof that every speaker of a dense language talks faster.

Conversation breadth was measured in only 14 languages

The conversational evidence came from the IARPA Babel program, which collected natural speech for language technology research. The study used thousands of conversations in Amharic, Bengali, Cebuano, Georgian, Guarani, Igbo, Kazakh, Lithuanian, Tagalog, Tamil, Telugu, Turkish, Vietnamese and Zulu. Together they represented eight language families.

These were real interactions recorded in homes, offices, streets, vehicles and public places, with an average of 78 turns. The researchers converted each transcript into a cloud of word vectors using pre-trained fastText models based on the relevant language’s Wikipedia.

They calculated breadth in several ways. One measure asked how widely the conversation’s unique words spread around their semantic center. Another compared the first few utterances with the last few to estimate how far the discussion moved. A third measured the average semantic distance between consecutive turns.

Across these measures, denser languages were associated with narrower conversations. The ending stayed closer to the beginning, and each turn made a smaller move through the modelled conceptual space. The authors interpreted this as a tendency to explore a subject more deeply before changing topics.

“Less ground” describes geometry in a model

The study did not ask bilingual judges to rate whether each exchange felt broad, focused, insightful or repetitive. “Ground” meant distance between vectors. “Depth” was inferred from staying within a denser neighborhood of related concepts.

That distinction prevents an easy but unsupported ranking. Narrower does not mean poorer, duller or less creative. A discussion may remain close to one subject while introducing evidence, qualifications and competing perspectives. Conversely, a broad semantic path may reflect productive synthesis, casual topic switching or simple distraction. The metric cannot tell those possibilities apart by itself.

The embedding models add another layer of dependence. They were trained on Wikipedia, whose size, editorial community and coverage differ enormously across languages. Word meanings and relationships learned from an encyclopedia may not perfectly represent how the same words function in a home or street conversation.

Wikipedia reproduced the pattern at a different scale

As a final comparison, the researchers sampled more than 95,000 Wikipedia articles in 140 languages. They measured the average semantic distance among word pairs in each article, again adjusting for the density of the language’s conceptual space.

Articles in denser languages tended to cover narrower conceptual terrain. This echoed the conversation result in a larger set of languages and a different form of collective communication.

It is not a clean replication. Wikipedia editions are separate social institutions. They differ in contributor numbers, editorial norms, topic selection, article length, translation practices and source availability. Those factors can shape an article’s vocabulary independently of grammar or information density.

The study cannot separate language from its speakers’ circumstances

Languages are inherited through histories of migration, contact, population change and cultural practice. They are not randomly assigned to otherwise identical groups. An association between linguistic density and conversational movement could reflect language structure, shared cultural conventions, recording circumstances or some combination.

The models accounted for language family and tested controls including population size, climate, morphology and some cultural variables. But cultural measures were unavailable for every conversation language, and no statistical control can guarantee that an observational comparison has removed all meaningful differences.

The authors say this directly in their limitations: their design does not permit causal claims. Controlled experiments could ask bilingual speakers to complete the same discussion task in different languages. Ethnographic research could examine how topic movement works in ordinary settings at opposite ends of the estimated density range. Both would test whether the computational association survives when content and context receive closer attention.

The study therefore offers a broad pattern, not a deterministic rule. Across 998 languages, words appear to package matched information at different densities. In the smaller datasets available for speech and conversation, that density tracked how quickly material was delivered and how far discussion moved through a model of meaning. The intriguing possibility is that language structure helps shape the route of conversation. The evidence so far cannot say how much of the route language itself chooses.