For most of the 20th century, one of the most reliable trends in psychology was that intelligence test scores kept rising. The New Zealand researcher James Flynn documented it across dozens of countries: on the same standardized tests, each generation outperformed the one before it by roughly three IQ points a decade, a pattern that became known as the Flynn effect. Then, in several of the countries where the rise had been most carefully measured, it stopped, and in some cases reversed.

The clearest evidence for the reversal comes from a 2018 study in the Proceedings of the National Academy of Sciences, led by Bernt Bratsberg and Ole Rogeberg of the Ragnar Frisch Centre for Economic Research in Oslo, and later summarized in more detail by the science outlet The Science Breaker. Norway conscripts nearly all of its young men, and gives each one a standardized intelligence test on entry, which means the country holds an unusually complete record of scores stretching back decades. Bratsberg and Rogeberg drew on 730,000 of those test results, covering men born between 1962 and 1991 and tested between 1970 and 2009.

It is worth being specific about what the test measures, since “intelligence test” can suggest more than the data actually supports. The Norwegian conscription test is a composite of three parts: arithmetic, word similarities, and figure completion, the kind of test designed to approximate general reasoning ability rather than any single narrow skill. It is not a clinical or diagnostic instrument, and the study does not claim to measure anything beyond performance on that specific composite, taken by young men at a specific age, under specific testing conditions, across five decades.

A turning point in 1975, then a steady decline

Scores rose steadily through cohorts born up to about 1975, consistent with the Flynn effect documented elsewhere. After that birth year, the trend reversed. Men born later scored lower, on average, than the men born just before them, a decline the researchers put at roughly seven points across a generation. This is one study, not settled consensus on its own, though it is among the largest and most carefully designed on the question, and its central finding, a rise followed by a fall within the same national testing program, has since been reported in Britain, Denmark, Finland and France as well, each using its own compulsory military testing records rather than a shared dataset.

What makes the Norwegian study particularly useful is not just its size but a specific piece of methodology aimed at a specific rival explanation. One theory of declining intelligence, sometimes called a dysgenic hypothesis, holds that if people with lower test scores have more children on average than people with higher scores, average scores could fall across generations for genetic reasons. Bratsberg and Rogeberg tested this directly by comparing brothers within the same family rather than comparing the population as a whole. A genetic, family-level effect of that kind would show up as a decline between families but not within them, since brothers share most of the same genetic and family background.

That is not what the data showed.

The decline showed up just as strongly within families as across the population, which the authors say rules out a genetic explanation and points instead to something in the environment that changed for everyone born after the mid-1970s, regardless of which family they were born into.

A different country, a more complicated picture

A separate, more recent dataset adds a useful complication rather than a simple confirmation. In 2023, Elizabeth Dworak, a research assistant professor at Northwestern University’s Feinberg School of Medicine, published an analysis in the journal Intelligence using data from the SAPA Project, a free online cognitive assessment. Her team looked at 394,378 people who took the test between 2006 and 2018 in the United States. Scores fell over that period on verbal reasoning, matrix reasoning, and letter and number series tasks. But scores on a fourth measure, three-dimensional spatial rotation, actually rose over the same years.

That detail matters, because it argues against a single tidy story of general decline. Dworak was careful not to claim a cause. Among the possibilities she raised were changes in what schools and workplaces emphasize, more time spent on visually complex media that could plausibly train spatial skills specifically, and a blunter possibility: that some of today’s test-takers may simply be less practiced or less motivated at the kind of test-taking the older instruments were designed to measure, a factor that would show up as a score change without reflecting a real change in underlying ability.

The two datasets are not really measuring the same thing, and the gap between them is worth sitting with rather than smoothing over. The Norwegian sample is close to a full population: nearly every eligible young man in the country, tested under the same conditions by the same authority, year after year. The American sample is self-selected: people who found the SAPA Project online and chose to take a free cognitive test, which is a very different group from a national conscription cohort, and one that could shift over time in ways that have nothing to do with intelligence, such as who is inclined to spend twenty minutes on an online reasoning test in 2018 versus 2006. Dworak’s team was upfront about this limitation rather than hiding it.

It does not make the American findings wrong, but it does mean the two studies sit at different points on the reliability scale, and a reader comparing them should weigh the Norwegian result more heavily than the American one on the strength of sampling alone.

What the reversal does not tell us

None of this supports a claim that people are becoming less capable in some general, folk sense of the word, and none of the researchers involved argue that it does. The original Flynn effect itself was never fully explained either. Leading theories for the rise, better nutrition, more schooling, more cognitively demanding jobs and leisure, more familiarity with abstract test-like reasoning, remain leading theories for what might now be fading or reversing, rather than settled facts on either side. I wrote recently about the century-old finding that performance on different mental tests tends to correlate, the basis for treating intelligence as measurable at all. The Flynn effect and its reversal are a reminder that the number the tests produce, whatever it is measuring, is not fixed. It moves with the environment people are tested in, generation to generation, for reasons researchers are still working out.

What would settle the question is the same thing that would settle most questions like it: more countries, more decades, and more datasets built with the same care Bratsberg and Rogeberg put into ruling out the explanation that seemed, on its face, like the simplest one.