I want to be upfront about something before I start: I am a writer, not a psychologist or a cognitive scientist. What follows is my reading of a long argument, not a verdict on it. The phrase “true intelligence” gets used as though everyone agrees what it points to, and the closer I look, the less settled that agreement turns out to be.
People have been trying to pin the word down for more than a century. The strange part is that the most durable result in all that time is not a grand theory of the mind. It is a correlation.
The finding that started the argument
In 1904, the psychologist Charles Spearman gave a group of schoolchildren a range of unrelated tests, from academic subjects to simple sensory discrimination, and noticed something he had not expected. The scores tended to move together. A child who did well on one kind of task was more likely than not to do well on the others, even when the tasks looked like they had nothing in common.
Spearman called this pattern of all-positive correlations the “positive manifold,” and he proposed that a single underlying factor, which he labeled g for general intelligence, was partly responsible. The specific skills sat on top of it.
What has kept psychologists interested is not Spearman’s explanation but the pattern itself. The positive manifold has held up across an enormous number of studies and populations, and reviews of the field describe it as one of the most replicated findings psychology has. A 2012 review in American Psychologist by Richard Nisbett and colleagues, summarizing decades of intelligence research, treats it as a starting point rather than a live question.
What g is, and what it is not
It is easy to hear “general intelligence” and picture a single substance in the head, like fuel in a tank. That is more than the evidence supports, and it is worth being careful here.
The g factor is a statistical regularity. It falls out of the math when you analyze how test scores relate to one another. It reliably predicts things, including performance in school and, more modestly, in some kinds of work. But the correlation itself does not tell you what mechanism produces it, and researchers genuinely disagree about that. Some read it as one central capacity. Others argue the correlations could arise from many small cognitive processes that reinforce each other during development, with no single master variable underneath. The pattern is solid. The story behind the pattern is not.
This is the sort of distinction that gets flattened in everyday talk. I have written before about the hobbies people like to associate with a higher IQ, and part of what makes that kind of piece fun to write is also what makes it slippery. The everyday signs we treat as intelligence are not the same thing as what a test measures, and neither is guaranteed to be what we actually care about in a person.
More than one kind, or one kind at different jobs?
Not everyone was satisfied with a single factor, and the pushback produced some of the most influential ideas in the field.
Raymond Cattell drew a line between fluid intelligence, the ability to reason through a genuinely new problem, and crystallized intelligence, the accumulated knowledge and skill you carry with you. The two behave differently over a lifetime. Fluid reasoning tends to peak relatively early and drift down, while crystallized knowledge can keep growing for decades. That split is now folded into the widely used Cattell-Horn-Carroll model, which keeps a general factor at the top but places a layer of broad, distinguishable abilities beneath it.
Then there is Howard Gardner, whose 1983 theory of multiple intelligences proposed separate and independent intelligences, including musical, bodily-kinesthetic, and interpersonal, alongside the usual verbal and mathematical kinds. The idea was popular in education, and it is easy to see why. It is generous. It hands more people a form of intelligence to claim.
The trouble is that it has not held up well against the data. The abilities Gardner described do not come apart cleanly the way the theory needs them to. They correlate, which is the positive manifold showing up again. Cognitive psychologists such as Lynn Waterhouse argued in the mid-2000s that there was little empirical support for treating them as distinct intelligences, and Gardner himself acknowledged at one point that the hard evidence was thin. Many of the categories look more like talents or interests than separate engines of thought. The related “learning styles” idea, the notion that each of us learns best as a visual or auditory or kinesthetic type, has been studied repeatedly and is now widely regarded as a myth.
None of this means the categories are useless as a way of noticing what people are good at. It means the science does not support the strong claim that they are independent.
The same argument, now about machines
The reason any of this feels current is that we have started building things that force the question again.
When people ask whether an AI system has “true” intelligence, they are reopening Spearman’s argument in a new setting. Is intelligence a general capacity that transfers to problems you have never seen, or is it a large collection of specific skills that only look general when you stack enough of them together?
The researcher François Chollet put this distinction at the center of a benchmark called ARC-AGI. In his 2019 paper On the Measure of Intelligence, he argued that intelligence should be measured not by the skills a system already has but by how efficiently it picks up new ones on tasks it was not prepared for. A calculator is not intelligent for doing arithmetic. The interesting question is how well something handles the genuinely unfamiliar.
The results so far are a useful reminder to stay humble about the word. According to the ARC Prize, a version of OpenAI’s o3 model reached roughly 76 percent on the original ARC-AGI puzzles under low-compute conditions and about 88 percent under high compute, the first time a system had crossed the human baseline on that particular test. That is a real result and it was widely discussed. But when Chollet and colleagues released a harder version, ARC-AGI-2, in 2025, frontier model scores fell sharply while ordinary people could still solve most of the tasks. A test built to resist memorization exposed how much of the earlier performance leaned on it.
I do not read that as proof that machines will never get there, and I would be wary of anyone who claims to know either way. What it shows is narrower. Skill on a set of tasks and the general adaptability we have in mind when we say “intelligence” are not the same thing, and the gap between them is easy to underestimate.
Why the word stays slippery
After a century, the sturdiest thing we can say is modest. Mental abilities tend to travel together, which is why a general factor keeps reappearing whether we are measuring children, adults, or, in a different form, software. Beyond that, most of the confident claims start to wobble.
Intelligence tests capture something real and stable, and they predict some outcomes. They also miss a great deal that people reasonably mean by the word, including judgment, curiosity, and the practical sense to know which problem is worth solving. Treating the test score as the whole of it is a category error, and treating it as meaningless is another.
The honest position, as far as I can tell from reading the field, is that “true intelligence” is less a thing we have located than a question we keep sharpening. That is not a failure. A question that has survived this long, and that now applies to our own machines, is probably pointing at something worth the trouble.