In nine preregistered experiments involving 313 people, participants appeared able to dynamically advance only one independent imagined moving object at a time. They could remember that two objects existed, but their response timing suggested that the internal motion was processed in sequence.
Halely Balaban of the Open University of Israel and Harvard psychologist Tomer D. Ullman reported the finding in a 2025 Nature Communications paper. The result surprised them because visible objects can often be tracked in small groups.
This is one study, not settled consensus. “One” does not mean that people can see, remember or picture only one object. It refers to the narrower act of updating independent imagined objects as they move through time.
What the study actually measured
The core task was simple to describe. Participants watched a short two-dimensional animation of one or two colored balls moving toward a line that represented the ground. After half a second, the animation paused and the balls disappeared.
Participants then had to continue the motion mentally. They pressed a key when each invisible ball should have reached the ground. Depending on the trial, the correct impact time was 1.0, 1.2, 1.4 or 1.6 seconds after motion began.
Those different times turned an invisible thought into a measurable pattern. If two imagined motions progressed together, each response should follow its own correct arrival time. If the mind processed them in sequence, work on the second ball should wait while the first simulation ran.
The researchers were separating two abilities that ordinary language tends to fold together. One is maintaining multiple representations, such as remembering that two balls exist and where they were. The other is repeatedly changing those representations to reflect what should happen next.
That second operation is mental simulation. It is the difference between holding a photograph in mind and running an internal animation.
The studies were conducted online with English-speaking adults in the United States recruited through Prolific. Each main experiment retained 36 participants after exclusions set in advance for such things as failing comprehension checks or not producing usable responses.
The 640-millisecond signature
In the first experiment, participants imagined a single ball. Their response times changed appropriately with the ball’s hidden travel time, evidence that they were not merely guessing or pressing after a fixed interval.
The important change came when two balls moved independently. The first response continued to follow the relevant trajectory. The second was delayed by an average of 640 milliseconds.
More telling than the delay alone was its shape. The researchers compared a parallel model, in which both imagined balls keep moving together, with a serial model, in which one simulation is completed before attention turns to the other. The serial model explained 96 percent of the variation in the observed timing pattern. The parallel model explained 20 percent.
The models did not claim that the second ball vanished from memory. Both assumed that people retained the objects and their starting conditions. They differed in whether the internal dynamics of both objects could be advanced simultaneously.
Ullman told the Harvard Gazette that he had expected a capacity of perhaps three or four objects, partly because earlier research found that people could keep track of several visible items. The answer of one surprised the researchers too.
Why two key presses were not enough to explain it
A two-response task creates an obvious alternative explanation. People may have simulated both trajectories at once but struggled to make two key presses in quick succession.
Experiment 2 kept the balls visible until impact while preserving the two-response requirement. The large serial signature almost disappeared. The average delay for the second response fell to 88 milliseconds, the parallel model explained 97 percent of the pattern, and the serial model explained 47 percent.
That control does not show that motor limits played no role at all. It shows that making two responses was not enough to produce the much larger delay seen when participants had to continue two motions in imagination.
The team also tested whether strong grouping cues would let two objects share one update. When the balls moved together and could be treated more like a common unit, the delay for the second response fell to 328 milliseconds.
The result sat between the clean parallel and serial predictions. Grouping made the task easier, but it did not fully erase the bottleneck.
The effect survived changes in physics and presentation
The team did not rely on one version of falling balls. In Experiment 4, the objects disappeared behind an occluder, making their absence look like a natural consequence of the scene rather than an arbitrary screen event. The second response still lagged by 364 milliseconds, and the serial model explained 96 percent of the timing pattern.
Experiment 5 removed falling, gravity and collision with a ground line. Participants instead continued simple straight-line motion. The second response lagged by 423 milliseconds. The serial model explained 98 percent of the pattern, compared with 27 percent for the parallel model.
Three supplementary experiments tested further alternatives. Added variability did not rescue a noisy parallel account. Extra financial motivation did not make the bottleneck disappear. Extended practice and more detailed timing measurements continued to support the serial prediction.
The authors made the preregistrations, data and analysis code available through their Open Science Framework project, allowing other researchers to inspect the decisions and rerun the analyses.
What “one” means, and what it does not
People can track several moving targets while they remain visible. Classic multiple-object tracking studies often find a capacity of roughly three or four under suitable conditions, although performance changes with speed, spacing and task design. That is a limit on selecting objects from continuing visual input, not on generating their unseen motion.
Visual working memory is different again. It asks how many items or features can be retained over a delay. These experiments were designed so that remembering two balls was not enough. Participants had to update where each ball ought to be at successive moments.
Nor is the result a claim about imagery vividness. Some people report little or no voluntary visual imagery, while others experience vivid mental pictures. Capacity, vividness and dynamic updating are related questions, but they are not the same measurement.
ScienceBlog has previously covered evidence that rats can mentally navigate to places they are not physically occupying. That work asked whether an animal could represent an absent route or destination. The new study asks how many independent moving representations the human mind can advance at once.
The result should not be turned into a general score of intelligence, creativity or imagination. Performance on brief animations cannot tell us how richly a person invents a story, plans a journey or understands a physical system.
Why imagination can still feel crowded
Grouping may help explain why imagination feels richer than the measured limit. A structured scene does not require every object to have a fully independent internal clock. Several objects can inherit movement from a common group.
A car carries its passengers. A turning Ferris wheel carries its cabins. If the mind can update the larger structure and let its parts follow, subjective detail can exceed the number of independent trajectories being calculated.
That interpretation remains an inference from behavior. The experiments used response timing and computational models. They did not use brain imaging to identify a neural gate, and they did not establish why such a gate might exist.
The experiments also deliberately simplified the world. Most tested objects moved independently and did not collide after disappearing. Real scenes contain interaction, prediction, prior knowledge and opportunities to group events.
Eye movements were not controlled in the online studies. Shifts of gaze between remembered paths could contribute to serial performance. The sample was also drawn from adults in one country using one platform. Laboratory replications, other populations and richer three-dimensional tasks would show how general the limit is.
“The imagination can track only one object” is memorable, but too broad without its verbs and conditions. The evidence supports a smaller statement: in these tasks, people advanced one independent imagined moving object at a time.
That still leaves room for crowded scenes, static images rich in detail, remembered layouts, grouped motion and quick switches of attention. We may hold a broad world in mind while its independent moving parts take their turns.