Trust usually makes a tool easier to use. A calculator that repeatedly produces the right answer earns a place on the desk. A navigation system that reliably finds the quickest route no longer needs to be checked against a paper map. Confidence removes friction, and much of modern technology is designed to earn exactly that response.
Generative AI introduces a harder version of the same bargain. Its answers can sound complete and assured even when the reasoning beneath them is weak, the evidence is missing or the context is wrong. Yet as people become confident that the system can handle a particular task, they may feel less reason to inspect what it gives them.
That pattern emerged in a survey by researchers at Microsoft Research and Carnegie Mellon University. Among 319 knowledge workers describing 936 real examples of AI-assisted work, greater confidence in the AI’s ability to perform a task was associated with less reported critical thinking. Confidence in their own ability moved in the other direction: people who felt capable of doing the task themselves reported more critical engagement.
What 319 knowledge workers were asked
The researchers recruited participants through the online platform Prolific and asked about situations in which they had used generative AI at work. The group spanned a variety of occupations and supplied three examples each, producing 936 task descriptions involving creation, working with information or seeking advice.
Before answering, participants were introduced to critical thinking through concrete examples drawn from levels of Bloom’s taxonomy. These included checking the tone of an AI-written email, verifying a code snippet and considering bias in a data analysis. The aim was to avoid treating critical thinking as only one kind of abstract reasoning.
For every example, the survey asked what the person was trying to accomplish, which AI tool they used and how they used it. Participants then rated factors including their confidence in handling the task themselves, their confidence in the AI, whether critical thinking was involved and whether the tool made different thinking activities more or less effortful.
The resulting Microsoft Research paper was presented at the CHI 2025 Conference on Human Factors in Computing Systems. It was written by Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks and Nicholas Wilson.
The confidence relationship ran in two directions
After accounting for task- and user-related factors in explanatory regression models, the researchers found that confidence mattered in two distinct ways. Higher task-specific confidence in generative AI predicted less reported critical thinking. Higher confidence in one’s own ability predicted more.
This is more informative than saying people either trusted or distrusted AI in general. A worker may have high confidence in a chatbot’s ability to rewrite a routine paragraph and low confidence in its ability to interpret an unfamiliar contract. The study examined confidence in relation to the particular task being described.
It also found that AI often reduced the perceived effort involved in critical-thinking activities. That can be useful. Breaking a blank page, collecting background information or converting notes into a first draft may genuinely require less labour with a capable assistant.
The concern appears when lower effort becomes lower scrutiny. If the output looks plausible and the user expects the system to be competent, independently checking evidence, considering alternatives or reconstructing the reasoning can feel redundant. Trust has removed the very friction that might expose an error.
An earlier ScienceBlog report on the study captured this central confidence effect. The fuller picture is not that thought simply vanishes, however. Some of it moves.
AI changed the location of the thinking
Workers described a shift away from producing every part of an artefact themselves and towards supervising what the system produced. The researchers call this a movement from material production to critical integration.
When creating content, participants spent less effort drafting and more effort defining what a suitable result should look like. When seeking information, the task moved from gathering material towards checking whether the returned material was accurate. When asking for advice, users had to decide how an answer fitted the constraints and consequences of the real situation.
The paper groups these newer demands around information verification, response integration and task stewardship. A person might check facts against external sources, edit generated language so it fits an audience, or remain responsible for the final decision even though the system supplied much of the apparent reasoning.
That is still critical thinking. In skilled hands, it may be a better allocation of attention than manually doing every routine step. A software developer does not need to type each predictable line if the saved effort is spent testing edge cases and understanding system behaviour. An analyst can accept help assembling a table while concentrating on whether its categories answer the right question.
The catch is that oversight is easy to name and hard to sustain. Reading a polished answer can create the feeling of having evaluated it when the reader has only recognised that it sounds reasonable.
What the study does not prove
The survey did not measure participants’ critical-thinking ability before and after months of AI use. It did not randomly assign some workers to use AI and others to work unaided. It also did not independently score the accuracy or quality of the 936 work products.
Its measures were largely based on participants’ recollections and perceptions of their own behaviour. Self-reports can reveal how work feels and how people explain their choices, but they cannot directly show every mental step that occurred. People may forget checks they performed, overstate their diligence or use different personal definitions of what counts as critical thinking.
The association also leaves room for several explanations. Confidence in AI may lead people to scrutinise less. Alternatively, people may become confident in AI precisely on routine tasks where they have already decided that intensive scrutiny is unnecessary. Both processes could operate at once.
For those reasons, the finding should not be translated into “AI makes people unable to think.” The study’s own title refers to self-reported reductions in cognitive effort. It identifies a workplace pattern and plausible risk, not cognitive damage or a proven long-term loss of skill.
Sometimes doing less thinking is rational
Critical thinking has a cost. It consumes time and attention that cannot be spent elsewhere. No professional independently derives every formula in a spreadsheet, verifies every word in a spellchecker’s dictionary or audits every database query after the surrounding system has proved reliable.
If an AI is genuinely dependable on a narrow, low-consequence task, reducing scrutiny can be efficient. Treating every generated sentence as a potential emergency would erase much of the benefit of using the tool. The goal cannot be permanent suspicion.
The better target is calibrated trust: confidence that rises and falls with the system’s demonstrated reliability, the user’s own knowledge and the cost of being wrong. A rough agenda for an internal meeting deserves a different review process from a legal filing, a safety calculation or an explanation sent to a client.
Calibration is difficult because fluency is not evidence. An AI can present a weak claim in clear prose, include invented detail and maintain the same composed tone it uses when correct. Interface cues can deepen that impression. As ScienceBlog reported in a separate experiment, people judged otherwise similar AI answers as more thoughtful when the system took longer to respond.
What feels considered may therefore reflect presentation as much as quality. A confident user needs evidence about performance, not merely the experience of smooth interaction.
Why expertise protects scrutiny
The positive relationship between self-confidence and critical thinking offers a second lesson. People who understand a task have more internal material against which to test an answer. They notice missing assumptions, strange terminology and conclusions that do not fit the setting.
A novice often faces the opposite problem. The reason for asking AI may be that they do not know enough to evaluate its response. If the answer is fluent, there may be no obvious foothold for disagreement. Telling that person to “check the output” is not a complete safeguard because checking requires knowledge, reliable sources or a second process that does not simply repeat the first answer.
This creates an awkward training problem. AI can help less experienced workers complete tasks that were previously beyond them, which may widen access and accelerate learning. But if it also removes the need to practise the underlying steps, the worker may struggle to develop the expertise required for future oversight.
The risk is not limited to factual error. Doing the work oneself teaches which details are difficult, which trade-offs recur and where a superficially good answer tends to fail. Those are the patterns that later support professional judgment.
How tools can keep the user engaged
The researchers argue that AI systems should be designed as tools for thought, not only engines for immediate answers. That does not require making every interaction slower or more annoying. It means placing useful friction where the stakes or uncertainty justify it.
A system could expose sources and uncertainty, separate observations from inference, invite the user to state criteria before generating a recommendation, or offer credible counterarguments rather than only strengthening the initial request. For complex work, it could pause at decision points and ask the user to approve assumptions instead of hiding them inside a finished response.
Organisations also shape the outcome. A workplace that rewards only speed sends a clear signal that checking is wasted effort. Review time, independent verification and responsibility for the final product need to be built into the workflow rather than left to individual discipline after a deadline has already arrived.
One practical approach is to separate generation from evaluation. The person can first define what a good answer must contain, use AI to produce or organise material, and then test the result against those criteria and an independent source. For important decisions, a colleague who did not see the prompt can review the output without inheriting its framing.
The quiet risk is misplaced confidence
The survey’s most useful contribution is not a verdict that AI is good or bad for thought. It shows why the same tool can lead to very different cognitive behaviour depending on where confidence sits.
When workers trusted their own understanding, they reported more critical engagement. When they trusted the AI to handle the task, they reported less. That can represent sensible delegation when the system is reliable and the consequences are small. It can also become over-reliance when an answer’s polish is mistaken for proof.
Independent scrutiny rarely disappears with a dramatic decision. It fades because each successful interaction makes the next check feel a little less necessary. The challenge is to gain the efficiency of AI without allowing familiarity to become evidence and convenience to become unquestioned authority.