Researchers conducted five experiments with 3,132 participants to test whether access to a language model changes how people handle uncertainty. Just having AI advice available nearly wiped out people's willingness to say I don't know. The researchers picked questions where the model they used, Step 3.5 Flash, was almost always wrong. These questions focused on fine visual details from movies, like the color of a team uniform in Bend It Like Beckham. According to the authors, such details rarely appear in online text, making them prime targets for hallucinations. GPT-5.5, Claude 4.6 Sonnet, and Gemini 3.5 Flash got most of the other questions right but sometimes failed on the harder ones.
Because the AI advice was mostly wrong, the effect cannot be explained as reasonable delegation to a reliable tool. The researchers still see a tendency for people to defer to AI answers, but whether this holds just as strongly beyond movie trivia remains an open question. In the first two studies, participants could choose whether to ask an AI for advice. In the control group without AI access, they withheld their judgment on 36 and 44 percent of questions. With AI access, those numbers dropped to 6 and 3 percent.
In Study 2, the researchers also measured participant confidence. Without financial incentives, confidence with AI access hit 75.9 points on a scale of 100, compared to 29.6 points without AI. That is roughly two and a half times higher. At the same time, the share of correct answers fell from 27.6 to 10.0 percent. Participants were far more confident but wrong much more often. Across all studies, participants without incentives who had AI access got 9.2 percent of questions right, versus 27.5 percent without AI.
Read nextGoogle Makes Free 1080p AI Video Generation Available Inside Google VidsStudies 2 through 4 added financial incentives. Participants earned 10 cents for each correct answer, lost 10 cents for each wrong one, and got nothing for I don't know. The researchers had pre-registered the hypothesis that incentives would increase willingness to abstain, but that AI availability would weaken that effect. None of the three studies showed a statistically significant interaction. AI access and financial incentives worked largely on their own, according to the study. Incentives led participants to seek AI advice less often and to answer correctly more often when AI was available.
In Study 4, the AI answer was shown to participants automatically without them having to ask. The researchers say this mirrors a growing reality where search engines display AI-generated summaries and writing assistants offer unsolicited suggestions. The effect barely changed. Without incentives, judgment suspension dropped from 35 percent without AI to 1 percent with automatically displayed AI answers. With incentives, it fell from about 39 to 7 percent.
The researchers frame their results in the context of Epistemia, the tendency to accept AI answers because they sound convincing rather than actually checking them. A language model always has to produce an answer and never pauses, even when it does not know something. When users hand off their judgment to such a system, they may adopt its lack of restraint, the authors write. A growing number of studies suggest that using AI affects human judgment and critical thinking. The next known step involves further research to see if this trend holds outside of movie trivia.



