A polished analysis lands in an inbox. Twenty pages, clear structure, specific numbers, sources and a short executive summary at the top. A document like this used to tell you something beyond what was written on the page. Someone had probably spent real time on it. Someone had tracked down the sources. Someone knew the subject well enough to organise the material properly. And there was a name at the top you could go and question if something didn’t add up.
None of that proved the analysis was actually good. But it gave you a way to judge what was likely sitting behind it. Generative AI has quietly broken that link. The same twenty-page document can now be produced in minutes by someone who knows very little about the subject and may never have opened the sources it cites. It still looks just as convincing. It’s just much harder to know what that appearance actually means anymore.
We still lean on the same old signals to decide who and what to trust: how polished something looks, whether a person has signed off on it, whether other people agree. But AI has quietly turned each of those signals into something that can be faked. There are three illusions worth pulling apart: the fluency trap, human in the loop and safety in numbers.
Illusion #1: The fluency trap
Most of the time, we can’t see competence directly, so we judge it from clues instead. Fluent writing suggests someone knows their stuff. So does specific detail. A bit of hesitation can tell you something honest is happening. Visible effort suggests care. Sources suggest someone looked beyond their own opinion. And when several knowledgeable people agree, that agreement carries extra weight.
None of these clues has ever been perfect but organisations need them anyway. Nobody has time to check every source or rebuild every argument from scratch. The real question now is whether these clues still mean what we think they mean.
A 2025 study in Nature Machine Intelligence found that longer explanations from large language models made people more confident in the answers they were given. But those longer explanations didn’t actually make people any better at telling right answers from wrong ones.[1] More explanation changed how sure people felt. It did nothing for how accurate they were.
The same pattern shows up elsewhere. A report with ten citations feels better researched than one with none, but AI can now generate those ten citations with almost no effort at all. A fast, fluent answer to a hard question used to be a reliable sign of real expertise. Now AI can produce that same fluent answer without any expertise behind it whatsoever. The clues haven’t disappeared. They’ve just stopped being trustworthy in the way they used to be.
Illusion #2: Human in the loop
One common response to all this is to put a person back into the process. The phrase is already familiar: AI generated, human reviewed. But that phrase can mean wildly different things depending on who’s doing the reviewing and how.
One reviewer might genuinely read the sources, challenge the assumptions, spot weak reasoning and send the work back for another pass. Another reviewer might have ten other things to approve that day, skim the summary, see nothing obviously wrong and click approve without a second thought. Both of these count, technically, as human review. But they are not remotely the same thing.
Ben Green looked at 41 government policies that required human oversight of automated decisions and found that most of them simply assumed people could provide useful oversight, without ever showing that those people actually had the time, the knowledge, or the authority to do it properly.[2] That exact problem is easy to spot in AI-assisted work today. A reviewer needs enough time to actually engage with the material. They need enough knowledge to challenge what they’re reading. They need real access to the underlying evidence. And, critically, they need the actual power to stop the process if something is wrong. Without all of that in place, “human reviewed” stops being a safeguard and starts being just a reassuring label. A senior person sees those two words on a document and quietly assumes someone else has already handled the hard parts.
Illusion #3: Safety in numbers
Agreement between people is another clue that’s becoming harder to read reliably. Suppose five people independently look at the same strategic question and arrive at the same conclusion. That has traditionally meant something, because those five people bring different experience, notice different things, and tend to make different mistakes along the way. Agreement, in other words, used to be evidence of real independent thinking.
Now imagine all five of them lean heavily on AI while doing their work: using it to search, to summarise research, to generate options, and to stress-test their first instincts. There are still five people in the room, and there may well be five separate documents on the table. But large parts of their thinking may now be running through the same models, drawing on similar sources and inheriting similar blind spots.
Research across more than 350 large language models found that their errors were often correlated rather than random. On one benchmark, when two different models both got a question wrong, they gave the exact same wrong answer 60 per cent of the time, and similar patterns showed up across different models and different providers.[3] That research is about models rather than teams of people but it’s a useful warning all the same. Five people agreeing no longer automatically means five genuinely independent paths led to the same place.
Cue vs. ground
Thinking all of this through, I kept coming back to one useful distinction. A citation can still be useful. So can human review. So can five people agreeing. I think of all of these as cues: something that suggests an answer is probably reliable, without actually confirming it. A ground, on the other hand, is something you can go and check for yourself.
A citation is only a cue when it’s sitting quietly at the bottom of a report. It becomes a genuine ground the moment you actually open the source and check whether it really supports the claim being made. “Human reviewed” is a cue on its own but knowing exactly who reviewed the work, what they actually checked and whether they had the power to reject; it turns that cue into something much sturdier. Five people agreeing is a cue too, until you know how independent their evidence and thinking really were.
Organisations can’t function without cues; there’s simply too much information and never enough time to check all of it from first principles. The trouble starts when a familiar-looking cue quietly stops carrying the meaning we’ve come to assume it has.
Earn it
I started out thinking about trust but I’ve come to believe the more useful question is confidence: specifically, what should earn it now. People will keep trusting their colleagues, their organisations and the systems they rely on, because work simply can’t happen otherwise. What’s changed is the evidence we should be using to decide when that trust is actually deserved.
AI has made most of the visible signs of competence cheap to produce. A polished report is cheap. A detailed explanation is cheap. A long list of sources is cheap. A confident, well-structured argument is cheap. Even a thoughtful-sounding list of objections can now be generated in seconds. None of that means everything needs to be checked from scratch, most work would grind to a halt if every claim had to be independently verified before anyone could move forward. The level of checking something deserves should scale with what’s actually at stake. A quick internal summary probably needs very little scrutiny. A decision about an acquisition, a medical diagnosis or someone’s job needs a great deal more. In those higher-stakes cases, it’s worth asking where a claim really came from, what assumptions sit underneath it, what evidence points the other way and who is still genuinely accountable for the final call.
That twenty-page analysis is still sitting in the inbox. It still looks good. The numbers are still specific, the sources are still listed and someone’s name is still at the top. Those signs still tell us something real. They just don’t tell us nearly as much as they used to.
References
1. Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, “What large language models know and what people think they know”, Nature Machine Intelligence, 7, 221–231 (2025).
2. Ben Green, “The flaws of policies requiring human oversight of government algorithms”, Computer Law & Security Review, 45, 105681 (2022).
3. Elliot Myunghoon Kim, Avi Garg, Kenny Peng and Nikhil Garg, “Correlated Errors in Large Language Models”, Proceedings of the 42nd International Conference on Machine Learning, PMLR 267, 30038–30066 (2025).
Human & Machine studies how judgement fails under complexity. This piece is part of that work.
Does this pattern feels familiar? You can explore it further through the Decision Audit.


