Three research papers landed in the last few weeks that, taken together, should unsettle anyone who runs a creative organization. They come from different fields: computational linguistics, rhetoric, and behavioral economics. And none of them cite each other. But they’re describing a syndrome in three different dimensions. I’m not talking about about efficiency or automation or any of the usual AI debates. It’s scarier than that.
What happens to the quality of your agency’s creative work when the people doing it have quietly lost the ability to generate truly original thoughts…and can’t tell.
1. The range of “arguments” is collapsing
In this case it means “logical argument” like a perspective or rationale. You know, like when in a spirited discussion someone says something that feels like left field…for a moment, but then you realize that they’re viewing this from a really interesting angle, and maybe you missed something? It’s how humans make things better together, collaboration and all those other buzzwords.
So, a team at the University of Maryland’s CLIP Lab collected over a thousand human-written opinion responses from New York Times debates and nearly five hundred from Boston Review forums. Then they generated more than 23,000 essays from leading LLMs on the same prompts and compared.
The finding, in the NYT corpus: 65.3% of human arguments were unique within a given debate. For LLMs, that number was 3.4%.
When humans take a position, they come at it from wildly different angles; whether from lived experience, unexpected analogies, domain-specific knowledge, idiosyncratic logic, and on. Needless to say, the variety and range of humans as a group is enormous.
But when LLMs take a position, they converge on a narrow set of “reasonable” framings almost every time. (I explain this LLM corpus and RLHF phenomenon in chapters 3-6 of my upcoming book,The Mostly Helpful Psychopath)
And like great researchers do, they looked a bit deeper. Among essays that shared the same main thesis, 41% of human supporting arguments were still unique. Different evidence, different structure, different reasoning. For LLMs, a paltry 9%.
Now let’s take it home: your agency.
The brief comes in. The team reaches for AI to brainstorm angles, develop concepts, draft positioning. What comes back sounds smart. It sounds thorough. But it’s drawing from a fraction of the range a room full of humans would have generated. And because the AI output is fluent and fast, (and for reasons I'll cover in paper 3 below) it's not so easy to see that the range has collapsed. The work looks elevated. It just isn’t…it’s commonly good, but only good.
The danger isn’t that AI produces bad work. Increasingly, the work that will be most competitive will be the work whose ideas weren't driven by an AI.
Question one: Where in your process does raw human idea generation happen before AI touches the brief — and who’s accountable for protecting that space?
2. The voice is flattening
A separate body of research, now substantial enough that multiple teams are building on each other’s findings, shows the narrowing goes beyond ideas to how they’re expressed.
Co-writing with an LLM leads to stylistic convergence across tasks, across users, and even across model families from different companies. One study on scientific writing found that native-language signals (the subtle markers that reveal a writer’s first language and cultural frame) are being smoothed out at a measurable rate, with detection dropping over 10% in the post-LLM era.
This isn’t grammar cleanup. It’s the slow erasure of the thing that makes one creative voice different from another. The odd rhythm, the unexpected register shift, the sentence that shouldn’t work but does — that’s where the best work lives. It’s also exactly what gets polished away when AI is in the loop, because the models optimize toward the middle of the distribution. They make everything sound like a better version of average.
Every agency says voice and tone are part of the value they deliver. What the research suggests is that the more AI touches the work, the more that voice converges toward a shared default that belongs to no one and distinguishes nothing.
Question two: If your best writer’s voice started flattening toward the median, how would you know, and who else would notice?
3. Your teams are probably blind to AI’s weaknesses.
And additionally, they get blinded to their own.
Researchers from Milano-Bicocca, École Normale Supérieure, and Sapienza ran five experiments with over 3,100 participants, published this month as a preprint. They designed questions where the AI advice was deliberately wrong, separating the effect of having an AI available from the effect of AI being accurate.
People’s willingness to say “I don’t know” dropped from 44% to 3%. Accuracy fell from 27% to 9%. And confidence nearly doubled, from 30% to 76%.
Let me play that back: 1) the willingness to say you don’t know is the moment that you might have started learning something; and 2) your (likely) over-self-assurance means that you’re answering when you shouldn’t (unknowingly), and 3)(OMG) being ever-more confident that your ignorance and incorrect answers are right!
This makes our first two problems invisible.
When the range of ideas narrows and the voice flattens, the people using the tools don’t feel like anything is missing. They feel more confident, not less.
The internal signal that would normally say (or even scream) “Wait, wait, have we actually pushed this far enough?!?”, the creative restlessness that drives the best work…it gets dampened. The AI tools have driven your brain to see something merely good enough as better than good enough.
Wharton researchers named a related phenomenon “cognitive surrender” earlier this year — people accepting incorrect AI outputs 80% of the time while reporting higher confidence than those working without AI. The surrender isn’t dramatic. It’s the opposite. It feels like flow.
Question three: When everyone on the team feels more confident about the work, what in your review process is built to distrust that feeling?
Not AI bashing
You need to be using AI tools. Try cutting down a tree without a chainsaw. Not worth it, except as a hobby. Like with chainsaws, you need to be sure they’re being used properly. You need to design process (yes, and structure) around what AI does to the people doing the work.
The agencies that see this first have a window. Not to reject the tools, but to build the practices the three questions point at: idea generation that preserves human range, voice that resists the pull toward the median, and judgment that stays calibrated under confidence inflation.
The three questions are process and structure questions — which is the work I do. I've spent 15 years inside 200+ agencies rebuilding how the work works, and lately, that means building the practices that keep human range, voice, and judgment intact while the tools do their part. If you're running an agency and you can't answer one of the three questions, that's worth a conversation. Tell me what you're seeing — no pitch, no pressure.