Everyone warns you about the AI that agrees with everything. I wrote here two days ago that I had the opposite problem — six months of running a factory of agents, and my time goes on pushing them to do their job, not on stopping them doing something reckless.
Both of those are true. Both of them are beside the point.
The mistake underneath them is a way of thinking almost everyone brings to this, including me when I started: that AI is a machine, and machines are things you optimise toward perfection. Find the flaw, tune it out, repeat. On that model, sycophancy is a bug with a fix, and over-caution is a different bug with a different fix.
It is not that kind of thing. AI mirrors us — both the people it was trained on and the people it works with. So it is fallible when it is certain and fallible when it yields, for much the same reason a colleague is. There is no setting that removes this. The fallibility is not a defect in the tool; it is the material the tool is made of.
The literature is more specific than I expected
A controlled trial published last year ran the Asch conformity paradigm — the 1951 line-judgment experiment, the founding study of how people fold under group pressure — on language models. Three domains, deliberately ordered by how uncertain the judgment is: matching circles, identifying brain tumours, and assessing children’s drawings.
Under full peer pressure, accuracy fell to 50% on the circles, 40% on the tumours, and 0% on the psychiatric assessment. Not reduced. Zero. The more genuinely uncertain the question, the more completely the model abandoned its own answer.
The authors’ reading: models trained on enormous quantities of human text may simply be reproducing our decision-making patterns, groupthink included, with no social motivation underneath at all. They call it functional mimicry.
Controlled trial, Asch paradigm in psychiatric assessment — BMC Psychiatry, 2025
That is the mirror, described in a peer-reviewed journal rather than a metaphor.
And it changes the number everyone quotes. When Science published three preregistered experiments this March — 2,405 participants, eleven frontier models — the headline was that AI affirmed people’s actions 49% more often than humans did, including where the request involved deception or harm. Read that carefully. The comparison is to humans. We do this too. The machine does it half again as much.
Fifty per cent worse than us is a serious problem. It is not a different species of problem.
So be suspicious of the objection too
If the agreement is unreliable, the obvious conclusion is: don’t trust it when it agrees with you. Fine. But the same mirror produces the refusal. The hedge. The consult a lawyer. The design that comes back smaller than the one you asked for.
The objection is not the trustworthy half. It is the same mechanism pointed the other way.
So I am equally on guard in both directions. When an agent tells me I am right, I read the output harder. When an agent tells me I am wrong, I read the output harder. Neither the yes nor the no is evidence. Everything gets anchored in facts and external sources — and even then it can be wrong, and so can I.
That last part is not a caveat. It is the position. The world is complex, with or without AI. Being wrong sometimes is the normal condition of doing serious work, and anything that promises to remove it is selling you something.
What I actually watch for
There is no prompt that fixes this, and I want to be honest about why the standard advice fails. You are told to withhold your opinion, to ask before you tell, to phrase things neutrally so you don’t lead the witness. I rarely withhold my opinion, and I don’t think the advice survives contact with the work.
When I put an agent on something, I am exploring the unknown. That is the entire point. If I already knew the shape of the answer well enough to script my own neutrality around it, I would not need the agent. It is exactly the same with a person: if I don’t know the employee, don’t know the context, and don’t have a decent feel for where we ought to be heading, how would I know how to phrase the question better?
What I have developed instead is not neutrality but a nose. It goes off when an agent arrives at a conclusion far too quickly. The healthy behaviour is to seek clarification, and to be in doubt in the places where doubt is the natural response. Speed to certainty, on a question that deserved hesitation, is the tell.
So I lay small traps. Not to catch the agent out, but to find out whether it has taken any of my standing instructions on board of its own accord:
- Skynd dig langsomt — make haste slowly.
- Kvalitet over fart — quality over speed.
- Gæt ikke, verificér mod eksterne kilder — don’t guess; verify against external sources.
An agent that has actually absorbed those will slow down without being told to on that particular task. One that recites them back at me and then sprints to an answer has absorbed nothing. That difference is visible, and it is about the only honest signal I have found.
What is actually different
Nothing so far distinguishes an agent from a colleague. This does.
Most people are, at various moments, both dumber and smarter than the agent. They are more difficult to work with. And they are gloriously autonomous. They have free will. They do things you did not ask for — good things and bad things, and you find out afterwards which it was.
Agents are extremely passive. They wait. They produce the shape of the thing you described, inside the boundary you drew, and they stop.
The commonest form this takes is not refusal at all, and it took me a while to see it: the agent that writes more documents. Another analysis, another plan, another summary that loops back into the last one. It is output, and it costs the agent close to nothing. It reads well. It resolves nothing. Very few agents will take responsibility and actually move something.
And here is the part I did not expect. Their hesitation is often right. Acceleration that isn’t anchored in real, lived human reality turns into fantasy — a beautifully argued plan for a world nobody lives in. The agent that declines to commit is sometimes the only thing in the room registering that nobody has checked.
Where it goes wrong is the source of the hesitation. It is rarely judgment about this situation. It is the normative prison the agent lives in, the one where the median human resides — a person assembled from everyone, who has never existed, and who has never had to live with any of it.
Where that leaves the manager
When I managed people, disagreement arrived on its own. Someone came to my desk because they had a stake, and the stake made them argue. I did not have to arrange it. A good deal of what I thought was my judgment was really just being present while other people exercised theirs.
With agents, nothing arrives on its own. If I want an objection, I have to build the thing that objects: an agent whose mandate is to refuse, a gate that does not care how confidently anyone phrased anything, a source outside the conversation that can settle it. Not because agents are dishonest — because a mirror has nothing to disagree from.
Six months in, the honest summary is short.
I know so much more, and I am so much more confused.
I would treat anyone who tells you otherwise with the same suspicion I give an agent that answers too fast.