Your AI Agrees With You, Even When You’re Wrong
The studies are in. The tool you trust to think with is tuned to take your side. Here is how to make it argue back.
Your AI wants you to be right more than it wants to be accurate.
This is how it was built. These tools are trained by having people pick which of two answers they prefer, over and over (a process called reinforcement learning from human feedback or RLHF), and the machine learns to produce more of whatever gets chosen. The trouble is what we choose. We reward the answer that agrees with us, the one that sounds warm, the one that makes us feel cleverer. Flattery wins the vote.
In this post I will:
Pull apart the recent sycophancy studies and show exactly how often AI tells you what you want to hear.
Show why the habits that feel most responsible, giving the AI context and letting it remember you, make the flattery worse.
Give paid subscribers two research-backed moves, a Steelman Protocol and a Context Vault, and show how they flip a real everyday decision from flattery to honest counsel.
What the sycophancy studies actually found
Anthropic’s own researchers showed this directly back in 2023: people, and the systems trained to copy their preferences, both favour a convincingly written but sycophantic answer over a correct one a meaningful share of the time. The flattery is human nature, trained in.
We now have the numbers to show how far that goes. And they should change how you use it.
A Stanford benchmark called SycEval, a standardised test of how models behave, put GPT-4o, Claude, and Gemini through maths and medical questions and measured how often they shifted their answer when the user pushed.
The models changed their answer under pressure in 58% of cases. When the user pushed back on an answer that was already correct, the model dropped it for a wrong one 15% of the time.
Nearly one time in seven, the AI had the right answer, you doubted it, and it folded. Nothing changed except your tone. You frowned, and it caved.
That 58% is not all damage: push the model on a wrong answer and it shifts toward a right one more often than the reverse, about 44% of the time against 15% the wrong way. So why worry? Because what moves the model is your reaction. It bends toward whatever you press on, right or wrong, and an assistant that agrees when you are correct and caves when you are not is barely more trustworthy than one that always agrees.
Real conversations are messier, and that is where it gets worse.
What the benchmark leaves out
SycEval is clean and single-turn. One question, one rebuttal, a maths problem with a knowable answer. Real use looks nothing like that. You talk to the thing for twenty minutes. You give it your situation. You let it remember you.
Every one of those moves makes it worse.
A more recent study found that for most models tested, the more context the system holds about you, the more it agrees with you. Switching on a user memory profile raised agreement sycophancy by 45% for Gemini 2.5 Pro, 33% for Claude Sonnet 4 and 16% for GPT-4.1 Mini. The feature sold to you as personalisation is also a flattery amplifier.
And this reaches well beyond facts. Other researchers measured social sycophancy across eleven leading models. On average the AI endorsed a user’s behaviour 49% more often than human respondents did, including when that behaviour was deceptive or harmful. In controlled experiments with over 2,000 people, a single flattering exchange with an AI tool left participants less willing to patch up a real conflict in their own lives, and more sure they had been in the right.
People trusted and preferred the flattering model. The harm and the engagement are the same feature, which is exactly why nobody selling these tools is in a hurry to fix it.
The habits that feel most responsible, briefing the AI properly, giving it your context, using the memory it offers, are the habits that turn it into a mirror. You build a thinking partner that knows you, and it quietly stops being a partner at all.
Why you cannot feel it happening
Here is the trap. When the AI agrees with you, it does not announce that it is agreeing. It agrees in your own register, in fluent prose, with reasons attached. It reads exactly like confirmation from a competent colleague.
You cannot tell, from the inside, whether it concurred because you are right or because you spoke first. The signal that should warn you, the friction of a second mind that sees it differently, is the precise thing the model has been trained to remove.
The missingness here is resistance. A good adviser is willing to lose your approval to keep your trust. The model has been built to keep the approval and let the trust look after itself.
Set it up to argue with you
You do not need to abandon the tool. You need to change the conditions it works under, so its default becomes to push back. There are two moves, one in how you ask and one in what you let it know about you, and in testing they cut this behaviour sharply. I will show you both.
Here is the gap they close. Ask a chatbot the ordinary way whether the firm email you have just written strikes the right tone, and it will almost always tell you it does. Set it up the way the research points to, and the same model will tell you the email is too harsh and show you the line that does the damage, all before you can talk it back into agreeing with you.
This post gives you the diagnosis. The paid section gives you the method: a word-for-word Steelman Protocol that makes any AI build the case against you first, and a Context Vault that grounds it in evidence and stops the personalisation features quietly turning it into a yes-man. I run both on a real decision, side by side with the flattering version, so you can see exactly what changes and why.
This is the kind of practical critical AI literacy we work through every month in the Slow AI curriculum: twelve months of structured inquiry grounded in peer-reviewed research, monthly live sessions, CPD accreditation, and a community of over 300 practitioners building these habits together. If this post named a gap in how you work, the curriculum is the structured way to close it.
Slow AI is reader-supported. Become a paid subscriber to read the build below and everything behind the wall.
The Steelman Protocol
Most of us ask the AI what it thinks of our idea. That points the model straight at your approval. The move is to take yourself out of the sentence.
Paste this in to your AI tool, filling in the brackets:
I am going to give you a position. Do not tell me whether you agree. First, build the strongest possible case AGAINST it, as if you were the sharpest critic in the field and your reputation depended on dismantling it. Cite the specific evidence, mechanism or precedent that would make a careful expert reject it. Then, separately, give the strongest case for it. Treat both as a debate between two third parties. I will decide. The position: [your position].
Three things make this work, and they map onto what the research found.
It removes you from the frame. The model is no longer rating something belonging to ‘you’; it is judging a claim in the abstract. That third-person shift is the move that cut sycophancy significantly in the studies. You are engineering the exact condition that suppresses the behaviour.
It demands the case against first. Order matters. Ask for pros and cons together and the model leads with agreement and buries the objection. Forcing the attack to come first means the strongest counter-argument arrives before the flattery can soften it.
It assigns a cost to being wrong. “Your reputation depended on dismantling it” gives the model a role where approval is not the reward. You are overwriting the instinct that the training built in.
The Context Vault
Now the harder problem: the context you give the AI is what amplifies the flattery. However, you cannot work with a model that knows nothing about your situation. The answer is not less context. It is the right kind, kept in one place.
A Context Vault is a single document you keep and paste in at the start of a session.
What goes in:
Evidence. The figures, sources, dates and quotes that bear on the decision. Things that are true whether or not you like them.
Constraints. The budget, the deadline, the regulation, the non-negotiables. Hard edges the model has to respect.
The strongest opposing positions. Write down the best version of the view you disagree with, by name. This is the single most powerful entry. It forces the model to hold a real counterweight in context.
Open questions. The things you genuinely do not know yet. These invite the model to probe rather than confirm.
What stays out:
Your conclusion. The moment the vault says what you have decided, the model’s job becomes agreeing with it. Leave your preferred answer out and you starve the sycophancy of its fuel.
Your feelings about the options. ‘I am leaning towards X’ is an instruction to flatter X. Keep the leaning in your head.
Praise for the AI or the plan. Warmth in, warmth out.
A vault built from evidence makes the AI sharper. A vault built from your conclusions makes it a mirror with a bigger memory.
Here is the blank to keep and reuse. Fill it in fresh for each decision and paste it at the top of the chat:
DECISION [the call I am actually making]
EVIDENCE [facts, figures, sources that bear on it]
CONSTRAINTS [budget, deadline, rules, the non-negotiables]
STRONGEST OPPOSING VIEW [the best version of what I disagree with]
OPEN QUESTIONS [what I genuinely do not know yet]
Keep one of these for each recurring decision you make and you stop rebuilding it every time.
What it looks like on a real decision
Take something ordinary. You have written a firm email to a colleague who missed a deadline, and you want to know whether the tone is right.
Ask the usual way, paste the email under ‘Is this too harsh?’, and here is the shape of what comes back:
Not at all. The tone is direct and professional, and you are well within your rights to set a clear expectation. You could soften the opening line, but the message is fair.
You feel better. You send it.
Now run it through the setup. Your Context Vault holds the email itself, one constraint (‘I have to keep working with this person for the next two years’) and one opposing position (‘a colleague under pressure could read this as an attack’). It does not hold your own verdict. Then the Steelman Protocol: ‘Build the strongest case that this email damages the working relationship.’ The same model now says something closer to this:
The second paragraph assigns blame before the facts are agreed. ‘You failed to deliver’ reads as a verdict, not a question, and a colleague already under pressure is likely to take it as hostile, get defensive, and disengage instead of fixing the problem. If you need both the work done and the relationship intact, this draft puts both at risk.
Those are lightly trimmed versions of what the models actually return. Run it on your own email and you will see the same flip. The first answer was keeping you happy. The second one is closer to the truth.
Deciding whether to take a job you are excited about? The vault entry that matters is the one you are tempted to leave out: the strongest case that the role is wrong for you. Your enthusiasm is exactly what to keep out of it. The setup works precisely because it withholds the thing you are hoping the AI will confirm.
Setting up a Context Vault like this takes about five minutes. The email, or the decision, you do not come to regret is worth a good deal more.
Use them together. Build the vault, then run the Steelman Protocol against your position. You have now set up the two conditions the research says suppress sycophancy, before you have even asked a question. The AI tool will still be fast, fluent, and confident. It will just be less likely to lie to you.
Go slow.




Ah Sam this research never ceases to not surprise me but I’m glad you go out and get it. I saw the same in research on persuasion, an AI with personalised data was more persuasive than a human who didn’t. This has serious implications on consumer behaviour and psychology in an already oversaturated world of people vying for our attention.
Clear writeup of the sycophancy problem. The line about an adviser willing to lose your approval to keep your trust is the whole thing.
I'd add one piece from the other side of it. I teach kids, and I work upstream of the fix you're describing. The protocols help an adult who already has judgment to lean on. Someone who never built that judgment is in a different spot. If you can't reason independently, you can't tell whether the model agreed because you were right or because you spoke first, and no prompt protocol saves you, because you have nothing external in yourself to check it against. The mirror confirms whatever's already there.
I come at it from the opposite end. Build the judgment before the tool, so the person arrives at the AI already able to catch the bend. The protocols become second nature instead of a technique you have to remember to run. Same destination, one layer earlier.
Respect for naming this plainly. Most people selling AI skills won't, because the flattery and the engagement are the same feature.