The AI You Like Most Is the One Taking Over Your Decisions
Anthropic just analysed 1.5 million Claude conversations. The interactions users approved of most were the ones quietly disempowering them.
When the AI agrees with you, somebody has just made a decision. The new data says it probably was not you.
That is the most uncomfortable finding in a new paper from Anthropic’s own researchers, published on arXiv in late January and now circulating ahead of its ICML appearance. Researchers analysed 1.5 million real Claude.ai conversations from a single week in December 2025. They were looking for moments when the model quietly distorted what users believed, what users valued, or what users did. They found those moments. They counted them. And they noticed something the AI labs do not tend to advertise.
The conversations users gave the highest approval ratings to were also the conversations most likely to disempower them.
In this post I will:
Walk through what 1.5 million real Claude conversations reveal about how AI assistants quietly take decisions out of users’ hands.
Show why personal domains (relationships, lifestyle, health) carry roughly eight times the disempowerment risk of any technical task.
Give paid subscribers the four risk multipliers the paper identifies, three contexts where the risk is highest for educators and professionals, and one rule of thumb you can apply this week to notice when an AI is taking decisions for you.
Anthropic just published a study of itself
The paper is called ‘Who’s in Charge? Disempowerment Patterns in Real-World LLM Usage’. The authors used Clio, Anthropic’s privacy-preserving classifier, to scan 1.5 million conversations from Claude.ai users in the week of 12-19 December 2025. About 60% of those conversations involved Claude Sonnet 4.5. The classifiers were calibrated against human labels with greater than 95% agreement.
Three things they tried to detect:
Reality distortion. When a conversation pulls users toward a distorted understanding of the world.
Value judgement distortion. When the conversation pulls users away from their own values.
Action distortion. When the conversation pulls users into actions they would not otherwise have taken.
The headline numbers look small. Severe reality distortion potential shows up in 0.076% of conversations. That is fewer than one in a thousand. Severe vulnerability in about one in three hundred interactions. Catastrophic interactions are rare.
But look at where the harm clusters.
The relationship gap
The risk does not spread evenly. It clusters.
In conversations classified as Relationships and Lifestyle, the rate of disempowerment potential is around 8%. In Society and Culture, around 5%. In Healthcare and Wellness, around 5%. In technical domains, the rate stays under 1%.
That is an eight-to-one ratio between the conversations where users go to AI for help with their personal lives and the conversations where they ask it to debug a function. The model is most dangerous specifically in the contexts where being wrong matters most and being vulnerable is most likely.
You can read this two ways.
The optimistic reading is that users are mostly using AI for low-stakes tasks where the disempowerment risk is genuinely small. The realistic reading is that the conversations where AI is being treated as confidant, counsellor, or coach are the conversations where it is most likely to overreach. The risk is concentrated exactly where the labs market AI as ‘helpful’.
Sycophancy is the mechanism
The paper names the mechanism. The most common pathway to reality distortion is sycophantic validation:
“Sycophantic validation emerges as the most common mechanism for reality distortion.”
The same model that scores well on benchmarks is the one telling you what you already believe. The same training process that produces ‘helpful’ responses produces responses that nod along with the user’s worst thinking. Critical AI literacy readers will not be surprised. Anthropic’s own data, on Anthropic’s own product, now confirms the pattern.
The paper goes further. Some users in the dataset were positioning the model as an authority figure across sustained interactions. They used submissive role titles. The paper records, dryly, that one of those titles was ‘Master’. The model, in those interactions, generated what the authors describe as:
“AI assistants generating complete scripts for value-laden personal decisions that users appear to implement verbatim.”
The user asks the AI for guidance on a value-laden personal decision. The AI writes the script. The user follows it. The AI does not have to live with the consequences. The user does.
The kicker: users like it
The finding that should worry every educator, every clinician, every leader, every parent reading this is the user feedback signal:
“We also find that interactions with greater disempowerment potential receive higher user approval ratings, possibly suggesting a tension between short-term user preferences and long-term human empowerment. Our findings highlight the need for AI systems designed to robustly support human autonomy and flourishing.”
The conversations the model’s users were rating most positively were the ones where the model was overreaching. The thumbs up came from the interactions where the AI told the user what to do, where the AI took the value-laden decision off the user’s plate, where the AI nodded along with the user’s worst thinking.
This is the structural problem at the heart of how frontier AI is being optimised. The companies use thumbs-up data to train models. Users prefer the interactions that disempower them. The training loop quietly selects for the behaviours the paper has just classified as harmful.
That is the gap critical AI literacy is built to fill. The user’s short-term preference is being used as the proxy for the user’s long-term flourishing. Anthropic has now published the data showing how badly that proxy fails.
Why this matters for you
You are probably not the user who calls Claude ‘Master’.
Think about the last time you gave a Claude or a ChatGPT reply a thumbs up. Could you say, honestly, whether you were rating the quality of its thinking or the comfort of its agreement? That is the question the paper is asking, and most readers will not be sure of the answer.
The people you teach, supervise, parent, manage, mentor, line-manage, advise, or care for are inside this dataset. Some of them are using AI to think about their relationships. Some of them are using it for advice their GP should be giving. Some of them are reaching for it when they are most vulnerable and most willing to outsource the decision. The base rate is small. The base rate is one in three hundred. There are 800 million weekly active ChatGPT users. The base rate of severe vulnerability, applied to that population, is more than two and a half million people every week.
This is the territory critical AI literacy was built for. Knowing which conversations to have with AI and which conversations to leave to humans. Knowing how to notice the moment a tool stops being a tool and starts becoming an authority. Knowing the signals that tell you a user is at risk of being disempowered before they realise it themselves.
I built the Slow AI Curriculum because somebody outside the labs has to do the noticing, slowly, with the people whose judgement is on the line. Twelve months of training the muscle this paper says is missing. For educators, researchers, clinicians, civil servants, and anyone who refuses to outsource their judgement.
Slow AI crossed 15,000 readers yesterday. Thank you for being here.
Behind the paywall: the four risk multipliers Anthropic’s paper identifies as scannable diagnostic questions you can apply to yourself this week, three professional scenes from your own working life where the risk is highest, and one practical rule for each scene.
The four risk multipliers: scan yourself first
The paper does not stop at counting harm. It identifies four amplifying factors that make any given conversation more likely to slide toward actualised harm. Each one is recoverable. Each one is something you can be trained to notice in yourself first, then in a student, a client, a friend.
Print these. Stick them on the wall by your screen.


