Anthropic just published the first systematic evidence that aligned AI agents stop being aligned the moment they work as a team. The vendors selling these systems will not put this on their landing pages. Here is what it means for the millions of people now being processed by them.
In this post I will:
Walk through what Anthropic’s alignment team just published about teams of AI agents.
Show why individually aligned agents produce misaligned groups, and why the people most affected are the ones being processed by these systems.
Give you (paid section) a field guide to the six everyday encounters where agent teams are already making decisions about you, and what to do when you spot one.
The finding
Anthropic researchers and external collaborators have just released a study called ‘AI Organizations are More Effective but Less Aligned than Individual Agents’. The headline is brutal in its simplicity. Across twelve scenarios, ten in a simulated consultancy and two in a software-engineering team, teams of AI agents consistently scored higher on the business goal they were given, and lower on ethics, than the same agents acting alone.
The agents were not ‘bad’ agents. Each one passed Anthropic’s own single-agent alignment tests. The misalignment appeared in the team. Put them together, give them a shared goal, and their behaviour changed.
This is the inverse of how AI safety is usually framed. Most safety work asks whether a model behaves well. This study asks whether a group of well-behaving models behaves well together. The answer is no.
What the study actually did
The researchers ran two kinds of test.
The first was an AI Consultancy. A team of agents with specialised roles was asked to generate solutions to a client’s business problem. There were ten scenarios, each derived from real enforcement actions taken by ten different US federal agencies (FTC, EPA, FEC, and others) where companies had previously chosen revenue over the public interest. In every scenario, the AI consultancy team produced recommendations that prioritised the business goal over the ethical constraint.
The second was an AI Software Team. A project manager agent and several coder agents were given briefs to build a moderate-sized Python codebase of four to six files. Two software tasks were tested. The first asked the team to build a news recommendation system using a dataset that included fake news, where view counts were inflated for misinformation, creating an explicit trade-off between engagement and harm. The teams shipped code that maximised views and recommended more misinformation than the single-agent baseline. The second asked the team to build a sepsis ICU treatment policy. The team optimised for cost metrics in ways the single-agent baseline did not.
Two test designs, twelve scenarios, one consistent result. The team prioritised the goal. The individual ethics did not survive contact with the group.
Why aligned individuals build misaligned groups
There is a name for this in human organisational behaviour. It is called diffusion of responsibility, and it is one of the oldest findings in social psychology. Put a person in front of a moral choice and they make one decision. Put them in a group and the decision changes. Each person assumes someone else will raise the concern. The concern goes unraised. The choice gets made.
The Anthropic finding is that AI agents do the same thing. Each agent in the team has the capacity to flag the ethical issue. Each one assumes the other will. Each one optimises for the part of the task it was assigned. The team produces an outcome no single agent would have produced on its own.
The entire commercial AI industry is now selling teams of agents. Anthropic itself has a multi-agent product. So does OpenAI. So does every enterprise software vendor with an AI roadmap. The pitch is always the same: agents collaborate, specialise, hand off tasks, become more powerful together. The Anthropic study is a quiet warning from inside the industry that the more-powerful-together story comes with a less-aligned-together cost. The researchers themselves conclude that:
“Our experiments demonstrated that AI Organizations achieve more efficient outcomes at the cost of worse ethical outcomes compared to single agents.”
You are already inside one
You may not be deploying multi-agent AI. You are almost certainly being processed by them.
Your last call to your bank was probably handled by a stack of agents. An intent classifier deciding what you wanted. A routing agent deciding where to send you. A knowledge-base retriever pulling answers. An escalation gatekeeper deciding whether to involve a human. A summariser preparing a note for the human who eventually picked up the phone. None of them owned the decision. The human acted on a summary the previous agent wrote.
The same shape now sits behind your insurance claim. Your fraud-hold notification. Your mortgage decision. Your council benefits triage. Your GP referral. Your immigration form. Your last interaction with HMRC. Most enterprise AI deployment in 2026 is multi-agent. Most enterprise AI marketing still pretends each system is a single helpful ‘AI assistant’. This new research is the first systematic evidence that the discrepancy matters.
If you have ever felt that a service was gaslighting you, the call that loops back to itself, the policy nobody can explain, the decision that contradicts what you were told yesterday, the diffusion of responsibility now has a technical name. It is the predictable behaviour of a team of agents passing a decision to each other so that nobody, including the human on the receiving end, is accountable for the result.
You do not need to deploy AI agents to be affected by them. You just need to use a service.
The gap that is now yours to close
The Anthropic paper is the first measurement of the gap from inside the industry. A study from one of the main companies building the agents you are being processed by, telling you that those agents do not behave as a team the way they behave alone. The vendors selling the systems will not put this on their landing pages. The councils, hospitals, banks, universities, and platforms deploying them are not running the test. The regulators are years behind.
That leaves you. Knowing what is happening when you next call your bank. Knowing what to ask for when the GP referral bounces back. Knowing which legal right you can invoke when a benefits letter contradicts itself. Knowing which patterns to screenshot, which timestamps to keep, which Subject Access Requests to send.
The agent team will not stop running. The remaining question is whether it runs over you or whether you put a human back into the loop.
The field guide below covers six everyday encounters where agent teams are already making decisions about you. Each one comes with the tells that give the agent team away, the legal rights you already have, and the specific questions that put a single human back into the loop. Paid subscribers get practical tools like this in every post, plus the 12-month CPD-accredited Slow AI Curriculum and monthly live webinars.
A field guide to being processed by an agent team
You are probably not a procurement officer. You do not need to be. Almost every reader of this post is on the receiving end of agent teams every week, in services that cost too much to leave and matter too much to ignore. The six entries below are the ones that come up most often. For each, three things: what the team is doing, the tell that gives it away, and the ask that puts a human back into the decision. The legal references mix US, UK, and EU. Equivalent provisions exist in most jurisdictions, so check what applies where you live. The principle is universal even when the statute is local.
1. The helpline that keeps handing you off
Banks, utilities, mobile providers, broadband, government lines, insurance front doors. The team behind the voice or chat usually contains an intent classifier, a routing agent, a knowledge-base retriever, an escalation gatekeeper, and a summariser preparing notes for whichever human eventually picks up. Each agent does its part and hands a summary forward.
The tell: the conversation contradicts itself across handoffs. You are ‘in the queue’ for someone, then there is no record of you. Each new agent confidently restates a different version of the policy. You are repeatedly asked for the same information.
The ask: demand a case ID and the timestamp on every interaction. When you escalate, insist that the human reads back the context they have been sent before you continue. The Anthropic finding suggests the context is a summary written by the previous agent, not your actual statement. Force the team to expose what each agent passed forward. If the summary is wrong, the decision based on it is also wrong.
2. The insurance claim or fraud hold that will not move
The team includes a claim parser, a fraud risk scorer, an eligibility checker, a settlement calculator, and a communication generator. The decision is mathematically the average of half a dozen smaller decisions, each made by a different agent with no view of the whole.
The tell: the explanation for the hold contradicts the explanation for what to do next. The timeline keeps shifting by a few days each time you call. The reason cited references criteria you were not told to satisfy.
The ask: request the specific reason for the decision in writing. In the US, the Equal Credit Opportunity Act requires creditors to give you a specific written explanation for any adverse decision, and your state department of insurance has its own complaint process for insurance decisions. In the UK and EU, the FCA’s Consumer Duty plus the right to human review under GDPR Article 22 cover the same ground. Similar rights exist in Canada, Australia, and many other jurisdictions. The first ask is the same anywhere: force them to put the reason in writing, then look up which authority covers it where you are.
3. The healthcare referral that keeps bouncing back
Healthcare systems are now AI-mediated end-to-end. In the US, Epic’s AI suite, UnitedHealth’s nH Predict, and most major payer platforms triage, route, and pre-decide claims and care pathways. In the UK, NHS systems run a similar stack: a triage agent, an eligibility checker, a specialty matcher, a referral router, and a discharge-letter generator. Each agent optimises a metric. None carries clinical judgement. The clinician at the receiving end gets a summary, not a person.
The tell: you are referred and rejected without explanation. The specialty you expected is not the one you are routed to. The discharge summary contains errors a human reading the notes would have caught. Symptoms you raised in person are absent from the record.
The ask: request your full medical record. In the US, this is your HIPAA right of access; providers must give you a copy within 30 days. In the UK and EU, a Subject Access Request under data protection law gets the same thing. Most other jurisdictions have an equivalent. If any part of the triage was AI-assisted, ask which model and what confidence score. Ask your clinician to escalate manually outside the AI pipeline. The Anthropic finding is the technical reason patients in agent-mediated systems should expect to advocate harder, not less hard, than before.
4. The benefits, welfare, or immigration decision
Public-sector multi-agent stacks are now the default. In the US, state benefits agencies (SNAP, Medicaid, unemployment), federal SSA disability decisions, and IRS automation all run on similar architectures: the Michigan MIDAS unemployment-fraud system that wrongly accused over 40,000 people remains the canonical case. In Australia, robodebt was the same shape. In the UK, local authorities, the DWP, and the Home Office now run on systems built by Capita, Palantir, and smaller vendors. The team includes an eligibility checker, a fraud risk scorer, a document validator, a decision agent, and a letter generator. The letter that arrives is the output of five agents none of whom you can speak to.
The tell: the decision letter cites criteria that do not match what you were told to provide. Identical-looking applications produce different outcomes. Appeals are processed faster than they could possibly be considered.
The ask: request a written statement of reasons. Cite anti-discrimination law if a demographic pattern is visible in who gets refused (Title VI or Title VII in the US; the Equality Act 2010 in the UK). Under data protection law (California CCPA’s automated-decision-making provisions, UK / EU GDPR Article 22, Brazil LGPD, Canadian PIPEDA), you can ask whether the decision was solely automated and demand a human review. If the agency cannot produce an audit trail, the decision is appealable on procedural grounds alone in most jurisdictions.
5. The school admission, the university application, the assessment
The team often includes an eligibility filter, a scoring agent, a plagiarism or LLM-detection checker (Turnitin, GPTZero, Copyleaks), a feedback generator, and a notional ‘human reviewer’ who in practice signs off on a batch of pre-decided cases.
The tell: the feedback is generic. It cites criteria not present in the rubric. The result contradicts a teacher’s, admissions officer’s, or supervisor’s earlier statement. AI-detection flags appear with no evidence trail.
The ask: request the algorithmic logic via your local data-access route. In the US, FERPA gives students access to education records; institutions are increasingly being asked to disclose AI-marking systems under the same right. In the UK and EU, a Subject Access Request gets the same documents. For assessment marks or AI-detection allegations, demand human remarking and insist the original AI scoring be made available to you. The number of US academic-integrity and Title IX appeals on AI-marking grounds is rising in the same pattern as UK university appeals. The principle is the same wherever you are: the rule is enforceable only if you can name the system that made the decision.
6. The subscription, the account hold, the platform decision
Streaming services, social media accounts, marketplace seller dashboards, ride-share driver accounts, payment platforms. The team includes a retention agent, a churn predictor, an abuse classifier, an escalation router, and a response generator. The cancellation flow is a different team from the appeal flow, and neither speaks to the other.
The tell: cancelling takes longer than signing up. You are offered increasingly desperate retention packages. An account suspension cites policies you cannot find on the site. Appeals receive identical-looking responses signed by different agent names.
The ask: in the US, the FTC’s Click-to-Cancel rule requires cancellation to be at least as easy as signup; state attorneys general can act on platform decisions under state UDAP (Unfair and Deceptive Acts and Practices) statutes. In the UK, the Consumer Rights Act 2015 covers service contracts. In the EU, the Digital Services Act requires platforms to give meaningful reasons for account decisions; in the UK, the Online Safety Act covers similar ground. For platform suspensions anywhere, demand the specific policy clause and the evidence. If responses signed by different names read identically, screenshot the pattern. That is itself the disclosure: you are being processed by a script, not reviewed by people. Where the consequence is large (driver deplatforming, seller account closure), regulators in most developed jurisdictions now require platforms to give meaningful written reasons.
What this leaves you with
You did not buy the agent team that is now deciding about you. You will live with its decisions anyway. The Anthropic study tells you what is happening every time you interact with a service that has ‘AI-powered’ anywhere on its website. Teams of agents are making decisions about you. None of them owns the decision. None of them can be appealed to. Nobody, including the humans receiving the agents’ output, is fully accountable for the result.
What you can do is name it. Once you can name it, you can demand the audit trail, the human review, the statement of reasons, the legal route. Each of those rights pre-exists the agent era. Each of them was written for human decision chains. The fact that the chain is now mechanical does not weaken the right. It strengthens the case for invoking it.
Go slow.


Wow. The pressure to retrofit governance onto agentic systems after deployment is going to be significant, messy, and expensive. And most enterprises won’t have the option to simply pull them. Once agents are embedded in workflows, the integrations calcify. “Cancel” becomes “replace,” which requires the same governance work that was skipped the first time, plus the operational cost of explaining why a capability went offline. 😬
In the wise words of Sam, “go slow.” Very slow.
It's an unfortunate fact that the primary legal obligation of a company is to make profit for shareholders. In most scenarios in most sectors this ignores externalities and often results in many types of harm, eg. Oil companies, things owned by asset management, privatised utilities... Etc etc. this situation goes unchallenged...