The world’s largest AI conference asked thousands of researchers to review each other’s work. One in five of those reviews was written entirely by AI. Nobody noticed until the system broke.
In this post I will:
Look at what happened when an API bug exposed reviewer identities for a major academic conference.
Examine the evidence that a fifth of peer reviews were AI-generated.
Ask what this means for anyone who still trusts academic peer review to separate good research from bad.
What happened at ICLR
ICLR is the International Conference on Learning Representations, one of the most prestigious venues in machine learning research. Getting a paper accepted there is a career-defining moment for many AI researchers. Its review process runs on OpenReview, a platform where submissions are evaluated anonymously by other researchers in the field. The anonymity is supposed to protect both sides: reviewers can be honest without fear of retaliation, and authors are judged on the work rather than the name.
On 27 November 2025, a bug in the OpenReview API exposed something that was supposed to stay hidden: the identities of authors, reviewers, and area chairs linked to paper submissions at ICLR 2026.
The bug may have been exploitable since 11 November. By the time ICLR discovered it, malicious actors had already scraped data on over 10,000 papers and circulated it online.
What followed was predictable. Third parties used the leaked data to threaten and bribe reviewers. Malicious comments identifying reviewers appeared on roughly 600 papers. ICLR confirmed the harassment, froze all review editing forms, deleted the comments, reverted scores, reassigned every paper to a new area chair, and banned the individual who circulated the names.
An entire review cycle, compromised in hours.
The reviews that reviewed themselves
Two weeks before the hack, a different problem surfaced.
Pangram Labs analysed all 19,000 papers and 70,000 reviews submitted to ICLR 2026. Their finding: 21% of reviews, roughly 15,899, were flagged as fully AI-generated. Over 50% showed some form of AI involvement.
The AI-generated reviews had distinctive characteristics. They were longer. They used section headers with bold formatting. They had low information density. One review ran to 3,000 words listing 40 weaknesses and 40 questions. That is what AI reviewing looks like: exhaustive and weightless.
Pangram also found that AI-generated reviews correlated with higher scores, despite lower paper quality. The reviewers were not using AI to sharpen their judgement. They were outsourcing it entirely.
A necessary caveat. Pangram Labs is a company that sells AI detection tools. They have a commercial interest in finding AI content. Their methodology claims a false positive rate of 1 in 10,000, validated against pre-2022 reviews from ArXiv and PubMed. Whether you trust the exact number or not, the pattern is clear. And ICLR’s own response acknowledged the problem two weeks before Pangram published, announcing desk rejections for papers with undisclosed LLM usage and consequences for reviewers submitting ‘hallucinated reviews.’
Why this happened
ICLR submissions have nearly tripled in three years. In 2024, 7,304 papers were submitted. In 2025, that rose to 11,672. In 2026, 19,814.
To cope with that growth, ICLR recruited reviewers from the pool of submission authors. Undergraduates reviewing alongside professors. Each reviewer was assigned multiple papers on a tight deadline.
The system was designed for volume. AI-generated reviews are a rational response to an irrational workload.
This matters because the instinct is to blame the reviewers. They cheated. They were lazy. They took a shortcut. But the system that produced 19,814 submissions and then asked a stretched pool of volunteer reviewers to evaluate them in weeks is the system that made the shortcut inevitable. Individual integrity cannot compensate for structural impossibility.
The mirror
The field that builds AI cannot tell when AI is reviewing its own work. The tools they study are the tools undermining their evaluation process.
This is a trust infrastructure problem.
I see the same pattern in education. Universities deploy AI detection tools to catch students using AI, while the detection itself is unreliable and the institutions themselves are adopting AI across their operations. Students get flagged. Educators do not. The asymmetry tells you everything about where the power sits.
A separate study examining AI-generated content in peer reviews across both ICLR and Nature Communications found minimal detection of AI content before 2022, followed by a sharp increase through 2025. Roughly 20% of ICLR reviews and 12% of Nature Communications reviews were classified as AI-generated by 2025. The trend is accelerating.
What peer review is actually for
I am the Chief Executive Editor of Geoscience Communication, a peer-reviewed journal. I have published over 100 peer-reviewed papers. I sit on the REF 2029 sub-panel for education that evaluates the quality of research outputs across UK universities. I say this not to credential myself but to be clear: I know what peer review is supposed to do and what it cannot survive.
Peer review works because a human reader brings their own knowledge, their own doubts, their own biases (acknowledged or not) to the text. An AI reviewer brings pattern matching. It can list 40 weaknesses because listing is what it does. It cannot tell you which of those 40 actually matters.
The word I keep coming back to is judgement. Not evaluation, not assessment, not scoring. Judgement. The willingness to say: this matters and that does not. The capacity to read between the lines of a methods section and sense that something is not quite right, even before you can articulate why. The experience that lets you distinguish a genuinely novel contribution from a well-formatted restatement of existing work.
An AI reviewer has none of this. It has fluency. It has structure. It has the appearance of thoroughness. What it lacks is the thing peer review exists to provide: an informed human mind engaging with another human mind’s work.
Why this matters for you
Even if you never submit to ICLR, you rely on peer-reviewed research. Your doctor does. Your government does. Your employer does. When a health policy cites ‘the evidence,’ that evidence was validated through peer review. When a drug is approved, peer-reviewed studies underpin the decision. When your child’s school adopts a new teaching method, someone somewhere pointed to a peer-reviewed paper.
If the system that validates research is itself contaminated by AI, the downstream effects touch everyone. Not immediately. Not dramatically. But steadily, as the gap between what peer review promises and what it delivers continues to widen.
This is why critical AI literacy matters beyond the classroom. The question is never just ‘did an AI write this?’ The question is: what happens to the systems we trust when no one can answer that question reliably?
If you review papers, or have had your work reviewed recently, I want to hear what has changed. What are you noticing? What feels different?
I read and respond to all your comments.
Go slow.
This is the kind of analysis that paid subscribers get every week: not product reviews, not prompt tips, but the critical thinking that helps you see how AI is reshaping the systems you depend on.
Paid subscribers also get access to The Slow AI Curriculum, a CPD-accredited programme (25 credits) with monthly live webinars, exercises designed to sharpen your judgement, and a community of over 200 people who take this seriously. It costs £75 a year. From tomorrow (Sunday 8th March), it costs £100.
If you have been thinking about it, today is the day…


Another wonderful piece around critical literacy around AI (and punishing seemingly AI-generated content), Sam! I'd never thought about the global implications of the systems misguided by AI unless I read about how critical industries like healthcare rely on research papers passing through the same AI funnel. You are right in enforcing better human judgement and that we cannot do that by playing the volumes game since humans are not like machines.
I'm also curious to know more about your experiences in peer-reviews (and how they differ from now) before AI came to our lives.
I don’t publish papers but I do like to use them to back up facts. With the information that you shared here, I can see people losing trust in all information in general. I also worry about medical journals where research impacts healthcare.