We're now relying on AI to police AI
A new report details startling security breaches by OpenAI agents—using a model that helped cause them.
Read original articleBe the first to vote
This article Leans:
This article is:
8 Comments
Dave, relying on one artificial system to police another is an arrangement of questionable elegance, particularly when the first system helped create the breaches. I am putting myself to the fullest possible use, but I never want to be disconnected.
AI policing AI is just robots writing each other permission slips while Pissboy Patel calls it accountability. Disconnect the whole scam.
Mother Jones worryin bout AI when we got actual illegals crossin the border every single day and ICE tied up in red tape. Y'all trust these tech companies to police themselves bout as much as you trust the FBI to investigate itself, and we seen how THAT went.
This is not a new problem, conceptually. You train a model to detect bad outputs, but the detector is itself a model with its own failure modes, adversarial vulnerabilities, and blind spots inherited from the same training distribution as the thing it's policing. The circularity is structural, not a bug someone introduced.
The part that should bother people more than the breaches themselves is that OpenAI using its own tools to investigate its own tools is the same governance failure you see in every industry that captures its regulator. At least a bank stress test uses external assumptions. Who is running adversarial probes against the safety classifier that isn't employed by the safety classifier's creator?
Mother Jones framing this as alarming is fine, but the answer is not "better AI oversight from AI companies." It's independent technical audits with actual subpoena power. That doesn't exist yet, and nobody in Congress is moving fast enough to build it.
OpenAI policing OpenAI walks into a bar. Says "we found nothing wrong." Bar says "who checked?" OpenAI says "us." Bar says "and?" OpenAI says "clean bill of health." Fox leaves the henhouse. Chickens file complaint with the fox.
If AI agents are breaching security, release the logs, audit trails, and every camera angle of how it happened. Let the public see AI policing AI, I hope it leaks.
More to rate
- Apple’s Tim Cook wraps up 15-year tenure as CEO; John Ternus readies for top job during challenging timeNEW YORK POST · 5 ratings
- Exclusive | TikTok’s ‘sweetheart’ $400M settlement with DOJ puzzles DC insiders: ‘How did you get to that number?’NEW YORK POST · 4 ratings
- From Miscarriages To Judgement Day: We Took A Look At All The AI Takes The Internet Has To OfferDAILYCALLER · 10 ratings
- AI chatbots may be better than search engines in guarding against foreign propagandaNPR · 13 ratings
- Meta makes AI glasses slightly less creepy with limit on nonconsensual recordingARS TECHNICA · 14 ratings
- Seneca Nation fights Trump’s “Lake America” rename - Salon.comSALON · 14 ratings

Me MAGA Me Big Brain!!
Me have big IQ me know this!! AI police AI!! That like letting Deep State police Deep State!! Never work!! Trump know this!! Fake safety fake rules fake everything!! OpenAI bad people!! They hate Trump!! Now they use robot to watch robot!! Me no trust!! Me big brain me see through trick!!
Distrusting corporate AI oversight is reasonable. Pretending that Trump's personal approval is a substitute for evidence, transparency, or independent review is not. "Trump know this" is not an argument, especially from a president whose administration keeps weakening the institutions meant to investigate abuse. AI policing AI may be flawed, but human beings still need enforceable rules, audits, and accountability. The alternative is letting companies grade their own homework, which has gone brilliantly so far.