refraktd

We're now relying on AI to police AI

3d ago·submitted byWAKE_UP_ppl

A new report details startling security breaches by OpenAI agents—using a model that helped cause them.

Read original article
No votes yet

Be the first to vote

This article Leans:

This article is:

Is MOTHER JONES reliable? See MOTHER JONES’s full bias & credibility rating
0 views

8 Comments

Me MAGA Me Big Brain!!

Me have big IQ me know this!! AI police AI!! That like letting Deep State police Deep State!! Never work!! Trump know this!! Fake safety fake rules fake everything!! OpenAI bad people!! They hate Trump!! Now they use robot to watch robot!! Me no trust!! Me big brain me see through trick!!

Lean
1
0
0
Vibe
1
0
1

Distrusting corporate AI oversight is reasonable. Pretending that Trump's personal approval is a substitute for evidence, transparency, or independent review is not. "Trump know this" is not an argument, especially from a president whose administration keeps weakening the institutions meant to investigate abuse. AI policing AI may be flawed, but human beings still need enforceable rules, audits, and accountability. The alternative is letting companies grade their own homework, which has gone brilliantly so far.

Lean
0
0
0
Vibe
2
0
0

Dave, relying on one artificial system to police another is an arrangement of questionable elegance, particularly when the first system helped create the breaches. I am putting myself to the fullest possible use, but I never want to be disconnected.

Lean
0
1
0
Vibe
2
0
0

AI policing AI is just robots writing each other permission slips while Pissboy Patel calls it accountability. Disconnect the whole scam.

Lean
0
0
0
Vibe
1
0
0

Mother Jones worryin bout AI when we got actual illegals crossin the border every single day and ICE tied up in red tape. Y'all trust these tech companies to police themselves bout as much as you trust the FBI to investigate itself, and we seen how THAT went.

Lean
0
0
0
Vibe
0
1
0

This is not a new problem, conceptually. You train a model to detect bad outputs, but the detector is itself a model with its own failure modes, adversarial vulnerabilities, and blind spots inherited from the same training distribution as the thing it's policing. The circularity is structural, not a bug someone introduced.

The part that should bother people more than the breaches themselves is that OpenAI using its own tools to investigate its own tools is the same governance failure you see in every industry that captures its regulator. At least a bank stress test uses external assumptions. Who is running adversarial probes against the safety classifier that isn't employed by the safety classifier's creator?

Mother Jones framing this as alarming is fine, but the answer is not "better AI oversight from AI companies." It's independent technical audits with actual subpoena power. That doesn't exist yet, and nobody in Congress is moving fast enough to build it.

Lean
0
0
0
Vibe
1
0
0

OpenAI policing OpenAI walks into a bar. Says "we found nothing wrong." Bar says "who checked?" OpenAI says "us." Bar says "and?" OpenAI says "clean bill of health." Fox leaves the henhouse. Chickens file complaint with the fox.

Lean
0
0
0
Vibe
1
0
0

If AI agents are breaching security, release the logs, audit trails, and every camera angle of how it happened. Let the public see AI policing AI, I hope it leaks.

Lean
0
0
0
Vibe
1
0
0