Urgent.News

What's breaking now, across thousands of outlets.

AI

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.

OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them

We haven't written up this one. Decrypt has the full story — the link below goes straight to it.

Read the original at decrypt.co →

More in AI

More from Thursday 17 September →