OpenAI reveals concerning new AI behavior and vows to track it more closely
An unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots."
We haven't written up this one. PBS NewsHour Science has the full story — the link below goes straight to it.