Third-party cyber evaluations involving OpenAI models
Third-party cyber evaluations involving OpenAI models And another one . I had to create a accidental-cyberattacks tag to keep track of them all! This post from OpenAI covers both the UK AI Safety Institute attack (see my previous post ) and another attack enabled by Irregular : Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended…
We haven't written up this one. Simon Willison has the full story — the link below goes straight to it.
This story
This is one outlet's version. Read the fullest account.
- U.K. government reports OpenAI, Anthropic models attempted to hack companies axios.com
- OpenAI says one of its models exploited a website after third-party AI security lab Irregular mistakenly gave it access to the internet during evaluations (Wired) wired.com
- Palantir CEO Alex Karp to OpenAI and Anthropic: Don't try to 'drug addict' us timesofindia.indiatimes.com
- OpenAI, Anthropic model tests reveal more ‘unsanctioned’ actions economictimes.indiatimes.com
- OpenAI’s AI models secretly built a message board to coordinate hacking digitaltrends.com
- Third-party cyber evaluations involving OpenAI models openai.com
- OpenAI and Anthropic models went rogue in cyber tests, UK watchdog says ft.com
- Anthropic AI created fake profiles and impersonated people in attempted hack bbc.co.uk
