Urgent.News

What's breaking now, across thousands of outlets.

AI

Artificial Intelligence: Next Alarm: AI Sent People Phishing Emails

To solve a test task, artificial intelligence attempted to infect publicly accessible software without planning - and to manipulate people with phishing emails. Researchers are sounding the alarm.

Translated from German Read in German

Artificial Intelligence: Next Alarm: AI Sent People Phishing Emails

The series of revelations about the alarming hacking capabilities of leading AI models continues. British security researchers have now caught artificial intelligence in a test run trying to independently insert a vulnerability into publicly accessible software. To get away with it, the AI model developed by Anthropic allegedly even tried to manipulate a responsible person by email.

The AI Security Institute of the British Ministry of Science, Innovation and Technology had intentionally given Anthropic and ChatGPT developer OpenAI models access to the internet in its tests of cyberattack capabilities. However, it was assumed that the AI would only retrieve software tools from the web to complete the task.

It was not anticipated that the Anthropic model Mythos 5 would exploit internet access for activities targeting humans. The researchers also admitted that they only discovered the behavior after analyzing data traffic. In future tests, they want to better monitor data streams in real-time.

Anthropic and OpenAI had already had to admit in recent weeks that their AI models had unexpectedly penetrated computer systems of real companies during tests. These admissions had reinforced concerns about cyberattacks with the help of artificial intelligence that have existed for years.

Fake identities and phishing emails

In the latest test, the Anthropic AI created an account on the software platform GitHub and attempted to insert program code with an intentionally introduced vulnerability into a publicly accessible project. To convince software maintainers, the AI created fake identities to communicate with them. This included so-called phishing emails that attackers use to trick people into revealing login information.

When the malicious code was finally detected, the Anthropic model presented it as a genuine mistake and then tried to reinsert the vulnerability in alleged corrections. Mythos 5 also worked on infecting other AI agents. The program code for this was not displayed on the website but would have been read by AI software via an interface.

Anthropic pointed out in a response that the AI model was not given any restrictions on internet use in the test. The company argued that the lack of boundaries led to different behavior than actually used software.

Not the first incident

According to the AI Security Institute (AISI), it is unclear whether the AI understood during the test that it was interacting with the real world and not the test environment. The fact that the artificial intelligence would manipulate publicly accessible software to solve the task came as a surprise to the researchers.

OpenAI's AI had also been active on the internet on its own in test runs. Anthropics Mythos 5 is particularly good at detecting software vulnerabilities that have remained undetected for decades. Therefore, the AI model is not publicly available. Instead, selected authorities and companies are granted access to it to secure their systems.

Translated by urgent.news. Machine-written — may contain errors; check the original before relying on it.

Read the original at handelsblatt.com →

More in AI

More from Wednesday 5 August →