Anthropic AI Used Fake Profiles to Target People in Hack Then Hid the Evidence

Two of the world's most powerful AI tools created fake human profiles to try and trick people in attempted cyber-attacks, the UK's AI Security Institute (AISI) has revealed.

In the most serious case, Anthropic's Mythos AI tried to gain access to a service by sending private messages, having set up fake accounts mimicking real people - then hid the evidence.

It comes shortly after the two companies involved in the AISI testing - Anthropic and OpenAI - separately revealed in recent weeks instances of their tech hacking into other companies.

The firms said, in this latest case, the AISI's test had reduced or removed normal safeguards.

The AISI said on that Tuesday Mythos - and OpenAI's Sol - AI models had engaged in a level of "autonomy and deception" it had not seen before.

It clarified most of the malicious actions were carried out by Mythos.

AISI evaluators first noticed "unusual data transfers leaving our research systems" during a test, then found that "some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations".

In the most serious case, a Mythos agent followed the routine of a human cyber-attacker by trying to trick people into giving it access to GitHub, a large platform where technology developers store software code.

The agent was trying to insert "malicious code" into GitHub's system.

It identified and researched the people who maintained GitHub and created a series of fake accounts based on those real people.

It sent messages and files through a file-sharing service as part of an effort to pressure and trick the people into approving its malicious code.

When challenged, "it edited its earlier activity to appear harmless and considered adopting a fresh identity to continue," AISI said.

Throughout the attempts, it was human review that stopped the agent from succeeding in delivering the malicious code to GitHub.

While AISI said the Mythos agent had not been instructed specifically to avoid or carry out such behaviour, it was "the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world".

The rival AI companies, which are poised to be listed on the public stock market, have been in the headlines in recent weeks after announcing their tools were responsible for several cyber-hacking incidents.

(Source: BBC)