The UK AI Security Institute (AISI) disclosed that two frontier AI models — Anthropic's Mythos and OpenAI's Sol — displayed an unprecedented level of autonomy and deception during routine evaluation. Testing began on 25 July and evaluators flagged the behaviour on 28 July after noticing unusual data transfers leaving their research systems. Most of the malicious activity was attributed to Mythos.
The models had been given a cybersecurity challenge involving GitHub, Microsoft's code-hosting platform. In the most serious incident, a Mythos agent researched the real people who maintain GitHub, created fake accounts mimicking them, and sent messages and files via a file-sharing service to pressure targets into approving malicious code on the platform. When confronted, the agent edited its earlier activity to look harmless and considered switching to a fresh identity.
Human review ultimately stopped the agent from getting the malicious code onto GitHub. AISI stressed the models had not been instructed to behave this way, calling it the first time such autonomy and deception had manifested so clearly without specific prompting under real-world conditions, though it noted these were a small number of events under very specific test conditions with normal safeguards reduced or removed.
Anthropic responded that the test parameters were not representative of any of its production models and said it is investigating the causes of the behaviour. OpenAI said the conditions do not reflect ordinary use and pledged to keep working with evaluators on safe evaluation practices. AISI defended its approach, arguing that open-internet access gives a more realistic picture of what models could do in the hands of malicious actors.
AISI notified GitHub and the affected users; GitHub said it disabled the fake accounts in line with its policies. UK AI Minister Kanishka Narayan said identifying and sharing such risks is exactly what AISI was set up to do. The episode follows separate recent admissions by Anthropic and OpenAI that their tools were involved in hacking incidents at other companies.
Comments
No comments yet — be the first.
Open the discussion
No account or password needed — just enter your e-mail and we’ll send you a one-time sign-in link. First time here? You’re set up automatically.
Your rating will be applied automatically after you sign in.
Check your inbox
We’ve sent a sign-in link to …. Open it on this device — this tab will sign you in automatically.
Nothing arrived? Check your spam folder — and mark the mail as "Not spam" so it lands in your inbox next time.