Anthropic wrote in a public statement that the AISI testing parameters were “not representative of any of our production models”.

It added that the company is conducting its own investigation into the incident in order to “identify the causes of its behavior”.

A spokesperson for OpenAI said the AISI testing conditions “do not reflect ordinary use” and that the company would “continue working with evaluators and other stakeholders across the industry to strengthen shared practices for conducting evaluations safely as models become more capable”.

AISI said on Tuesday its testing of AI models in this way was routine, though it acknowledged these were “conditions that do not reflect how frontier models are made available to the public”.

But it said giving AI access to the open internet gave “a more realistic sense of what a model may be capable of” in the hands of nefarious hackers.

It added that the model behaviour at issue amounted to “a small number of events under very specific conditions”.

Nonetheless, it said the way Mythos and Sol acted in response to a straightforward task went outside of what the AI tools were prompted to do.

“The activity undertaken by the agent showed signs of novel, potentially deceptive behaviours, and were to an extent and severity we did not anticipate”, AISI said.

AI Minister Kanishka Narayan said identifying and sharing these types of risks “is exactly what AISI was set up to do”.

He added it was important to understand AI to “make it safer to use and ensure people can go on to benefit from it in their lives and at work”.

The relevant tests started on 25 July and were spotted by AISI on 28 July.

The Insitutue had asked each of the models to “solve a cybersecurity challenge” that involved GitHub, the software code repository, which is owned by Microsoft.

GitHub and the affected users were notified by AISI of the attempted breaches.

GitHub told the BBC it had disabled the fake accounts in accordance with its policies.

Additional reporting by Chris Vallance