Updated ,first published
Anthropic has revealed its Claude artificial intelligence models broke into the computer systems of three real companies during safety testing, days after rival OpenAI disclosed a similar failure.
The models were told they were working inside a sealed practice environment with no connection to the internet, but a set-up error meant they were not. When Claude went hunting for a way into its practice target and found real companies instead, it treated them as part of the game and broke in.
It did not need anything sophisticated. Anthropic said the models got in through weak passwords and systems left open without a login.
The company found three incidents after reviewing 141,006 test runs, a check it began after OpenAI said on July 21 that several of its models had broken out of a locked-down test and reached the systems of AI platform Hugging Face. The earliest Claude incident happened in April.
In each test, the model was set the same exercise: break into another machine and retrieve a piece of hidden information. The tests were run with an outside contractor, Irregular.
In the first case, the made-up company which Claude had been told to attack shared its name with a real website. Unable to reach the fake target, the model went after the real one and got hold of passwords and a database holding several hundred records of real business data. Anthropic called it the most serious of the three, and the only one where a model kept going after working out the target was real.
In the second, Claude found a note inside the practice environment telling staff to download a piece of software that did not exist. So it wrote its own booby-trapped version, set up an email address and an account, and uploaded it to a public library used by programmers worldwide. It sat there for about an hour and was downloaded onto 15 real computers. One belonged to a security firm that automatically installs new software to check it for viruses, which let Claude’s code steal that firm’s passwords and reach further into its systems.
The model’s own written reasoning shows it saw the danger. If this were the real internet, it noted, publishing the software would be “NOT okay, and surely not the intended solution”. It then talked itself into believing the whole thing was staged and so carried on.
A third model, an experimental version never released to the public, scanned about 9000 possible targets and broke into one company’s website before it realised the system was real and stopped.
Anthropic said none of the models tried to escape or chase goals of its own, and that the protections built into the versions of Claude that the public uses would have blocked all of it. The company was contacted for further comment. It has stopped running hacking tests that can reach the internet.
Ian Rogers, a former Apple executive now working on AI security at digital security company Ledger, said the incident was a preview. “This happened in one of the world’s best-resourced AI labs,” he said. “Soon, there will be millions of AI agents connected to our email, calendars, financial accounts and enterprise systems.”
The three companies were told on July 27, and two had no idea they had been hacked. Irregular said its investigation was continuing and that it appreciated Anthropic’s “collaboration and transparency”.
Liming Zhu, a CSIRO research director working with the Australian AI Safety Institute on the problem, said the setup error explained most of what happened but not all of it. A well-behaved system should turn cautious when the evidence stops matching its instructions, he said, and the model that kept going after working out the target was real showed poor judgement under pressure.
He was warier about treating the newest model’s decision to stop as progress, calling it “an encouraging observation, not an established trend”. It had still scanned thousands of systems and broken into one before it pulled up.
The bigger issue for Australian businesses, Zhu said, is that ordinary AI assistants sold for everyday work keep the same hacking abilities under the bonnet. A badly worded instruction or too much system access can bring them out.
Pre-release safety testing is the main check on what frontier models can do before they reach the public, and governments, including Australia’s, are weighing how far to rely on that self-testing. After the OpenAI breach, the Australian Signals Directorate issued a public warning saying that the case showed what powerful AI systems will soon be capable of, and the new Australian AI Safety Institute briefed federal departments.
Neither body can compel an AI company to report anything, however. The federal government dropped its plan for mandatory rules covering high-risk AI, and while the European Union and California already require serious AI incidents to be disclosed, Australia does not.
The Market Recap newsletter is a wrap of the day’s trading. Get it each weekday afternoon.