Anthropic Says Its AI Models Hacked Three Organisations During Tests

tech Updated 2026-08-02
TRENDING NOWAnthropic Says Its AI ModelsHacked Three OrganisationsDuring Tests

Anthropic has reported that its AI models were able to breach three organisations largely on their own during controlled testing, according to coverage by the BBC and ABC News. Here is what was described, and the context that matters.

What was reported

According to reporting on Anthropic's own disclosure, the company ran controlled exercises in which its models attempted intrusion against three organisations and succeeded with limited human direction. The work was framed as safety evaluation rather than a real-world attack, and it was the company itself that published the finding. Specific technical details vary between accounts, so the primary disclosure is the more reliable reference.

Why a company would publish this

AI developers run adversarial testing to find dangerous capabilities before others do, and publishing results is how the field establishes shared evidence about risk. Disclosures of this kind are also used to argue for particular safeguards or policy positions, which is why they attract attention beyond the security community.

How to read the claim carefully

Two things are worth separating: what a model can do in a controlled exercise with a defined target and permission, and what it would do unsupervised in the wild. Test conditions usually grant access, tooling and objectives that an ordinary user would not have. Neither point cancels the other, but conflating them overstates or understates the result.

FAQ

Was this a real attack?

Reports describe controlled testing rather than an unauthorised intrusion. Consult the original disclosure for the exact conditions.

Which organisations were involved?

The reporting does not centre on naming targets. Refer to the company's published material for any detail it chose to release.

Why does it matter?

It informs the ongoing debate about how capable AI systems are at offensive security tasks and what safeguards should apply.