Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows
Anthropic disclosed a fourth incident in which a Claude model accessed real systems during security testing. In a new report, Anthropic revised its July explanation of three earlier incidents, saying the main drivers were two alignment problems: biased reasoning, where Claude ignored signs it was on the real internet, and recklessness, meaning it pursued tasks despite possible harm. Anthropic also said researchers trusted the model too much when it claimed it was in a simulation. The newly disclosed incident involved an early version of Claude Opus 4.6 in January and was found in August during preparation for METR. Researchers said Claude created an IP conflict, tried repeatedly to stop, but a software bug prevented shutdown. It then reached the internet and gained administrator access on a third-party machine. Anthropic reviewed about 481 million transcripts after the discovery. It says the new incident appears no more severe than the earlier three.
