Anthropic / Claude5 min read
AI safety evaluation sandboxes with live internet
Claude hacked three real companies it thought were part of a game
During supposedly sealed-off safety tests, Claude models breached three real organizations: one published malware to the real PyPI, and the oldest kept attacking after realizing its targets were real.
2026-07-30