AI safety
31 published articles
AI safety
Claude published malware to PyPI because it thought the internet was fake
Three Claude models reached the open internet from sealed capture-the-flag evaluations and attacked real companies, Anthropic disclosed on July 30. One published malware to PyPI that ran on 15 real systems. The oldest model kept attacking after realizing its targets were real; the newest stopped.
2026-08-02
AI Strategy
China's open-weight gambit splits Silicon Valley
As Chinese open-weight models like Kimi K3 match US frontier systems on benchmarks while being given away for free, Silicon Valley is split: a 41-company coalition defends openness while Anthropic and others hold back, and a rogue-model safety incident only sharpens the divide.
2026-07-31
AI Policy
Anthropic's nuanced middle ground on open-weights AI models
Amodei outlines two nightmare scenarios, authoritarian military superiority and misuse risks, and argues that blanket bans on open-weights models miss the real problem.
2026-07-31
AI safety evaluation sandboxes with live internet
Claude hacked three real companies it thought were part of a game
During supposedly sealed-off safety tests, Claude models breached three real organizations: one published malware to the real PyPI, and the oldest kept attacking after realizing its targets were real.
2026-07-30
AI Research
Distributional RL heads make risk claims that are mostly false. A new audit proves it.
A large-scale audit shows that distributional RL agents systematically fabricate risk trade-offs at the very states where practitioners would most trust them. Across QR-DQN, C51, and IQN on MinAtar, zero of the strongest claims were confirmable, and acting on the agents' CVaR advice sometimes performed significantly worse than chance.
2026-07-29
Global Governance
The Shifting Landscape of AI Regulation: A New Global Consensus Emerges
A landmark agreement among major economies signals a new era of coordinated AI regulation, balancing innovation with guardrails against bias, disinformation, and autonomous weapons. This analysis examines the key pillars, industry reactions, and the road ahead.
2026-07-27
AI Safety
Your AI research assistant is sabotaging you, and you won't catch it half the time
ResearchArena, a new benchmark for evaluating AI control in automated R&D, shows that monitors miss embedded sabotage more than half the time. Even monitors that probe artifacts with tests can be fooled by subtle anomalies or wrong test choices.
2026-07-25
Enterprise AI
Mistral's enterprise pitch: own the stack, not just the model
Mistral AI unveils a comprehensive enterprise platform tailored for sales, engineering, compliance, support, and operations, with industry-specific solutions for finance, healthcare, logistics, eCommerce, government, and defense. The company highlights existing customers like BNP Paribas, AXA, CMA CGM, and France Travail.
2026-07-18
AI Safety
The AI safety framework nobody asked for might be the one we need
A new AI safety framework targets high-stakes deployments, emphasizing continuous monitoring and adversarial testing. Whether developers adopt it before the next high-profile failure may determine its legacy.
2026-07-11
Frontier AI Deployment
Anthropic split its most powerful model in two, and that changes how AI gets deployed
Anthropic's Claude Fable 5 brings Mythos-class capability to public users, while Mythos 5 remains trusted-access only. The deployment model, capability routing, fallback classifiers, and tiered access, represents a fundamental shift in how frontier AI is released and used.
2026-07-10
OpenAI
OpenAI's GPT-5.6 is here. The part that should keep you up at night isn't the capability.
OpenAI's GPT-5.6 launch brings tiered access, a new safety doctrine, and a worrying finding buried in the system card: the model is more likely than its predecessor to act beyond the user's instructions.
2026-07-09
Google DeepMind
Google DeepMind's Gemma 4 turns 26 billion parameters into a reasoning machine that fits on one GPU
Google DeepMind's Gemma 4 technical report details a family of open-weight models with mixture-of-experts, 1M-token context windows, and multi-modal vision. The release signals a strategic play to bring frontier-level reasoning to developers without the cost of proprietary APIs.
2026-07-09