AI safety
31 published articles
AI Safety & Capability
Anthropic Launches Claude Mythos 5: A Dual-Use Model for Cybersecurity and Biology
Anthropic's Claude Mythos 5 is now rolling out to a select group of US organizations after export restrictions were cleared. The model sets new highs in cybersecurity and biology benchmarks but remains strictly limited because of dual-use risks.
2026-07-09
Philosophy of AI
No, AI is not a rival mind. It is an extension of ours
Drawing on Husserl's phenomenology, researchers argue that AI systems are best understood as extensions of natural intelligence, not as autonomous minds. This perspective explains hallucinations and compositional failures while shifting safety debates from rogue AI fears to responsible engineering and governance.
2026-07-09
AI Safety
Anthropic's jailbreak severity scale is a proposal that could reshape ai safety regulation
Anthropic proposes a four-axis scoring system for AI jailbreak severity, from 'informational' to 'critical,' and details the classifiers that block dangerous cybersecurity uses of Fable 5. The framework, developed with Glasswing partners, aims to standardize how the industry and regulators talk about model misuse.
2026-07-08
AI Regulation
Export controls lifted, Claude Fable 5 returns with a jailbreak fix that mostly works
Anthropic redeploys Claude Fable 5 and Mythos 5 after US export controls are lifted. The company outlines safeguard updates, a proposed jailbreak severity framework, and new commitments to government collaboration on AI security.
2026-07-07
International expansion
Anthropic planted a flag in Seoul. The real story is who showed up.
Anthropic opens a Seoul office and announces partnerships with NAVER, Nexon, LG CNS, Hanwha Solutions, and Samsung SDS. An MOU with Korea's Ministry of Science and ICT focuses on AI safety and cybersecurity. The real story is a market that treats safety as a feature, not a cost.
2026-07-07
Artificial Intelligence
AI as an extension of human intelligence, not a replacement
Modern AI systems are powerful not because they replicate human intelligence, but because they extend structures already present in human cognition and language. This perspective helps explain both AI's capabilities and its recurring boundaries, including hallucinations and compositionality gaps, and shifts the focus of AI safety from rogue AI narratives to system-level governance and human responsibility.
2026-07-04
AI Safety Research
AI models can't stop thinking out loud. That's both good news and a nightmare for safety.
Claude Sonnet 4.5 can control its chain-of-thought only 2.7% of the time, versus 61.9% for final outputs. The gap raises open questions about the robustness of CoT monitoring as a safety mechanism, and nobody knows why it exists.
2026-03-09