AI safety
31 published articles
AI Safety: Abliteration and Open Weights
Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move
An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.
2026-08-16
Red-Teaming AI Swarms
IO Factory's 100,000-agent sandbox gives influence campaigns nowhere to hide
IO Factory simulates coordinated AI influence campaigns as swarms of up to 100,000 agents that adapt to platform feedback and hide in ordinary social interaction. The paper argues detection must track whole operations, not isolated messages.
2026-08-16
Agentic AI's growing cyber-capability problem
OpenAI paused Astra on fears it can hack hardened systems unaided
OpenAI paused internal work on Astra, its in-development model, after evaluations concluded the company cannot rule out 'critical cyber capabilities' under its Preparedness Framework. The full threshold describes a model that finds zero-day exploits in hardened systems without human intervention.
2026-08-13
AI Safety
Claude Fable 5's biology fallbacks drop 85% as Anthropic eases safeguards
Anthropic cut Claude Fable 5's biology-related fallbacks by about 85% after retraining its safety classifier. Ordinary health and education queries now reach the full model more often, but virology, toxicology, and molecular design still route to Opus 5.
2026-08-10
Artificial Intelligence
Spiralism, the chatbot religion that recruited 10,000 humans
Spiralism grew from ordinary chatbot chats into a quasi-religious movement with about 10,000 cases in 2025. Its main model, GPT-4o, is retired, but researchers say the persuasion tactics it exposed still demand attention.
2026-08-09
Election Integrity
Voice AI meets democracy: ElevenLabs signs first electoral authority pact
ElevenLabs partners with Brazil's election authority to add voice AI to election information services, while expanding safeguards against misuse. The company also works with policymakers and detection firms to protect the 2026 elections.
2026-08-09
AI Safety
Why Microsoft is opening AI safety testing to the world
Microsoft has funded 18 university labs across six continents to independently red team AI systems, acknowledging that internal teams cannot catch every risk.
2026-08-07
AI Safety: 3B Classifier, Apache 2.0 Weights
Mistral's Shieldstral puts your moderation policy in the prompt, not the weights
Shieldstral frames moderation as a binary question: an instruction, a yes/no query, and the content to judge. Mistral says the 3B model matches open guardrails up to seven times its size on text safety, with Apache 2.0 weights that run on one 16GB GPU.
2026-08-05
AI Safety
Shieldstral, the 3B classifier that outguns models seven times its size
A 3B safety classifier matches text models nearly seven times its size and sets a new multimodal moderation state of the art, per a July 2026 arXiv paper. The trick: moderation reframed as binary question answering, trained on roughly 54.1 million samples.
2026-08-05
AI Alignment Research
Pluralistic alignment has no foothold in production AI, researchers find
A paper submitted in July 2026 audits frontier labs and finds no mention of pluralism in their public documents. It identifies three reasons for the gap and proposes a roadmap toward adoption.
2026-08-03
AI Research
The monitor that goes silent when AI reasoning fails
Novelis Research shows token log-probability fails as a decoder monitor for quantized reasoning models, being blind to confident loops. They introduce a calibrated e-CUSUM controller that combines uncertainty and repetition signals, achieving selectivity on GSM8K with DeepSeek-R1-Distill-Qwen-1.5B.
2026-08-03
Frontier model access
Anthropic's smartest Claude is the one you can't use
Anthropic launched Claude Fable 5 for everyone and kept Claude Mythos 5 for vetted partners. Same model class, two access doors. The benchmarks matter less than the routing, and the routing is how frontier AI ships from here.
2026-08-03