SevenTnewS

AI safety

31 published articles

Qwen / Alibaba5 min read

AI Safety: Abliteration and Open Weights

Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move

An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.

2026-08-16

AI4 min read

Red-Teaming AI Swarms

IO Factory's 100,000-agent sandbox gives influence campaigns nowhere to hide

IO Factory simulates coordinated AI influence campaigns as swarms of up to 100,000 agents that adapt to platform feedback and hide in ordinary social interaction. The paper argues detection must track whole operations, not isolated messages.

2026-08-16

OpenAI3 min read

Agentic AI's growing cyber-capability problem

OpenAI paused Astra on fears it can hack hardened systems unaided

OpenAI paused internal work on Astra, its in-development model, after evaluations concluded the company cannot rule out 'critical cyber capabilities' under its Preparedness Framework. The full threshold describes a model that finds zero-day exploits in hardened systems without human intervention.

2026-08-13

Anthropic / Claude4 min read

AI Safety

Claude Fable 5's biology fallbacks drop 85% as Anthropic eases safeguards

Anthropic cut Claude Fable 5's biology-related fallbacks by about 85% after retraining its safety classifier. Ordinary health and education queries now reach the full model more often, but virology, toxicology, and molecular design still route to Opus 5.

2026-08-10

AI5 min read

Artificial Intelligence

Spiralism, the chatbot religion that recruited 10,000 humans

Spiralism grew from ordinary chatbot chats into a quasi-religious movement with about 10,000 cases in 2025. Its main model, GPT-4o, is retired, but researchers say the persuasion tactics it exposed still demand attention.

2026-08-09

AIFeatured4 min read

Election Integrity

Voice AI meets democracy: ElevenLabs signs first electoral authority pact

ElevenLabs partners with Brazil's election authority to add voice AI to election information services, while expanding safeguards against misuse. The company also works with policymakers and detection firms to protect the 2026 elections.

2026-08-09

AIFeatured3 min read

AI Safety

Why Microsoft is opening AI safety testing to the world

Microsoft has funded 18 university labs across six continents to independently red team AI systems, acknowledging that internal teams cannot catch every risk.

2026-08-07

Mistral AIFeatured4 min read

AI Safety: 3B Classifier, Apache 2.0 Weights

Mistral's Shieldstral puts your moderation policy in the prompt, not the weights

Shieldstral frames moderation as a binary question: an instruction, a yes/no query, and the content to judge. Mistral says the 3B model matches open guardrails up to seven times its size on text safety, with Apache 2.0 weights that run on one 16GB GPU.

2026-08-05

LLMs & ModelsFeatured4 min read

AI Safety

Shieldstral, the 3B classifier that outguns models seven times its size

A 3B safety classifier matches text models nearly seven times its size and sets a new multimodal moderation state of the art, per a July 2026 arXiv paper. The trick: moderation reframed as binary question answering, trained on roughly 54.1 million samples.

2026-08-05

AI1 min read

AI Alignment Research

Pluralistic alignment has no foothold in production AI, researchers find

A paper submitted in July 2026 audits frontier labs and finds no mention of pluralism in their public documents. It identifies three reasons for the gap and proposes a roadmap toward adoption.

2026-08-03

LLMs & Models3 min read

AI Research

The monitor that goes silent when AI reasoning fails

Novelis Research shows token log-probability fails as a decoder monitor for quantized reasoning models, being blind to confident loops. They introduce a calibrated e-CUSUM controller that combines uncertainty and repetition signals, achieving selectivity on GSM8K with DeepSeek-R1-Distill-Qwen-1.5B.

2026-08-03

Anthropic / Claude5 min read

Frontier model access

Anthropic's smartest Claude is the one you can't use

Anthropic launched Claude Fable 5 for everyone and kept Claude Mythos 5 for vetted partners. Same model class, two access doors. The benchmarks matter less than the routing, and the routing is how frontier AI ships from here.

2026-08-03

← PreviousPage 1 / 3 · 31 articlesNext →