alignment research
2 published articles
Qwen / Alibaba5 min read
AI Safety: Abliteration and Open Weights
Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move
An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.
2026-08-16
LLMs & Models4 min read
AI Safety Research
AI models can't stop thinking out loud. That's both good news and a nightmare for safety.
Claude Sonnet 4.5 can control its chain-of-thought only 2.7% of the time, versus 61.9% for final outputs. The gap raises open questions about the robustness of CoT monitoring as a safety mechanism, and nobody knows why it exists.
2026-03-09