abliteration
2 published articles
Qwen / Alibaba5 min read
AI Safety: Abliteration and Open Weights
Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move
An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.
2026-08-16
AIFeatured5 min read
local ai
The uncensored model paradox: one knob removes both the annoying refusals and the safety guardrail
Benchmarking five locally run uncensored LLMs shows abliteration cuts over-refusal from 44% to near zero with no hit to reasoning, but the same edit collapses safety refusals from 41.5% to 9.5%, because both ride on the same internal direction. The real reason to run uncensored may not be what you think.
2026-07-19