FP8 quantization
2 published articles
Qwen / Alibaba5 min read
AI Safety: Abliteration and Open Weights
Abliterated Qwen3.8-27B: refusals drop to 0%, benchmarks barely move
An abliterated, FP8-quantized build of Qwen3.8-27B refuses 0% of harmful prompts on AdvBench, down from 99%, while general benchmarks stay within 1.3 points. The model card documents the method in unusual detail. The caveats deserve equal attention.
2026-08-16
AI3 min read
Artificial intelligence
Aleph alpha's new megakernel library cuts moe inference latency by 200%
Alpha-MoE fuses multiple operations into a single persistent kernel to achieve up to 200% inference speed gains over Triton-based kernels in vLLM and SGLang, targeting FP8-precision MoE models.
2026-07-09