quantization
4 published articles
Open source AI
Muse Glimmer: Meta's 30B agent fits under 20GB, cloud optional
Meta open-sourced Muse Glimmer, a 30B agentic model that runs offline on a single consumer GPU. Quantization keeps it under 20 GB; a DFlash drafter delivers up to 3.1x faster decoding on an RTX 5090, and the weights are on Hugging Face under Apache 2.0.
2026-08-10
Qwen / Alibaba Cloud
Qwen2.5-Omni outperforms Gemini-1.5-Pro on OmniBench, fits under 12GB
Alibaba's open-source Qwen2.5-Omni outscored Gemini-1.5-Pro on OmniBench and topped the MMAU audio reasoning leaderboard. Quantized builds cut VRAM below 12GB and MNN support brings real-time voice chat to phones.
2026-08-07
AI Research
The monitor that goes silent when AI reasoning fails
Novelis Research shows token log-probability fails as a decoder monitor for quantized reasoning models, being blind to confident loops. They introduce a calibrated e-CUSUM controller that combines uncertainty and repetition signals, achieving selectivity on GSM8K with DeepSeek-R1-Distill-Qwen-1.5B.
2026-08-03
Model Quantization
Inkling was a 1.9 TB model. Unsloth just squeezed it into a desktop.
Unsloth's dynamic GGUF quantization shrinks Inkling, a 975B-parameter open model, from 1.9 TB to 270 GB at 1-bit with 74.2% accuracy retention. The method selectively preserves high-precision layers, enabling local inference on machines with 290 GB of combined RAM and VRAM.
2026-07-16