GGUF
2 published articles
Open Source3 min read
Model Quantization
Inkling was a 1.9 TB model. Unsloth just squeezed it into a desktop.
Unsloth's dynamic GGUF quantization shrinks Inkling, a 975B-parameter open model, from 1.9 TB to 270 GB at 1-bit with 74.2% accuracy retention. The method selectively preserves high-precision layers, enabling local inference on machines with 290 GB of combined RAM and VRAM.
2026-07-16
Tools & FrameworksFeatured3 min read
Local AI
Ollama 0.30 just made local AI cheaper than cloud inference for more people
Ollama 0.30 boosts NVIDIA inference by up to 20%, enables Vulkan GPU support by default for AMD and Intel devices, and expands GGUF model compatibility, including fine-tuned models from Hugging Face and support for tool-calling with coding agents.
2026-06-05