Inkling
2 published articles
Open Source3 min read
Model Quantization
Inkling was a 1.9 TB model. Unsloth just squeezed it into a desktop.
Unsloth's dynamic GGUF quantization shrinks Inkling, a 975B-parameter open model, from 1.9 TB to 270 GB at 1-bit with 74.2% accuracy retention. The method selectively preserves high-precision layers, enabling local inference on machines with 290 GB of combined RAM and VRAM.
2026-07-16
LLMs & Models7 min read
Artificial Intelligence
Inkling is open-source AI's $53 billion reality check
Inkling is the first openly available near-1T parameter model with native audio, image, and text input alongside a 1M-token context window. The raw benchmark scores are strong. The real story is how the open-source ecosystem has moved from playing catch-up to competing at the frontier, and where Inkling fits on that new map.
2026-07-15