on-device AI
4 published articles
On-device AI agents
LFM2.5-2.6B: the tiny agent that outruns models 4x its size
Liquid AI's LFM2.5-2.6B fits an agentic model into 2.6B parameters and under 2.5 GB of memory, topping every instruction-following benchmark it was tested on. It runs 220 tokens/s on a laptop; coding is the one clear gap.
2026-08-12
Open source AI
Muse Glimmer: Meta's 30B agent fits under 20GB, cloud optional
Meta open-sourced Muse Glimmer, a 30B agentic model that runs offline on a single consumer GPU. Quantization keeps it under 20 GB; a DFlash drafter delivers up to 3.1x faster decoding on an RTX 5090, and the weights are on Hugging Face under Apache 2.0.
2026-08-10
Artificial Intelligence
The 7B model that beats bigger ones, runs on a phone, and changes the AI cost equation
Alibaba Cloud's Qwen2.5-Omni is a 7B model that handles text, images, audio, and video, generating speech and text in real time. Its Thinker-Talker architecture and TMRoPE position embedding enable streaming interaction, and the 3B variant runs on mobile SoCs like Snapdragon 8 Gen 1 at usable speeds.
2026-07-23
Mobile AI
Gemma 4 goes fully offline on mobile, no cloud required
React Native developers can now embed Gemma 4 for offline inference with hardware acceleration on both Android and iOS. The model handles vision and tool-use tasks locally, as demonstrated by reading a flyer and scheduling a calendar event entirely on-device.
2026-07-09