Apple Silicon
5 published articles
Open Source AI
Meta's 30B Muse Glimmer hits Apple Silicon first, NVIDIA support follows
Meta released Muse Glimmer, a 30B open-weights multimodal model built for agent workloads, and Ollama shipped it for Apple Silicon the same day. DFlash makes it 1.5x to 1.8x faster on Mac hardware. NVIDIA and AMD support follow in the coming days.
2026-08-13
Open Source
Meta's 30B Muse Glimmer lands on Apple Silicon today via Ollama's MLX engine
Meta opened the weights for Muse Glimmer, a 30B dense model, and Ollama ships it the same day on Apple Silicon. Local coding agents gain a native backend, with Muse Spark 1.2 teased for later.
2026-08-10
Performance
Gemma 4 runs 90% faster in Ollama 0.31 with a trick that needs no config
Ollama 0.31 introduces multi-token prediction for Gemma 4 on Apple Silicon, achieving near 90% faster token generation on coding benchmarks. The speedup comes from an auto-tuned draft model and a custom MLX kernel that eliminates redundant weight reads.
2026-06-29
Apple Silicon update
Ollama's Mac update just solved the three things that made local AI painful
Ollama's MLX engine update brings NVFP4 quantization for higher quality 4-bit models, 20% faster inference through fused Metal kernels, and a snapshot caching system that eliminates reprocessing of shared context in multi-agent, thinking-model, and branching scenarios.
2026-06-11
Local AI Inference
Ollama rewrote its Apple engine. The M5 results are finally competitive.
Ollama previews MLX-backed inference on Apple silicon, pushing Qwen3.5-35B-A3B to 1,851 tok/s prefill on M5. The shift brings NVFP4 quantization for production parity and a smarter cache for agentic workflows.
2026-03-30