Fetching from the wire…
Public story · 2026-08-17 · high
Version 1.1.10 shipped 32 MLX benchmarks on an Apple M4 Pro, the project's first, and fixed an Ollama bug that mismarked model families as installed.
Why now: llmfit published its first MLX benchmarks and shipped v1.1.10 on August 17.
Local-inference tool llmfit shipped version 1.1.10 with the project's first MLX benchmark results, run on an Apple M4 Pro, per its GitHub release.
That matters for anyone picking a local-inference runtime on Apple Silicon. MLX and llama.cpp's Metal backend rarely get compared on identical hardware, and 32 runs now do exactly that.
Those 32 results merged into llmfit's benchmark suite as new entries, giving developers a same-chip reference point instead of scattered numbers from different machines. The GitHub release doesn't say which backend won more of the 32 runs, only that the entries now exist side by side.
Version 1.1.10 also expanded llmfit's MCP server with RamaLama runtime discovery, and added the Qwen3.8 model family alongside vision capability exposure for MiniMax M3.
A separate fix stops a single sized Ollama install from marking its whole model family as installed.
The benchmark numbers matter less than the precedent they set. Once one project publishes a same-hardware MLX-versus-Metal comparison, framework maintainers lose the excuse that unfavorable results came from mismatched test rigs. Watch whether other local-inference tools start citing llmfit's numbers instead of running their own one-off comparisons.
llmfit shipped this comparison on August 17.
Each link below shares sources, entities, or timing with this story.
Three moves, two days, no coordination between them. August 10–11: GitHub shipped Ollama as a BYOK provider inside Copilot for JetBrains (GitHub Changelog). Unsloth released Unsloth Desktop with a command literally named unsloth start claude, which points Claude Code and Codex...
The July 6 release delivers nearly 90% faster Gemma 4 token generation through multi-token prediction with automatic draft-length tuning, on by default, output-preserving, no config (Ollama). It also adds MLX-engine support for more model families and flash attention for older...
Released August 26, it brings MLX support for Qwen3.8 Flash Next, adds structured output to mlxrunner, and stops Metal GPU timeouts when loading models from slow storage (GitHub). Structured output on the MLX path is the practical unlock: Apple Silicon local inference can now...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
v0.1.803-beta, released August 25 with 170+ PRs, lets long local chats continue past a model's context limit by rolling older turns into fresh context epochs rather than permanently trimming, with evicted conversations still searchable (GitHub). It also fixes MLX and Mac runti...
The repo appeared August 24, opening the weights of a multimodal MoE that Qwen frames explicitly as an architecture preview, the same role Qwen3-Next played for Qwen3.5 (GitHub). The hybrid Gated DeltaNet plus Gated Attention design it previews already carried through the Qwen...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.