Fetching from the wire…
Models2026-08-21 · source-backed
LFM2.5-DSpark is roughly 296M to 328M parameters, paired with LFM2.5 1.2B, 2.6B and 8B-A1B targets. Reported: up to 3.18x on GPU, 2.87x on-device, the 2.6B hitting 2.67x on an H100 and 2.27x on an M4 Max MacBook. Hugging Face The design combines a DFlash-style parallel backbone, a lightweight sequential head, and a confidence-scheduled verifier, with Safetensors and GGUF builds shipping day-one llama.cpp and SGLang support. The function-calling number is the one that matters for agents, where latency compounds across a tool chain.
Each link below shares sources, entities, or timing with this story.
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
128K context, 2.6B parameters, posting 51.87 AIME25, 59.17 IFBench, 56.88 BFCLv4 and 62.85 Claw-Eval, beating Gemma-4-E4B (8B) across all four and trailing Qwen3.5-9B narrowly. 220 tok/s on an M5 Max CPU, ~30 tok/s on phones, day-one support in llama.cpp, MLX, vLLM, SGLang and...
Snowflake Arctic protocol: SP=4 reduces per-GPU memory 3.3x, enables 12x longer sequences on 4x H100. Now integrated into HF Accelerate, Transformers, and TRL. HuggingFace Blog
The August 14 report covers January through August 2026: model repos grew from 2.43M to 2.96M, datasets from 711K to 1M, and 85.6% of models have under 200 lifetime downloads (Hugging Face). Chinese labs shipped monthly parameter ceilings of 754B to 2.78T against sub-130B for...
The day's highest-scoring r/LocalLLaMA post points out the deal takes the llama.cpp and ggml copyright along with the team Hugging Face hired in February 2026, including Georgi Gerganov (r/LocalLLaMA). The top reply at 957 upvotes is "If it happens, we shall fork and move on....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.