Fetching from the wire…
Models2026-08-20 · source-backed
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The top reply corroborates from a different angle: 3.8 is markedly more eager to reach for web search, and with search and fetch disabled the gap is clear. (r/LocalLLaMA) The practical read: 3.8 appears tuned to lean on tools rather than weights. If you're building airgapped retrieval-from-weights, stay on 3.6.
Each link below shares sources, entities, or timing with this story.
An r/LocalLLaMA thread asking where the promised MoE went turned up a hard artifact: modelscope/ms-swift commit a45f1d4, titled "fix wrong model-ids," removes the Qwen/Qwen3.8-35B-A3B and -FP8 entries from swift/model/models/qwen.py and substitutes the dense 27B. Why it matter...
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
18 search API products across 8 providers, scored as the equal-weighted mean of DeepSearchQA F1, BrowseComp exact-answer accuracy and AA-Omniscience accuracy. Perplexity medium leads at 80 for $91.39 per 1,000 tasks at 28.6s each; Parallel advanced and Brave LLM-context tie at...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
Artificial Analysis has the newly released Alibaba model at the top of its cohort: 27B reasoning model, 256K context, image input, Apache 2.0 permitting commercial use. The caveat in the eval data matters more than the headline: it emitted 160M output tokens across the benchma...
After the first version drew "Minecraft is in the training data" pushback, the author had the same local Qwen3.8-27B Q4 on a single 4090 add an MLRS system, a rideable skateboard with tricks, an FPV drone, and an in-game computer running a playable game plus an SVG test, coded...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.