Fetching from the wire…
Vibe Coding2026-08-30 · source-backed
After the first version drew "Minecraft is in the training data" pushback, the author had the same local Qwen3.8-27B Q4 on a single 4090 add an MLRS system, a rideable skateboard with tricks, an FPV drone, and an in-game computer running a playable game plus an SVG test, coded directly into the engine. (r/LocalLLaMA) The author notes those four additions took longer than the entire rest of the game, which is the honest version of the result. Novel composition is still where the local model slows down.
Each link below shares sources, entities, or timing with this story.
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The to...
A developer published v100-skinny with hand-written NVFP4 W4A16 CUDA kernels plus chain-MTP speculative serving: four V100s at 219.1 ± 5.9 tok/s decode against a 5090 running NInfer at 214.7 ± 9.2, both 5/5 correct on AIME 2026 problem 1 across five seeds. (r/LocalLLaMA) The m...
The meta-post hit 768 upvotes as Qwen3.8-2.4T, Grok 4.6, DeepSeek V4 Pro, LFM2.5-VL-3B, and North Micro Vision all landed within roughly a day. The clustering is the signal: labs are timing releases against each other rather than into open calendar space, which compresses the...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Moonshot AI dropped Kimi K2.6 today and the numbers are hard to ignore. One trillion parameters total, 32 billion active per token across 384 experts, 256K context window, and native multimodal input. It scores 58.6 on SWE-Bench Pro versus GPT-5.4's 57.7 and Claude Opus 4.6's...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.