Fetching from the wire…
Models2026-08-29 · source-backed
English and Chinese, with reference-free voice design from a natural-language description, reference-guided cloning and low-latency streaming, currently first among open-weight models on the Artificial Analysis TTS leaderboard. The 6GB/4x-realtime figure comes from testers using the audio.cpp dev branch. Caveats from the same thread: the supported tag set is limited and not always followed, quality varies noticeably by seed, and the multilingual variant on the playground hasn't been released. (r/LocalLLaMA)
Each link below shares sources, entities, or timing with this story.
An r/LocalLLaMA thread asking where the promised MoE went turned up a hard artifact: modelscope/ms-swift commit a45f1d4, titled "fix wrong model-ids," removes the Qwen/Qwen3.8-35B-A3B and -FP8 entries from swift/model/models/qwen.py and substitutes the dense 27B. Why it matter...
An r/LocalLLaMA post reports a working deployment on 80x RTX 5090 connected over 25 gigabit Ethernet rather than NVLink or InfiniBand — roughly 2.5TB aggregate VRAM against a ~594GB MXFP4 weight file, surplus absorbed by activation and KV overhead. Expert-parallel MoE over com...
PR #19378 landed in llama.cpp this week, and I think most people are underselling what it means. Backend-agnostic tensor parallelism via --split-mode tensor makes multi-GPU inference work across AMD, Intel, and Apple Silicon. Not just CUDA. Everything. For context: llama.cpp h...
A practitioner running a private trivia set found 3.8 failing questions 3.6 answered reliably, at every quantization and sampling setting tried, then checked Artificial Analysis' Omniscience evaluation and found the same regression in offline no-tool knowledge accuracy. The to...
owensong released Inflect-Nano-v2 (3,966,721 deployable params) and Inflect-Micro-v2 (9,356,513) under Apache-2.0, VITS-family end-to-end text-to-waveform with 128 latent channels, 3 encoder layers, 4 flow coupling blocks, 24 kHz mono. Nano-v2 runs at 0.0933 RTF (10.72x real-t...
A user moved from UD-Q3_K_XL at 140k context to UD-IQ3_XXS and cleared 200k on a 16GB eGPU over Thunderbolt 4, with KV cache at q5_1 and llama.cpp built with DGGML_CUDA_FA_ALL_QUANTS=ON (r/LocalLLaMA). Prompt processing fell from 700-800 tok/s to 400. A commenter on an RTX 508...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.