Fetching from the wire…
Public story · 2026-08-31 · high
A Reddit demo squeezed Qwen3.8-Flash-Next onto a $400-500 Android handset with 12GB of RAM by offloading most of the model to storage.
Why now: Posted to r/LocalLLaMA on August 31, 2026.
Someone ran Qwen3.8-Flash-Next, an 80GB model, on a mid-range Android phone with just 12GB of RAM. The video post shows it generating text at 3.5 tokens a second, slow enough to watch each word land but fast enough to prove the thing runs at all.
The top comment flags the real cost: the CPU reached 80C running this. That's not a number you build a product around. It's a number you stop a demo at before something throttles or degrades.
The trick is the same one that makes Flash-Next usable on constrained desktops, sparse mixture-of-experts routing plus offloading. The phone only has to hold and compute a fraction of those 80GB at any moment. The post adds aggressive quantization on the dense part and unspecified extra optimizations on top, though it doesn't say what those are.
A 12GB phone holding an 80GB model at all means the routing and offload math scale down further than a lot of people assumed. Whether it scales down to something you'd run during an actual phone call is a separate question, and this demo doesn't answer it. Nobody's shipping a chat app on this setup at 80C.
Each link below shares sources, entities, or timing with this story.
Google dropped Gemma 4 and it's not incremental. The 31B dense model ranks #3 on Arena AI with an ELO of 1,452, scores 85.2% on MMLU Pro, 89.2% on AIME 2026, and 80.0% on LiveCodeBench v6. It outperforms models 20x its size. Under Apache 2.0. At $0.20 per run. Only Opus 4.6 an...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Google DeepMind released Gemma 4 on April 2 with four model sizes (E2B, E4B, 26B MoE, 31B Dense) under Apache 2.0. Multimodal (text, vision, audio). 256K context. Native thinking and tool-calling optimized for agentic workflows. Day-zero ecosystem support across vLLM, llama.cp...
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
Alibaba released Qwen3.6-27B on April 22. Dense architecture. Open weights. 77.2% on SWE-bench Verified, within 3.7 points of Claude Opus 4.6. On SkillsBench, it scores 48.2% versus its own 397B MoE predecessor's 30.0%. That's a 77% improvement with 14.8x fewer parameters. Let...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.