Fetching from the wire…
Infra2026-08-09 · source-backed
PR #24448 adds Q2_0 to ggml for CPU (ARM NEON plus scalar fallback), completing the Q1_0/Q2_0/Q4_0/Q8_0 family, primarily to serve PrismML's Apache-2.0 Ternary Bonsai models. Format packs 2 bits per weight with one fp16 scale per 64 weights mapping {0,1,2,3} to {-1,0,+1,+2}·d. On an M4 Pro at 8 threads the 1.7B goes from 3.20 GiB/48.70 t/s at F16 to 461.79 MiB/117.20 t/s; the 8B from 15.25 GiB/14.88 t/s to 2.15 GiB/30.14 t/s. Mean KLD 0.00012–0.00020, 99.31–99.39% identical top-1 tokens. The asymmetry to check before swapping: token generation roughly doubles but prefill is slightly slower than F16 (170.51 vs 200.15 t/s at pp512). x86, Metal, CUDA, and Vulkan staged for later PRs.
Each link below shares sources, entities, or timing with this story.
Can a model that fits on a Raspberry Pi do reliable tool calling? Two independent labs just answered yes. PrismML emerged from stealth March 31 with Bonsai, the first commercially viable 1-bit LLMs built on Caltech research. The 8B model fits in 1.15GB (vs 16GB for FP16), runs...
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
Ternary lands at 5.9GB, 1-bit at 3.9GB, both Apache 2.0, with claimed preserved multimodal and agentic behavior. (Latent Space) A genuinely agent-capable model at consumer-hardware footprint under a permissive license is a real shift in what runs off-cloud, and it pairs with t...
750 points on HN. A ggml-based ASR inference library built as a drop-in whisper.cpp replacement, shipped through Mozilla.ai's Builders in Residence program by the maintainer of the Handy speech-to-text app. GPU acceleration via Vulkan, Metal, CUDA and TinyBLAS, and every suppo...
2,823 stars, Apache-2.0, three releases across two days. It wraps the open Demucs htdemucs_6s model to split vocals, drums, bass, guitar, piano and other, auto-selecting CUDA, MPS or CPU, then opens a browser multitrack mixer for solo, loop and per-stem export. Built by one de...
At its June 24 Investor Day, Qualcomm agreed to acquire Modular (Mojo language, MAX inference engine, founded by LLVM/Swift creator Chris Lattner) all-stock at $3.92B, and unveiled the Dragonfly C1000 data-center CPUs with Meta as launch customer. It's a ~$14B RISC-V-plus-open...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.