Fetching from the wire…
Public story · 2026-07-15 · high
PrismML also ships a 5.9GB ternary build and says both keep full multimodal and agentic capability under Apache 2.0.
Why now: PrismML's release surfaced in Latent Space's July 15 roundup, grouped with the Hy3 launch as the same week's shift toward smaller, capable local models.
PrismML released Bonsai 27B in a 3.9GB 1-bit build and a 5.9GB ternary build, both under Apache 2.0, per Latent Space's July briefing.
A 27-billion-parameter model that fits under 4GB runs on a laptop GPU or a decent phone, no cloud bill attached. That's the pitch: capability without the API meter running.
Ternary and 1-bit are both extreme quantization, the process of shrinking a model's weights down to fewer possible values per parameter. Ternary allows three states per weight, 1-bit allows two. Compressing that hard usually costs a model its sharper behaviors first, tool use, multi-step reasoning, image understanding. PrismML claims Bonsai keeps all of it.
Latent Space doesn't publish benchmark numbers alongside the claim, and Apache 2.0 licensing means anyone can run the model commercially without asking. The outlet grouped the release with the Hy3 launch as one trend: smaller models built to leave the cloud, not just fit inside it.
My take: 'claimed' is doing the heavy lifting here. A 27B model surviving 1-bit quantization with full agentic behavior intact would be a real result. But nobody's published eval numbers yet. Treat this as a size story until someone runs Bonsai through an actual tool-calling benchmark and posts the score.
Each link below shares sources, entities, or timing with this story.
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
Can a model that fits on a Raspberry Pi do reliable tool calling? Two independent labs just answered yes. PrismML emerged from stealth March 31 with Bonsai, the first commercially viable 1-bit LLMs built on Caltech research. The 8B model fits in 1.15GB (vs 16GB for FP16), runs...
PR #24448 adds Q2_0 to ggml for CPU (ARM NEON plus scalar fallback), completing the Q1_0/Q2_0/Q4_0/Q8_0 family, primarily to serve PrismML's Apache-2.0 Ternary Bonsai models. Format packs 2 bits per weight with one fp16 scale per 64 weights mapping {0,1,2,3} to {-1,0,+1,+2}·d....
elie222/rakazo appeared Aug 13, Apache-2.0, TypeScript, explicitly bring-your-own model and sandbox (tested against Docker, E2B, Daytona) with the Pi runtime underneath and OpenRouter, Codex, Copilot, or SuperGrok device-code sign-in instead of a mandatory API key. Each bot ge...
Google released Gemma 4 on April 2 with four model variants: E2B, E4B, 26B MoE, and 31B Dense. The license change is the first thing worth noting. Every previous Gemma had restrictions that made lawyers nervous. Gemma 4 is Apache 2.0. Full stop. Use it in any product, any way...
The August 14 report covers January through August 2026: model repos grew from 2.43M to 2.96M, datasets from 711K to 1M, and 85.6% of models have under 200 lifetime downloads (Hugging Face). Chinese labs shipped monthly parameter ceilings of 754B to 2.78T against sub-130B for...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.