Fetching from the wire…
Public story · 2026-08-07 · high
It runs on 1.3B active parameters out of 7.9B total, keeps 256K context and function calling, and bumps Ling 3.0 Flash from Vercel's free tier.
Why now: The free window runs from the August 6 release through 8:00am PT on August 14, the only stretch it costs nothing to test.
InclusionAI released Ling-3.0-tiny on August 6, and Vercel dropped it straight into the free slot on its AI Gateway, per the Vercel changelog. For builders testing agent loops on a budget, that means a tool-calling model with 256K of context is free to try through August 14.
Ling-3.0-tiny runs on 1.3B active parameters out of 7.9B total, a size built for constrained hardware and fully local deployment. It switches between a Thinking mode for reasoning and an Instant mode for fast replies, and it keeps native function calling and prompt caching even at that small active-parameter count. Vercel demoed it running obsidian-cli against a local notes repo, an agent task that usually needs a bigger model to hold context and call functions reliably.
That's the part worth noticing. Most models this small give up tool calling or context length to hit a low parameter count. Ling-3.0-tiny keeps both, so it can run an actual agent loop on modest hardware instead of only answering chat prompts.
Vercel is giving it away through 8:00am PT on August 14; the changelog doesn't say what fills the slot after that. The real test isn't the benchmark numbers, it's whether builders actually run a 1.3B-active-parameter agent loop while it costs nothing, or wait to see if it holds up once the meter's back on.
Each link below shares sources, entities, or timing with this story.
The changelog offers Z.ai's open-weights coding model with a 1M-token context free via Blackbox AI on AI Gateway, default for new eve agents, switchable for existing ones with eve set --model zai/glm-5.2. Excludes Fast mode and the glm-5.2-fast variant. Separately, Gemini 3.7...
Simon Willison found it in the OpenRouter price list. 1.6T MoE, ~49B active, 1M context, up to 384K output, at $0.435/M input (cache miss) and $0.87/M output. The agentic-coding deltas versus the preview are the story: DeepSWE 12.8 → 62.7, CyberGym 52.7 → 83.3, Terminal Bench...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Two adapters on August 25: the Notion one lets an agent already running on Slack, Discord, GitHub, Teams or WhatsApp join comment discussions on Notion pages with no separate codebase, and the XChat one handles encryption, key management and signature verification for E2E-encr...
Satya Nadella said companies routing everything through a single proprietary lab may not survive. His argument: you hand that lab your most sensitive business context, and the lab can turn it against you as a competitor. His prescription is an orchestration layer — keep the ha...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.