Fetching from the wire…
Models2026-08-18 · source-backed
Tencent published UI-Mate-27B on Hugging Face, built on Qwen3.6-27B, reading live screenshots and emitting pyautogui-style tool calls. The paper (arXiv 2608.15930) reports 77.0 on OSWorld-Verified and 66.2 on WindowsAgentArena, beating its base by 17.7 and 24.5 points, and introduces OSWorkerBench (100 office tasks, 41 applications) where a single in-context demonstration lifts strict success from 17.2% to 35.4%. That demonstration result is the real headline: an open computer-use model whose reliability you more than double by recording one successful run instead of fine-tuning.
Each link below shares sources, entities, or timing with this story.
On July 14, llama.cpp merged native support for Tencent's Hunyuan Hy3 architecture (PR #25395), a 295B-parameter, 21B-active MoE. Any recent master build can load it now. Community GGUF quants (Q2_K, IQ2_M, Q4_K_M) from AngelSlim and others already ship on Hugging Face, and so...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
DeepSeek released V4 Preview on April 24 with two open-weight variants: V4-Pro (1.6T total parameters, 49B activated via MoE) and V4-Flash (284B parameters, 13B activated). Both support 1M-token context windows. Both are Apache 2.0 licensed. Both are live right now on Hugging...
Announced July 27 with Microsoft, IBM, Red Hat, Palantir, CrowdStrike, Cloudflare, Databricks, Hugging Face, LangChain, Nous Research, Reflection AI, Thinking Machines Lab, SpaceXAI and the Linux Foundation. Huang's framing is pointed: during the Hugging Face incident "closed...
The attackers didn't use agents to help. They used agents to do the whole thing. Hugging Face disclosed that attackers chained a remote-code dataset loader with a template-injection flaw in dataset configuration to land on processing workers, then escalated to node-level acces...
Released August 4 under Apache 2.0, reframing moderation as policy-adaptive question answering: write your rule in plain language, get a calibrated safety score from a single token, no retraining, one interface for text and images. Mistral claims it matches open guard models u...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.