Fetching from the wire…
Research2026-07-13 · source-backed
Test-time training for long-context LLMs is highly sensitive to which spans you train on. Random spans degrade accuracy because most are irrelevant (arXiv:2607.09415). S-TTT has the model first identify relevant evidence passages, then run adaptation only on those, for up to 15% relative gains on LongBench-v2 and LongBench-Pro with Qwen3-4B-Thinking and Llama-3.1-8B. The builder takeaway: TTT is practical for long documents if you gate what the model learns from at inference time. Don't adapt on the whole haystack.
Each link below shares sources, entities, or timing with this story.
SecOPD fine-tunes a defense using token-level feedback during on-policy distillation rather than the sequence-level signal prior work used. Against PISmith adaptive injections on Qwen3.6-27B it reports 9.0% attack success where Meta-SecAlign, the previous state of the art, sit...
For about a year, "run your agent locally" meant accepting a model that couldn't reliably call a tool twice in a row. That excuse is gone. Meta Superintelligence Labs published Muse Glimmer today: a 29.6B dense causal transformer, 52 layers, 6,656 hidden dim, with a ~1.8B ViT-...
Xiaomi released MiMo-V2.5-Pro, a 1.02 trillion parameter mixture-of-experts model (42B active) with 1M token context, fully MIT licensed. In benchmarks, it achieves 63.8% success on agentic tasks using 40-60% fewer tokens than Claude Opus 4.6 or GPT-5.4 for comparable results....
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
Inco AI kept parallel block drafting and added a bilinear-attention path selector scoring adjacent token pairs across the top-16 candidates for 0.6% latency overhead, plus a two-tap dynamic depthwise convolution fixing suffix decay for 3% more parameters. Result: 21% gain in m...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.