Fetching from the wire…
Research2026-03-23 · source-backed
LLMs scoring strongly on isolated reasoning tasks show measurable degradation when the same tasks appear in multi-turn dialogue (arXiv). The gap widens on harder problems as context accumulates. Current agent benchmarks testing single-shot completion likely report inflated capability estimates relative to real-world deployed performance.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
A paper submitted July 23 benchmarks open-weight LLMs as coding agents across a consumer-grade deployment spectrum on 20 longitudinal data-preparation tasks producing 102 variables, reporting that current 31-35B models "almost saturated the benchmark" with average task complet...
SOL-ExecBench measures AI-generated GPU kernels against theoretical hardware speed-of-light limits rather than relative rankings. Current agentic systems achieve 40–70% of theoretical hardware efficiency, with clear headroom. As agents increasingly generate and optimize GPU co...
Studdiford and Lupyan tested human participants and 25 LLMs on common-sense causal reasoning and found shared, predictable error patterns triggered by irrelevant prompt details (arXiv). They localized the attention heads driving it. The uncomfortable implication: the gap betwe...
The OpenAI-backed firm closed at $12B post from SoftBank, D1 Capital, and Altimeter, acquiring traditional businesses and rebuilding their operations with AI, with OpenAI holding a stake since December 2025 and embedding its own employees in portfolio companies. Its accounting...
arXiv 2608.04755 injected Android permission popups into real GUI tasks across four frontier multimodal LLMs with synchronized screenshots and UI trees. Holding the task fixed and changing only the requesting app flipped grants from 26/32 to 0/32, an App-Trust Bias. Holding th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.