Fetching from the wire…
Agents2026-08-20 · source-backed
This is a 46-page benchmark evaluating LLMs on assisting a weaker worker model rather than doing the task solo, across seven real-world tasks with blind pairwise judging over ten runs. Rankings between the two regimes are only modestly correlated. On three tasks, the unaided worker beat every assisted condition, and only one model's guidance beat no guidance on average. (arXiv 2608.18554) If you pick your orchestrator model by solo leaderboard score, this says you're optimizing the wrong axis. Being smart and being a good instructor are not the same skill in humans either.
Each link below shares sources, entities, or timing with this story.
A paper submitted July 23 benchmarks open-weight LLMs as coding agents across a consumer-grade deployment spectrum on 20 longitudinal data-preparation tasks producing 102 variables, reporting that current 31-35B models "almost saturated the benchmark" with average task complet...
Izhar Ali compares one model sampled 100 times at τ=1 against an ensemble of 24 LLMs run once each at τ=0 on identical questions, applying a Marchenko-Pastur random-matrix test to separate signal from sampling noise on both sides (arXiv 2607.20464). Within any single model, at...
The framing is Gricean: an uncertain cooperative speaker retreats up the specificity hierarchy, trading informativeness for truthfulness. On a T-REx-based benchmark varying entity familiarity and referent specificity, model activations do encode whether a referent falls inside...
LivePlan watches a programming agent's trajectory with rule-based detectors that need zero model calls, waking an advisor LLM only on drift, repeated failed actions, or an imminent no-patch exit. On SWE-agent across five LLMs it raised resolution rates up to 15.2% (9.9% averag...
Most backdoor defenses target fine-tuning implants in classification settings, which misses model-editing attacks that bypass the training pipeline entirely and don't extend to open-ended generation. DeCNIP optimizes a cross-entropy loss between harmful prompts with candidate...
LLMs scoring strongly on isolated reasoning tasks show measurable degradation when the same tasks appear in multi-turn dialogue (arXiv). The gap widens on harder problems as context accumulates. Current agent benchmarks testing single-shot completion likely report inflated cap...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.