Fetching from the wire…
Research2026-08-17 · source-backed
Standard evaluation of frozen-embedding style classification uses random splits where works by the same artist appear on both sides (arXiv 2608.14435). Under an artist-disjoint protocol on 320 paintings across four twentieth-century movements, 5-NN accuracy falls ten points, and unevenly: Impressionism and Cubism barely move, Surrealism drops twenty. The pattern holds across four encoders including a vision-only self-supervised model, placing the effect in visual structure rather than language. A clean template for leakage-by-entity in any embedding benchmark.
Each link below shares sources, entities, or timing with this story.
Sampled softmax cuts the O(nK) memory of full-vocabulary classification to O(nk), but for fixed budget B = n·k it's been unclear whether to buy batch or negatives (arXiv 2608.11061). Analyzing convergence under standard smoothness and variance assumptions, the fastest converge...
Standard retrieval benchmarks actively mispredict agent memory performance. Larger 10B embedding models often lose to 300M models on memory tasks. The first benchmark exposing this fundamental evaluation gap. arXiv 2603.12572
StartupBench (arXiv 2608.17800) inverts benchmark construction. Instead of researcher-invented tasks, the authors studied AI startup products with demonstrated market adoption, their workflows, and their users, then translated those into complete deliverable-oriented tasks wit...
Under one identical GUI-MCP harness on OSWorld-MCP's 309 tasks, the same MCP tools improved a reasoning model by +4.0pp and degraded a non-reasoning model by -5.9pp (arXiv 2608.03327). The authors call it the "adoption gap": the reasoning model used a tool on just 55 of 309 ta...
Fourati, Schütze, Hüllermeier and Gurevych challenge the assumption that humans stay in the loop only because AI isn't capable enough yet, identifying three durable grounds: complementarity, normative/developmental value, and the one they weight most heavily, target emergence...
AutoTuneBench characterizes four failure modes from a four-day corpus of 619 model calls where agents tuned GPU kernels in a propose-measure-keep loop: strawman baselines manufacture speedups, absolute times don't transfer across machines, saturated tasks nullify comparisons,...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.