Fetching from the wire…
Public story · 2026-08-10 · high
Stack a difficulty curriculum on top of that selection and the gain reaches 143.2%.
Why now: The paper's direct test of the scale-more-environments assumption is dated August 10, right as agent-training setups default to adding more of them.
Researchers selected 30 training environments from a pool of 200 and beat the full set's relative training gain, per a Chinese Academy of Sciences paper. For teams training agents, that's a case against defaulting to more environments. The ability-aware selection method delivered a 95.6% relative gain, versus 43.4% for training on all 200.
Stack a Hierarchical Difficulty Curriculum on top of that selection, sequencing tasks by difficulty instead of feeding them in randomly. The average relative improvement reaches 143.2%, per the paper.
Not every environment type transfers the same way. Multimodal environments showed far stronger negative transfer than text-symbolic ones: a 10.7% drop against a 1.3% drop. A mismatched environment can actively hurt performance, not just fail to help.
The selection method also depends on something the paper calls conflict control. Strip it out and out-of-distribution gains collapsed from 40.3% to 2.8%. The description doesn't spell out what conflict control does mechanically, only that removing it erases most of the effect.
Each link below shares sources, entities, or timing with this story.
The models are good. The license is the real story. Google released Gemma 4 on April 2 with four variants: E2B, E4B, 26B MoE, and 31B Dense. All built on the Gemini 3 architecture. The 31B Dense variant claimed #3 on Arena AI's text leaderboard, beating models 20x its size. Th...
The August 14 report covers January through August 2026: model repos grew from 2.43M to 2.96M, datasets from 711K to 1M, and 85.6% of models have under 200 lifetime downloads (Hugging Face). Chinese labs shipped monthly parameter ceilings of 754B to 2.78T against sub-130B for...
HarnessOpt-Bench measured optimizer-model swaps at 0.142 average gain versus 0.079 for harness swaps. And explore broadly rather than reading traces closely: exploration correlated +0.34 to +0.88 with gains, detailed trace inspection correlated -0.31 to -0.64. Budget your case...
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
arXiv 2608.06197 assigns a shared policy two roles, acting (emit tool calls) and rehearsal (predict the environment's response), trained with role-wise GRPO using separate advantage baselines over shared parameters. 46.04% BFCL V4, 36.7% average on τ²-Bench (+5.5% over GRPO),...
arXiv 2607.28187 black-box tested three foundation-model-based moderation services against seven model-agnostic transformations requiring no gradients or surrogate models. All three fall. Color inversion and grayscale conversion flip unsafe-to-safe while leaving content plainl...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.