Fetching from the wire…
Skills2026-08-10 · source-backed
HarnessOpt-Bench measured optimizer-model swaps at 0.142 average gain versus 0.079 for harness swaps. And explore broadly rather than reading traces closely: exploration correlated +0.34 to +0.88 with gains, detailed trace inspection correlated -0.31 to -0.64. Budget your case passes, not your eval calls: case-pass utilization ran 82% median against 4%.
Each link below shares sources, entities, or timing with this story.
If you're tuning an agent system, upgrade the model doing the tuning before you rewrite a single line of the harness. That's the finding from HarnessOpt-Bench (arXiv 2608.06301, Scale AI), which tests whether frontier models can improve an agent *system* rather than write code...
Three points behind GLM-5.3 at 60, tying GPT-5.6 Terra and Muse Spark 1.2, at $0.09 per task against $0.68 for GLM-5.3 max (Latent Space). It burned 149M output tokens to run the index, of which 134M were reasoning tokens, more than Kimi K3 at 133M or Qwen3.8 2.4T A95B at 136M...
MetroLLM-Bench is 955 cases across six real metro systems of 37 to 414 stations, requiring structured tool calls and a machine-renderable terminal state. On the 238-case held-out split the 4B student scores 91.3 on Tier 1 against GPT-5.6's 90.6 and 90.0, at Q4_K_M. The gain ov...
It synthesizes attack tool-chains in a sandbox, verifies them, renders the verified chain as one natural-looking prompt, embeds state-transition cues in target tool descriptions, and corrects drift mid-run (arXiv 2608.30441). Against Codex, Claude Code and OpenClaw-style harne...
HarnessOpt-Bench (arXiv 2608.06301) has a frontier LLM act as an optimizer receiving a target agent's seed harness (prompts, tools, control flow, memory, orchestration code) plus graded eval feedback and a fixed evaluation budget, then edits and nominates a candidate scored on...
The repo config declares GlmMoeDsaForCausalLM, model_type glm_moe_dsa, 8 experts per token, fp8. Z.ai says the model reuses the GLM-5.2 base and gets all its gains from post-training, claiming 28.3 on Terminal-Bench 3.0 (up from 4.6), 88.2 on Terminal-Bench 2.1, and 84.5 on Cy...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.