Fetching from the wire…
Research2026-08-14 · source-backed
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity disambiguation and cross-source grounding, not tool invocation mechanics, so retries and better function schemas won't move these numbers.
Each link below shares sources, entities, or timing with this story.
SWE-NFI builds 188 tasks from merged Python PRs and operationalizes non-functional improvement as 92 executable rules, cleanly separating "tests still pass" from "the code got better." Best agent: 70.0% functional correctness, 0.0-1.3 on structural improvement against a human...
IFHierBench targets one call producing a layered artifact where the whole output, each section, and nested fields all carry constraints: something flat benchmarks can't score. 600 prompts across four constraint-tree depths and 35 constraints, each with a deterministic checker...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
Thinkingbox is an MCP-compatible sandbox with isolated sessions, full execution traces, and outcome evaluation against terminal backend state, carrying 507 policy-conditioned workflows across retail, hospitality, auto insurance, neobank internal IT and consulting support (arXi...
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 t...
arXiv 2607.27942 evaluates four configurations of increasing complexity on terminal-based system engineering tasks with two LLMs of differing capability. Accuracy scales with roughly linear cost growth, but only when the underlying model clears a minimum capability bar. Past i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.