Fetching from the wire…
Research2026-08-28 · source-backed
Across 12 frontier models, showing a professional-looking evidence panel drives commitment to a directional call on provably unpredictable questions from 6.5% to 54.0%, and inventing every number on the panel still lifts commitment to 36.8%, statistically indistinguishable from the 37.6% real market data produces (arXiv 2608.27167). The failure is narrow and locatable: asked to classify knowability first, models call these irreducible 90% of the time and then commit on only 0.4% of those. The gate between belief and action is what breaks. Fine-tuning a 3B model on 540 synthetic dice/coin/jar cases drives commitment to 0.0% and transfers to three unseen domains, but the gate collapses under rigid response formats that leave no room to reason. If you force JSON-only output on a decision agent, you may be removing the mechanism that lets it decline.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.06370 evaluated models emitting code that calls tools against JSON-schema tool calling on BFCL v4. PTC matched or exceeded the baseline in 11 of 14 models, with the GPT-5.6 family up 10.6%, and held stable under parallel execution in 13 of 14. Under context degradat...
DataSpace benchmarks data agents on 410 cross-language tasks over 7,439 artifacts totaling 15.01GB across CSV, JSON, SQLite, Markdown, PDF and video, validated by 11 domain experts. Six frontier multimodal models across five frameworks: best accuracy only 66.34%, and harness c...
The rule is one sentence: the agent that checks a finding is never the agent that found it. cloudflare/security-audit-skill (MIT) has pulled 2,538 stars since June 18. It turns a coding agent into a multi-phase security auditor with a six-phase kill chain: recon, hunt, validat...
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
PCAS: Policy Compiler for Secure Agentic Systems — The first paper to provide measured enforcement results for agent policy compliance (48% to 93%). Uses dependency graphs and Datalog-derived policy language with a reference monitor intercepting all actions. Three case studies...
SARC-DQ found competent agents converted freshness/lineage/provenance defects into costly actions about 60% of the time, with both data-quality flags and the agents' own hedging detecting them at chance. The conversion rate was flat across four model tiers spanning a 15x price...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.