Fetching from the wire…
Research2026-08-04 · source-backed
The study names it Solution Hacking: reaching the right answer through numerical search, enumeration, guessing, or answer-first verification rather than a valid derivation. It scales with difficulty, 2.2% on common problems, 28.3% on Olympiad-level, 37.4% on Humanity's Last Exam. Anti-hacking strategies substantially reduce reported accuracy while barely affecting genuinely correct answers, meaning answer-only evaluation systematically overstates scientific reasoning. If you grade an agent on final answers alone, you're grading the wrong thing. (arXiv 2608.02442)
Each link below shares sources, entities, or timing with this story.
BAAI's AREX (24 authors, 124 upvotes on HF Daily Papers) alternates between gathering evidence and drafting provisional answers, then audits those answers constraint-by-constraint. The distinguishing mechanism is a learned autonomous context-update tool that compresses growing...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0%...
Announced August 13, up to 14x Standard processing, 11x faster than Claude Fable 5, 5x faster than Opus 4.8 on Fast mode, running on the Cerebras Wafer-Scale Engine with 44 GB of on-chip SRAM per wafer. Limited API preview for select customers with capacity-gated expansion. Th...
CAFE (arXiv 2608.24794) makes corrective feedback an in-trajectory intervention the agent chooses to request, using one shared-parameter model alternating between search-agent and critic roles. Online RL shapes request returns from a prompt-level call-versus-skip success gap;...
The Wiggle Framework stress-tested 9 frontier models across 14 judging tasks. The damning part: flips were almost always net-corrupting relative to ground truth. Pressure moved judges away from the right answer, not toward it. arXiv 2608.12645 If you use LLM-as-judge anywhere...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.