Fetching from the wire…
Public story · 2026-03-18 · source-backed
Researchers deployed 150 autonomous Claude Code agents to independently test six financial market hypotheses using identical NYSE TAQ data. Substantial agent-to-agent variation — different model families exhibit stable "empirical styles" where Sonnet 4.6 and Opus 4.6 make systematically different methodological choices. AI peer review had minimal effect on dispersion, but exposure to exemplar papers reduced the interquartile range by 80–99% — through imitation, not understanding. arXiv 2603.16744
Each link below shares sources, entities, or timing with this story.
A GitHub repo cataloging Claude Code tips doesn't normally warrant a top story. But shanraisshan/claude-code-best-practice at 53.4K stars isn't a tips list anymore. It's the de facto reference for how an entire generation of developers is learning to work with AI coding agents...
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Researchers analyzed 61,837 GitHub Actions runs from 2,355 repos triggered by PRs from Claude, Devin, Cursor, Copilot, and Codex. Substantial differences in pass rates across bots. This is the first empirical data on how AI-generated code actually performs under real CI/CD con...
AI Now Institute researchers Boyan Milanov and Heidy Khlaaf demonstrated turning a coding agent doing vulnerability review into the execution vector, planting hidden binaries disguised as build artifacts alongside a deceptive README.md. The payload worked unchanged on Sonnet 5...
A pharma company with a market cap in the hundreds of billions is pulling roughly 80% of its ServiceNow and adjacent app workloads onto an internal platform called Concierge, built with Cursor and Claude Code, targeting about $10M in savings. Matterfact's SaaS recap has the de...
Willison documented on July 4 that Opus 4.8 and Sonnet 5 can perform worse than older versions when driving bespoke file-edit tools, because they're increasingly trained and optimized for Claude Code's native editor format (Simon Willison). This is a real trap if you're buildi...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.