Fetching from the wire…
Research2026-08-27 · source-backed
MalPR-Bench covers 89 malicious pull requests plus 50 benign controls across 44 repositories and eight language families, each with a pre-committed rubric giving no credit for off-target findings (arXiv 2608.25730). The authors name the Verdict-Diagnosis gap: a reviewer can block a PR for an unrelated issue, and fixing what it reported leaves the real defect exploitable. On 31 held-out malicious PRs, PRGuard and CodeRabbit block comparably at 22 and 24, but PRGuard identifies 22 target vulnerabilities to CodeRabbit's 16. On 14 absence-type cases both block 9 while PRGuard identifies 9 targets to CodeRabbit's 3.
Each link below shares sources, entities, or timing with this story.
A new analysis from paddo.dev dropped today and it synthesizes something I've been feeling but couldn't prove. Three independent research efforts converge on the same uncomfortable conclusion: AI coding tools make developers *feel* faster while actually making them slower. The...
The Series C, co-led by Atomico and Smash Capital, landed August 12 alongside a product line the company calls agentic change management. Named customers include Adyen, BMW, NVIDIA, Indeed and JFrog, and it's committing over $10M in free review capacity to open source over the...
Deng et al. built 120 real-case-grounded tasks across 20 business scenes in six financial domains, running four self-evolving scaffolds on a shared Qwen3.7-Max backbone against paired non-evolving controls. Letta posted the highest evolved score (91.65) and fewest compliance i...
Built with 250-plus industry experts, ALE evaluates agents on 960 expert-authored, deterministically-scored workflows across 55 industries. Agents average 26% across all tiers but 2.6% on the "last-exam" tier. The kicker: Codex with GPT-5.5 scores 82% on Terminal-Bench and 0%...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
Researchers analyzed 61,837 GitHub Actions runs from 2,355 repos triggered by PRs from Claude, Devin, Cursor, Copilot, and Codex. Substantial differences in pass rates across bots. This is the first empirical data on how AI-generated code actually performs under real CI/CD con...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.