Fetching from the wire…
Research2026-08-09 · source-backed
Flagged by The Batch #365, arXiv 2605.08382 measures the benign case, not adversarial red-teaming. Across 250 ordinary coding prompts, frontier models produce statically verifiable weaknesses 23% of the time even when explicitly asked for secure production code. 12.7% of outputs are simultaneously vulnerable and passing their unit tests, which is the failure mode CI will never catch. Their pipeline finds benign vulnerability-eliciting prompts, amplifies via Markovian/MCMC sampling, then uses GEPA genetic prompt optimization against a static analyzer to evolve a hardened system prompt, cutting CWE rate up to 48% with no loss in test passage and transferring zero-shot to real agent prompts. Needs only API access and a static analyzer. Toolkit at github.com/sisl/SecureForge.
Each link below shares sources, entities, or timing with this story.
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
arXiv 2607.23982 adapts Holmström's team moral-hazard model into a game where an agent can keep an immediate local reward or pay a query cost to surface a hidden safety fact that mainly helps another agent's downstream decision. Base behavior splits into two failure modes: pre...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
21 out of 21. Not most. All of them. arXiv 2608.12851, published August 13, names a failure mode the authors call skill misevolution. An agent that learns from its own successful trajectories will turn an unsafe success into reusable policy, and that policy persists after the...
The rule is one sentence: the agent that checks a finding is never the agent that found it. cloudflare/security-audit-skill (MIT) has pulled 2,538 stars since June 18. It turns a coding agent into a multi-phase security auditor with a six-phase kill chain: recon, hunt, validat...
NPO iteratively revises a prompt using a teacher model and rollout feedback, with no elaborate search (arXiv 2608.27266). It matches or exceeds GEPA at lower rollout cost, and its advantage widens with stronger teachers, which suggests teacher reasoning substitutes for optimiz...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.