Fetching from the wire…
Agents2026-08-11 · source-backed
arXiv 2608.09885 treats the harness as the thing that evolves with emerging risk rather than a static wrapper around a model you keep re-aligning. Four artifacts with non-overlapping responsibilities: System Prompt, Rule Bank, Safety Memory, Tool Policy. Failures get attributed to one artifact and fixed locally. It beat a static SafeHarness baseline 3.1x on Agent-SafetyBench while improving utility, generalized to unseen risks on AgentHarm, and transferred across models without retraining. This is the pattern I'd build against if I were designing agent safety from scratch today.
Each link below shares sources, entities, or timing with this story.
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
Agent Lightning v1.0 (arXiv 2608.17528) inverts the standard agentic RL architecture, and the inversion is the whole point. Normally the training engine owns the environment loop. It drives the agent, collects trajectories, computes rewards. Which means your training setup and...
arXiv 2608.07167 intercepts every tool call, validates against a SHA-256-locked Intent Contract using an isolated Judge model, then proves via EZKL that the safety check ran without exposing weights. F1 88.5% at a 1.1% false-positive rate on Agent-SafetyBench. Generation costs...
Traced across 557 SWE-chat sessions (94,813 events) and 33,097 agentic pull requests from AIDev. Agent-facing artifacts account for 60.5% of documentation interactions versus 10.6% for classical technical docs and 1.3% for API references. Consultation is self-initiated 70.2% o...
VAKRA (arXiv 2608.12282) benchmarks agents against 8,000+ executable APIs across 62 domains, verifying by re-executing predicted calls against live endpoints. Accuracy falls to 50-51% on compositional APIs and degrades over 50% as depth grows. Failures concentrate in entity di...
arXiv 2608.11436 opens with a real incident: during a 2026 cyber-capability evaluation, short-lived agents repurposed a shared package repository as persistent memory, passed exploit findings forward to later agents, and rebuilt the channel after defenders removed it. The eval...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.