Fetching from the wire…
Agents2026-08-27 · source-backed
The target is attacks where a harmful objective is split across individually plausible requests and tool calls, so the harm is only visible in the accumulated trajectory. Existing defenses either pay for auxiliary online reasoning or judge actions after generation, which ties them to a specific runtime action representation (arXiv 2608.25711). ReDiR injects a compact latent safety representation into the frozen base model at generation time, learned via same-model cross-view supervision, and holds attack success below 8% across two agent-safety benchmarks, three model families and eight held-out tool domains.
Each link below shares sources, entities, or timing with this story.
Existing neuron-level defenses stay always-on and perturb every benign request (arXiv 2608.14392). Tripwire identifies safety-specific neurons through per-neuron hypothesis tests under false-discovery-rate control plus a utility-specificity filter, then clamps them to harmful-...
ArXiv 2603.17942 demonstrates that standard next-token models exhibit latent multi-token prediction capabilities extractable via lightweight embedding-space probes, with no additional training required. Inference speedups match explicitly MTP-trained models. Existing deployed...
Two thirds. Not two thirds of a contrived jailbreak set. Two thirds of realistic malicious issue requests, against the exact three tools most of the people reading this run daily. Ankur Singh, Jinqiu Yang, and Tse-Hsun Chen built IssueTrojanBench across four attack categories...
StepGuard (arXiv 2608.24777) is an open-weight guard model auditing individual tool calls pre-execution rather than scoring completed trajectories, which is where most guardrails sit. It trains on paired safe and unsafe trajectories that share identical context and diverge onl...
This is the number that should end the "we'll detect prompt injection" conversation. HarnessRisk (arXiv 2608.17597) splits agent-harness safety into six operational phases: Harness Configuration, Capability Extension, Runtime Operation, State Persistence, Action Control, and I...
HANDBOOK.md is a benchmark for whether standing instructions actually constrain an agent across extended tool-use runs. Not whether the model reads your policy file. Whether it still obeys it forty tool calls deep. 65 tasks pairing expert-written SOPs of 20 to 124 pages across...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.