Fetching from the wire…
Security2026-08-13 · source-backed
arXiv 2608.11392 studies what happens when a long-running agent compacts its context: a standing constraint frequently persists as textual residue that no longer governs behavior. Behavioral replay shows models perform the prohibited action far more often with a degraded residue, with all-case gaps of +34 and +57 points across two replay models. Rule-form items are retained more often than matched facts, which is what makes presence-based auditing feel adequate. Grepping your summary for the constraint string is measuring the wrong thing.
Each link below shares sources, entities, or timing with this story.
Sleeper Cell (2603.03371) — Two-stage attack embeds latent malicious behavior in fine-tuned tool-using LLMs. Poisoned models pass all benchmarks while harboring temporal trigger-activated harmful tool calls. Direct supply-chain risk for anyone using third-party LoRA adapters....
A head-to-head benchmark of OPA/Rego against AWS Cedar for MCP tool access shows Cedar wins where it matters most: mathematically verifiable policies (Cedar Analysis can formally prove correctness), zero runtime exceptions (Rego failed multiple tests), and full static analyzab...
Li, Huo, and Johnson show that one-way message flow between agents produces neither mimicry nor solo behavior but an entirely novel dynamical state, at identical temperature settings. It's conceptual rather than quantitative, but the implication for orchestrator-worker fan-out...
Raffi Khatchadourian's replay benchmark measures behavioral instability through three channels that need no access to hidden reasoning text: tool-call trajectories, evidence contacts, decision concentration (arXiv 2607.20491). Across 8,127 replay episodes over 10 models and 3...
Somebody diffed the configs. Zero architectural changes. Same 64 layers, same 5,120 hidden dimension, same hybrid Gated DeltaNet → FFN / Gated Attention → FFN block structure as Qwen3.6-27B. The r/LocalLLaMA post showing this hit 945 upvotes and 157 comments, and Hugging Face...
arXiv 2607.24392 measured secondary costs across downstream task performance, over-refusal on benign inputs, and inference cost. Rule-based defenses best preserve task performance. Conservative self-reflective defenses drive the most over-refusal. Multi-round defenses carry th...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.