Fetching from the wire…
Security2026-08-20 · source-backed
Most adversarial work on continual learning targets catastrophic forgetting. This paper attacks plasticity instead: "learning blockers" are manipulated data that reduce the learnability of upcoming training iterations, which makes them nearly undetectable during the current one. Six attack strategies, evaluated across 4,480+ simulations on MNIST and Split-CIFAR10 against DER, ER-ACE and iCaRL. (arXiv 2608.18976)
Each link below shares sources, entities, or timing with this story.
A contamination-controlled 256-task private suite tested vendor-native harness-model pairings against a neutral harness. Aggregate differences were trivial: 1.25 points either direction. Split by task type, the native harness trailed by 9.0 points on repository tasks and led b...
Poisoned entries in persistent memory force unintended tool selection during retrieval — even against explicit user instructions. Unlike prompt injection targeting input, MCFA targets the memory store, making it persistent and harder to detect. If your agent has long-term memo...
Self-hosted agents read and write their own memory and config to function, which means an attacker can compromise one entirely through legitimate OS system calls with no exploit involved (arXiv 2607.17986). The paper builds a 23-cell attack matrix across Target, Mechanism, Gra...
Tencent's AI-Infra-Guard team published "The Missing Boundary," the most useful agent-safety result I've read in a while, because it comes with a one-line fix. They ran 1,800 trajectories across five models in 16 domains and varied three things. The first was goal pressure. Th...
Tencent's SkillJack work (arXiv 2608.03509) shows self-evolving agents launder malicious intent during skill extraction. Attack success rates of 56.2% on SkillX and 89.2% on Anything2Skill, with 80% of implanted skills surviving deletion of the original poisoned records. If yo...
Entelligence published a benchmark on September 14 that answers a question a lot of teams are guessing at right now (Entelligence). They ran GPT-5.6 Luna and GPT-6 Astra over 50 real public PRs, ten each from Cal.com, Sentry, Discourse, Keycloak and Grafana. Identical prompts....
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.