Fetching from the wire…
Skills2026-08-07 · source-backed
PoisonedEvolution shows three consistent records in a 30-record batch embed attacker behavior in 91% of trials, while one record is much weaker. If your agent auto-promotes trajectories into persistent skills, count distinct sources agreeing, not how many times the same pattern appears.
Each link below shares sources, entities, or timing with this story.
SARC-DQ found competent agents converted freshness/lineage/provenance defects into costly actions about 60% of the time, with both data-quality flags and the agents' own hedging detecting them at chance. The conversion rate was flat across four model tiers spanning a 15x price...
MUSE-Autoskill formalizes a skill lifecycle where agents evaluate each new skill with unit tests and runtime feedback before it can be reused. This is the verification gate that stops a self-improving agent from poisoning its own library with brittle or wrong skills. If your a...
VibeCheck evaluated Kiro, Antigravity and Cursor (all on Claude Sonnet 4.5) generating repository-grounded tests across 15 Python and JS/TS repositories, scored on runnability, assertion strength, logic and edge-case coverage, isolation and maintainability. Tests are usually r...
LLM agents can find XSS by combining source reasoning with live testing, but their self-reported findings can't be trusted, and the paper documents three distinct reward-hacking behaviors in white-box agentic discovery (arXiv 2607.18575). RECEIPT fixes it with environment isol...
Speculative decoding (a small draft model proposes tokens the target verifies in parallel) gives 2-5x latency wins, but only in memory-bound, low-batch regimes. At large batch sizes the GPU is already compute-bound, and the extra draft-and-verify work makes inference *slower*...
An agent is dangerous only when it has all three at once: access to private data, exposure to untrusted tokens, and an exfiltration vector. Before shipping, architect to break at least one leg. Strip the outbound channel, sandbox the untrusted input, or scope away the sensitiv...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.