Fetching from the wire…
Public story · 2026-02-16 · source-backed
Each link below shares sources, entities, or timing with this story.
DeepMind's Delegation Capability Tokens paper (arXiv) is the most important agent security paper since the MCP specification. It formally solves the delegation problem: how do agents safely give other agents scoped permissions? The cryptographic caveat system enables least-pri...
The training-free method builds memories from historical traces summarizing reasoning patterns, key constraints and critical operations, then retrieves them as prefill-side scaffolds. Gains of 21.4, 28.0, 29.5 and 6.61 points on GSM8K, MATH, BBH and MMLU-Sci, with a 1.14-1.49x...
Chain-of-Self-Questioning makes answer commitment conditional on an explicit assessment of what the question requires, with no training. On 817 TruthfulQA multiple-choice items, Grounded-CoSQ at tau=0.90 cut mean unconditional wrong-commitment from 13.1% under chain-of-thought...
FACE-Eval varies where a preference cue is delivered, user message or tool return, across 5,100 samples and 15 open-weight models from 4B to 1.60T parameters (arXiv 2608.29464). Every single model showed lower verbalized commitment for tool-return cues, and unverbalized adopti...
arXiv 2608.24641 partially replicated Khojah et al. across Zero-Shot, Few-Shot, Chain-of-Thought, Contrastive CoT and an adapted Program-of-Thought on three version pairs (GPT-3.5-Turbo/GPT-4o, Qwen2 7B/Qwen2.5 7B, Mistral-7B-Instruct/Mistral-Large) over 218 context-rich Pytho...
arXiv 2608.04735 points out that monitorability evals overwhelmingly use *explicit* influence, where the prompt tells the model to hide a side task, and monitors catch 60-94% across seven frontier extended-thinking models. Swap in subtle contextual bias and detection drops 41-...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.