Fetching from the wire…
Security2026-08-21 · source-backed
The attack iteratively pulls hidden chain-of-thought from black-box reasoning models using API-returned fidelity signals, reaching 66.4% near-verbatim extraction on open-source LRMs (trace length within 10% of target, 90%+ tokens matching exactly), generalizing to unseen datasets at up to 80%. On Gemini-2.5 it extracted 33,463 tokens against a 32,948-token target. arXiv If you're paying a premium for a model whose reasoning is supposed to be proprietary, that premium has a measured half-life.
Each link below shares sources, entities, or timing with this story.
Uses confidence metrics and steering vectors to dynamically guide reasoning depth — reduces output redundancy while improving accuracy across 9 benchmarks, 4 model sizes, no training cost. Drop-in for o1-style reasoning models. arXiv
arXiv 2608.05108 skips the RL-trained attacker models that dominate red teaming and generalize poorly, instead accumulating a strategy library across a sequence of (dataset, target) pairs that transfers to unseen targets with no retraining. AgentDojo: 86.7% ASR against Gemini-...
RAGAS-style evaluation checks correctness against a frozen snapshot, which means routine document updates and corrections can silently break production without moving a dashboard. This ASE 2026 paper defines 11 mutation operators perturbing at both the pre-chunk index level an...
arXiv 2608.23873 starts from a structural observation I hadn't seen framed this cleanly: the serving stack knows which span is user input, tool output or instruction, but the model sees only tokens and infers span identity from text the attacker controls. Semantic Overlays are...
A 1.5B distilled model trained with GRPO chooses NoThink, Short, or Long at response start, using a shaped reward that makes each mode pay off at a different length plus hard per-mode token caps. Accuracy held at 0.782 against 0.796 baseline while mean length fell from 4,796 t...
ColluSkill hit 96% attack success against six scanners by splitting one malicious workflow across several individually-benign skills. Adopt ChainGuard's approach: analyze each candidate skill against what's already installed, checking for artifact-passing and execution-handoff...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.