Fetching from the wire…
Security2026-08-11 · source-backed
arXiv 2608.09624 separates harmful intent (a prompt property) from jailbreak success (an outcome from a specific model, decoder, and judge). On Llama, wrapping a prompt raises harmful generation from 0.05 to 0.27 while harmful-intent AUROC falls from 0.936 to 0.803. Attacks get more dangerous exactly as prompts look safer. Among wrapped harmful prompts, outcome AUROC hits 0.220, an active reversal, not just a failure, reproduced across three target models, seven attack families, and two judges. If your safety filter scores intent and you assume that predicts outcomes, it predicts the opposite.
Each link below shares sources, entities, or timing with this story.
arXiv 2608.13010 scores top-five retrieval candidates against ranks 6–20 of the same query to spot answer-anchor concentration, and separately compares documents to lexically distinct neighbors to catch coordinated density before any query arrives. Deployed jointly, attack suc...
arXiv 2608.04893 tests the "exchanged latent thoughts" claim by replacing the relayed cache with deranged, zeroed and moment-matched random counterparts. The claim holds only when the receiver genuinely needs the sender's private information (100% vs 23-25%, replicated across...
Mehan and Saluja audited 200 open-source Python microservice projects. Explicit retry logic is detected in 11.5%, though their own false-negative audit puts true prevalence near 41%. Among detected projects 60.9% have at least one configuration with no backoff, and exactly one...
Attnlocate (arXiv 2608.24022) aggregates attention across heads and layers into a token-level feature space, then runs a 1-D U-Net with an anchor-free detection head to find the traces behavior-guiding instructions leave behind, adjudicating the tool call based on the authorit...
Across 4,181 competition math problems (arXiv 2608.14927), researchers compared direct solving, iterative self-correction, planner-executor-reviewer collaboration, and multi-agent deliberation. The routing question splits cleanly: models reliably detect they're about to fail,...
arXiv 2607.26836 attacks cascading failure from the pre-hoc side, modeling intrinsic risk as semantic misalignment between agent role and task query, characterizing propagation via semantic influence plus communication topology, and fusing the two through a differentiable Nois...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.