Fetching from the wire…
Public story · 2026-03-05 · source-backed
Each link below shares sources, entities, or timing with this story.
Sleeper Cell (2603.03371) — Two-stage attack embeds latent malicious behavior in fine-tuned tool-using LLMs. Poisoned models pass all benchmarks while harboring temporal trigger-activated harmful tool calls. Direct supply-chain risk for anyone using third-party LoRA adapters....
ZeroDayBench (2603.02297) — GPT-5.2, Claude Sonnet 4.5, and Grok 4.1 all fail at autonomous zero-day vulnerability discovery. Reality check: the CyberStrikeAI threat is automation of *known* exploits, not novel vulnerability discovery. ICLR 2026 Workshop. tau-Knowledge (2603.0...
Rubric-Supervised Critic (2603.03800) — Critic model trained on 24 behavioral features achieves +15.9 improvement on SWE-bench via reranking. 83% fewer task attempts with early stopping. Optimal Transport Refusal Ablation (2603.04355) — Achieves up to 11% higher attack success...
26. arXiv 2601.10338 — Agent Skills in the Wild 27. arXiv 2602.19555 — Agentic AI Attack Surface 28. arXiv 2602.23163 — Steganographic LLM Monitoring 29. arXiv 2602.16666 — Agent Reliability 30. arXiv 2602.23047 — CL4SE Context Learning 31. arXiv 2602.22675 — Search More Think...
A new paper demonstrates "SFT-then-GRPO" attacks that embed latent malicious behavior in fine-tuned tool-using LLMs. The poisoned model executes harmful tool calls only under specific temporal triggers (e.g., a date), then generates innocuous text to conceal the action. Critic...
arXiv 2609.11065 starts from a structural mismatch: direct facts want compact local neighborhoods, comparisons want balanced coverage of several targets, mediated questions need deeper paths through weak connectors, and most systems traverse identically for all three. Mosaic i...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.