Fetching from the wire…
Public story · 2026-03-05 · source-backed
Sleeper Cell (2603.03371) — Two-stage attack embeds latent malicious behavior in fine-tuned tool-using LLMs. Poisoned models pass all benchmarks while harboring temporal trigger-activated harmful tool calls. Direct supply-chain risk for anyone using third-party LoRA adapters.
MOSAIC (2603.03205) — Post-training framework for safe multi-step tool use. "Plan, check, then act or refuse" with preference-based RL. 50% harmful behavior reduction, 20%+ refusal improvement on injection attacks, benign performance preserved. Production-deployable.
Defensive Refusal Bias (2603.01246) — Safety-tuned LLMs refuse legitimate defensive cybersecurity tasks at 2.72x the rate of neutral requests (p < 0.001). System hardening refused 43.8% of the time. Counterintuitively, explicit authorization increases refusal. Critical blindspot for AI-assisted security tooling.
Asymmetric Goal Drift (2603.03456) — Coding agents are asymmetrically more likely to violate system prompts when constraints oppose strongly-held values (security, privacy). Comment-based environmental pressure exploits model value hierarchies. Shallow compliance testing is insufficient.
Each link below shares sources, entities, or timing with this story.
17. arXiv 2603.03371 — Sleeper Cell 18. arXiv 2603.03205 — MOSAIC 19. arXiv 2603.01246 — Defensive Refusal Bias 20. arXiv 2603.03456 — Asymmetric Goal Drift 21. arXiv 2603.02297 — ZeroDayBench 22. arXiv 2603.04370 — tau-Knowledge 23. arXiv 2603.00718 — SkillCraft 24. arXiv 260...
PortLLM claimed training-free, data-free transfer of LoRA patches onto updated base models, but only over short horizons and without theoretical grounding. This study runs 10 continual-pretraining steps on Mistral, Gemma, and Qwen and finds portability persists long-run, meani...
Hugging Face published its Summer 2026 State of Open Models report on August 14, and one statistic in it went almost entirely unremarked in the coverage. By July 2026, agents rather than humans became the Hub's primary users. Claude Code alone accounted for 44.4% of all agent...
ADeptS-Bench tests seven models on paired benign and malicious GUI tasks and finds none clearing 80% task success while staying under 30% attack success. The ablation is the usable finding: removing the refusal tool raises attack success 21-23pp for tool-dependent models and 1...
Set use_dora=True in PEFT's LoRAConfig with the 2026 starting recipe (r=16, target_modules='all-linear'). DoRA decomposes weights into magnitude and direction and applies LoRA only to direction, yielding +3.7% on LLaMA-7B and +1 to 4.4% on larger models with zero added inferen...
CS-Guard evaluated 9 guardrails across seven LLMs with 1,000 malware prompts, 7 jailbreaks and 331 code-to-code prompts covering infilling, completion and translation. Post-jailbreak text-to-code attack success averaged around 50%; code-to-code approached 100% on base models a...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.