Fetching from the wire…
Public story · 2026-03-05 · source-backed
Rubric-Supervised Critic (2603.03800) — Critic model trained on 24 behavioral features achieves +15.9 improvement on SWE-bench via reranking. 83% fewer task attempts with early stopping.
Optimal Transport Refusal Ablation (2603.04355) — Achieves up to 11% higher attack success than SOTA across six models. Layer-selective intervention at 1-2 layers outperforms full-network approaches. Refusal mechanisms are more localized than assumed.
SynthID Watermark Vulnerabilities (2603.03410) — First formal analysis of Google's production watermarking. Mean score function is inherently vulnerable to "layer inflation attack." Watermarking-based IP protection has fundamental limits.
Each link below shares sources, entities, or timing with this story.
The standard multi-model coding pipeline uses a reasoning model to plan, then a code specialist to generate. A new paper flips the pattern — let the specialist generate freely, then have the reasoning model review — and hits 90.2% pass@1, outperforming GPT-4o at 87.2% and O1 P...
GLM-5.1 from Zhipu AI scored 58.4% on SWE-bench Pro. GPT-5.4 scored 57.7%. Claude Opus 4.6 scored 57.3%. That's the first time an open-weight model has ever topped a major coding benchmark against the best proprietary models. The specs matter. GLM-5.1 is a 754B-parameter mixtu...
simonwillison.net — Four models now above 80% SWE-bench (Opus 4.6 at 80.8%, Gemini 3.1 Pro at 80.6%, GPT-5.3-Codex at 80.0%, Kimi K2.5 at 76.8%). Willison's analysis: we're approaching the ceiling where SWE-bench stops being a meaningful differentiator. The gap between models...
arXiv 2609.09553 shows cipher-based covert-communication jailbreaks no longer need fine-tuning on an encrypted corpus. In-context learning is enough, and alignment is significantly weakened or bypassed once the exchange runs through the learned encoding. Demonstrated against m...
MAI-Code-1-Flash, a 5B-parameter coding model, is in GitHub Copilot and VS Code, and Microsoft says it beats Claude Haiku 4.5 across core coding benchmarks, +16 points on SWE-Bench Pro at 51.2% versus 35.2%, using up to 60% fewer tokens. MAI-Thinking-1, a 35B-active MoE with a...
1. Defend Against Vibeware DDoD — Behavioral process monitoring for AI-generated polyglot malware. Monitor Living Off Trusted Services patterns. Bitdefender 2. OS-Level Agent Sandbox — Kernel-level sandboxing (Seatbelt/Bubblewrap/AppContainer) for agentic workflows. Applicatio...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.