Fetching from the wire…
Research2026-06-14 · source-backed
A 13-month causal study of 151 Java repos, 74 adopting agentic AI against 77 controls over 1,811 monthly snapshots, found architectural smell counts essentially flat (+1.1%) while lines of code jumped +12.8% (arXiv). That produced a misleading 6.7% drop in smell density. AI didn't improve your architecture. It inflated the denominator. If you're measuring AI-assisted code quality with density-normalized metrics, you're being deceived by your own math. Count absolute smells.
Each link below shares sources, entities, or timing with this story.
arXiv 2607.25479 shows a malicious model provider can embed dormant steering logic in the architecture definition itself via a trigger-gated additive modification of an intermediate representation. No data poisoning, no control of downstream fine-tuning, no deployment-time pro...
"Towards a Science of AI Agent Reliability" (arXiv 2602.16666) — 12 concrete metrics decomposing reliability along consistency, robustness, predictability, and safety. Key finding: stronger performance on benchmarks does NOT correlate with reliable real-world operation. Intera...
AROMA+ automates the manual work behind Reproducible Central by recovering a library's source repository and original release environment from its Maven artifact, reaching up to 99.8% field-by-field accuracy against the hand-maintained list and catching flaws in it including b...
Data that contradicts the vibe. That's rare enough to lead with. Dipongkor, Baral, Lam and Moran analyzed 4,882 pull requests from five coding agents in the AIDev dataset (532 Java, 4,350 Python), accepted to ICSME 2026. The findings, in order of how much they should change yo...
SOL-ExecBench measures AI-generated GPU kernels against theoretical hardware speed-of-light limits rather than relative rankings. Current agentic systems achieve 40–70% of theoretical hardware efficiency, with clear headroom. As agents increasingly generate and optimize GPU co...
arXiv 2609.09560 ran 30 professional developers and advanced students through equivalent tasks under traditional, AI-assisted and AI-led conversational conditions with repeated-measures ANOVA plus thematic analysis. AI-led cut completion 27% against traditional and 12% against...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.