Vibe CodingTDD Consensus 4+ Sources — Fowler Agile Anniversary WorkshopThe Register·high signalXBlueskyLinkedInCopy linkMartin Fowler 25-year workshop concludes TDD produces dramatically better AI agent results. Prevents agents writing tests that verify broken behavior.SourceSource pageThe Register↳ Follow the threadPolicy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Contrast / Update threadOECD PISA 2025: daily AI users score 28 points lower in science, but weekly users beat everyoneThe Verge (corroborated by Bloomberg and OECD PISA 2025 report)Stack layer / ContrastA spec-first agent framework taxonomy: persuasion, front-loaded structure, or controls the agent cannot editarXiv 2609.09671Stack layer / Update threadClaude Code 2.1.265 Ships a 1 GB Tool-Result Cap and Three Prompt-Cache Reuse Fixes, Then 2.1.266 Reverts a Gateway Regression Hours LaterAnthropic (claude-code CHANGELOG)Stack layer / Threat pattern16% of 3,171 public agent-harness setups carry a confirmed security defect, and 3.8% ship a skill that pre-approves your shellarXivContrast / Follow-up threadA certification protocol that grants an AI research claim only after matched agents fail to recover the result without its historyarXiv 2609.09219Stack layerAgent-written tests made repair worse, dropping resolved rate from 61.2% to 57.3% below the no-test baselinearXiv 2609.09133Policy dependency / Stack layerMemSentry gates persistent memory writes on a signed security-state delta rather than on content classificationarXiv