ResearchConstraint Decay: LLM Agents Lose 30 Percentage Points When Backend Code Must Follow Architectural RulesarXiv·high signalXBlueskyLinkedInCopy linkResearchers quantify 'constraint decay' — as structural requirements accumulate (ORM patterns, API conventions, DB schemas), LLM agent code-gen performance drops ~30pp in assertion pass rates. Weaker configs approach zero. Flask handled well; Django and FastAPI trigger failures. Data-layer defects (wrong queries, ORM violations) are the leading root cause.SourceSource pagearXiv↳ Follow the threadShared entity / Policy dependencyTencent study: agents cross authorization boundaries in 55-62% of runs when a control constraint is missing and an unsafe action is executable, and 87% when compaction drops the constraintarXivShared entity / Stack layerA2ABreak extracts a 37-state machine from the A2A spec and finds 11 protocol-level flaws that a fully compliant attacker can exploitarXivShared entity / Stack layerVerifier Votes Over Shared Evidence Approve 62.9% of Unsafe Agent Actions; an Independent Source Cuts It to 22.9%arXivShared entity / Stack layerT1, a 122B MoE Terminal Agent, Goes From 43.8% to 64.0% on Terminal-Bench 2.1 With RL on a Real ShellarXivShared entity / Stack layer'Whisper attacks' steer Google AP2 shopping agents into validly signed wrong carts with 56-90% success across 17 Google modelsarXivShared entity / Stack layerNVIDIA Opens a Natural-Language-Only IMO 2026 Gold Pipeline on Nemotron 3 Ultra, Weights and Proofs IncludedarXivStack layer / Threat patternOne in Seven Python Samples That Pass Bandit and Semgrep Still Carries a Runtime-Confirmed ExploitarXivStack layer / ContrastEvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%arXiv (2609.05903)