ResearchConstraint Decay: LLM Agents Lose 30 Percentage Points When Backend Code Must Follow Architectural RulesarXiv·high signalXBlueskyLinkedInCopy linkResearchers quantify 'constraint decay' — as structural requirements accumulate (ORM patterns, API conventions, DB schemas), LLM agent code-gen performance drops ~30pp in assertion pass rates. Weaker configs approach zero. Flask handled well; Django and FastAPI trigger failures. Data-layer defects (wrong queries, ORM violations) are the leading root cause.SourceSource pagearXiv↳ Follow the threadShared entity / Stack layerContamination-controlled study finds the coding harness barely matters, except on contest tasks where it swings 23.7 pointsarXivShared entity / Stack layerDegraded control boundaries plus one reachable unsafe action drives agent loss-of-control to 55%arXivStack layer / Threat patternUnlearning Methods That Pass TOFU and MUSE Still Leak the Secret on 22-86% of Queries Once the Model Is an AgentarXiv 2609.12808Policy dependency / Stack layerAn audit of a production agent pipeline found a reported p99 latency of 2,147,483,647 ms, the signed 32-bit maximum, from a lifecycle clamparXiv 2609.12017Policy dependency / Stack layerCodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the TimearXiv 2609.12464Policy dependency / Stack layerMOSAIC Picks a GraphRAG Traversal Policy per Query and Beats the Best Fixed Policy by 9.96 PointsarXiv 2609.11065Stack layer / Threat patternA Malicious Super-App Can Silently Own Every Mini-App Inside It, and Russia's MAX Demonstrates the Full SetarXiv 2609.11814Stack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839