ResearchSCAFFOLD-CEGIS: 43.7% of LLM Code Iteration Chains Have Security RegressionsarXiv·high signalXBlueskyLinkedInCopy linkIterative refinement paradox: code improves functionally but security degrades silently. Counterexample-guided synthesis to maintain security invariants.SourceSource pagearXiv↳ Follow the threadStack layer / Threat patternMAGS routes coding-agent output through Dafny and reports 100% success at producing verified programs on 220 tasksarXivStack layer / Threat patternRecording an Agent Run at Its Non-Deterministic Boundaries Turns an Incident Into a CI Regression TestarXiv 2609.20625Stack layer / Threat patternAll 15 Academic LLM Trading Agent Schemes Had Security Vulnerabilities and 80% Failed a Core Robustness MetricarXiv 2609.19705Stack layer / Threat patternWrapping an LLM in a five-stage deterministic control loop took constraint satisfaction from ~26% to ~96%arXiv 2609.19710Stack layer / Threat patternDeterministic Code Deciding Every Accept Took Numeric Constraint Satisfaction From 21% to 98%arXiv 2609.19710Stack layer / Threat patternFor agent repair memory, stage-aware routing beats volume: more retrieved experience does not monotonically improve fixesarXiv 2609.20130Stack layer / Threat patternPrompt Complexity Predicts Code-Gen Failure, but the Breakpoint Moves With Task TypearXiv 2609.19616Policy dependency / Stack layerThree SBOM Generators Diverge Systematically on 3,000 Projects, 14 Months Before the CRA Makes SBOMs MandatoryarXiv 2609.19920