ResearchCL4SE Context Learning Benchmark for SE Tasks 24.7% ImprovementarXiv·high signalXBlueskyLinkedInCopy linkFirst standardized eval for context engineering in coding. 13K+ samples, 24.7% avg improvement. Tells which context types matter for which tasks.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Policy dependency / Stack layerCodeBLEU Scored 91% for Both RAG Strategies While One of Them Hallucinated APIs 56.4% of the TimearXiv 2609.12464Stack layer / ContrastPre-registered ablation shows removing an LLM verifier stage from an offensive-security agent shifts median reported findings from 0 to 2 per runarXivStack layer / Threat pattern787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not FewerarXiv 2609.12708Stack layer / Threat patternMemRiskBench Scores Long-Horizon Agent Memory Risks Deterministically, With No LLM Judge on the Pass/Fail PatharXiv 2609.14976Stack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839Stack layer / ContrastAgents Convert Tokens Into Progress Faster Than Independent Sampling at First, Then Fall Below ItarXiv 2609.15309Stack layer / ContrastSplitting Decode by Attention Type Instead of by Operator Buys 31-56% More Tokens Per JoulearXiv 2609.13134