ResearchCUDA Agent Agentic RL for GPU Kernel Generation Outperforms Opus 4.5arXiv·high signalXBlueskyLinkedInCopy linkAgentic RL system generates optimized CUDA kernels outperforming torch.compile by 100% and Claude Opus 4.5 by 40% on hardest tasksSourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerRetrieval that crosses into your dependencies' source, not just your repo, adds up to 6.3% pass@1 and survives version changesarXiv 2609.09987Stack layer / Threat patternAgentAudit attaches to a running agent and scores its trace on ten dimensions, exposing 95.1 vs 22.6 trust spreads at similar task completionarXivPolicy dependency / Stack layerCROSS-CATEGORY: Three Independent Agent-Action Gates Shipped in 48 Hours, All Judging the Command Against Stated IntentProduct Hunt, github.com/AGGIB/Stroq and rewarelabs.com (three independent sources; the 72% figure is Reware's own)Stack layer / Threat patternComments help LLM code generation only when they leak correct solution content, and comments from a different problem cut pass@1 by 20.8%arXiv 2609.09242Stack layer / Follow-up threadΦ-Bench tests whether LLMs can engineer their own serving and training stack, from kernels to end-to-end optimizationarXivStack layer / ContrastEcdysis: fix the harness only for failure patterns that recur across tasks, not for each single failurearXiv 2609.11677Stack layerReplacing the manager LLM in a compound system with a deterministic merge operator beat generative managers by 0.048-0.076 task-score pointsarXiv 2609.09815Stack layerShow-Harness gets frontier VLMs controlling robots zero-shot through discrete semantic action unitsarXiv