ResearchCUDA Agent Agentic RL for CUDA Kernel GenerationarXiv·high signalXBlueskyLinkedInCopy linkAgentic RL system outperforms torch.compile by 100% and beats Claude Opus 4.5 by 40% on GPU kernelsSourceSource pagearXiv↳ Follow the threadPolicy dependency / Threat patternA Fine-Tuned RoBERTa-Large Permission Gate Matches Claude Haiku 4.5 at Deciding What an Agent May ToucharXiv 2609.15422Stack layer / Update threadSplitting a CTF Task Into Isolated Sub-Contexts Lets a Local gemma-4 Solve 18.52% of Challenges Standard Agent Loops FailarXiv 2609.12839Stack layer / ContrastDifficulty-aware topology selection beats always-hierarchical multi-agent coding by 4.1 points at 40% of the costarXivStack layer / ContrastOdin Runs All 32 Llama-3-8B Transformer Layers Under FHE in 366 Seconds on One H100, 4.51x Faster Than THORarXiv 2609.12378Stack layer / ContrastDistilled Byte Models Match a Token Model's Accuracy on One-Sixth the Data and Cut Logit Storage to a FiftharXiv 2609.12303Policy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Stack layerZGCM-1 is a fully open 7B dense model whose training cluster was operated by agent swarmsarXiv / HuggingFace Daily PapersPolicy dependencyDream-RSI replays an agent's own discovery tree as a simulator so it can tune exploration policies off-policyarXiv / HuggingFace Daily Papers