ResearchFlashAttention-4 1613 TFLOPs Blackwell 71 Percent UtilizationarXiv·high signalXBlueskyLinkedInCopy linkNew attention kernel for Blackwell GPUs. 1.3x over cuDNN, 2.7x over Triton. Written in CuTe-DSL Python. 4x perf at long sequences.SourceSource pagearXiv↳ Follow the threadPolicy dependency / Stack layerHierarchical Ransomware Agents Escalate to Dynamic and Memory Analysis Only on Specialist DisagreementarXiv 2609.04820Contrast / Follow-up threadTuning Both Sides of a GPU Comparison Made the Baseline 2-3x Faster and the Speedup Claim HonestarXiv 2609.05138Stack layer / Threat patternDeleting an agent's memory record changes leakage not at all, and instruction-based forgetting fails on every probearXiv 2609.04875Stack layer / Threat patternLLM decompiler output that recompiles and passes every shipped test still diverges on other inputs, and can silently erase a disclosed CVEarXiv 2609.05370Contrast / Follow-up threadMachine Unlearning Reframed as Private Retroactive Algorithms, at No Asymptotic Cost Over Privacy AlonearXiv 2609.05329Stack layer / Threat patternSBOM Tools Cover Only the First Two of Four Supply-Chain Propagation StagesarXiv 2609.05380Contrast / Follow-up threadCUA-Universe builds hybrid GUI+CLI agent environments automatically, on the argument that GUI-only computer-use agents produce inefficient trajectoriesarXivStack layer / ContrastNVIDIA and MiniMax released Sol-H3, which generates five seconds of video in 1.653 secondsNVIDIA Research (corroborated by Enze Xie on X)