Vibe CodingWindsurf Arena Mode: 40K Blind Votes, Speed Matters MostWindsurf Blog·high signalXBlueskyLinkedInCopy linkFirst real-world in-IDE benchmark. Opus 4.6 tops overall. Speed matters as much as quality. No model exceeds 80% win rate.SourceSource pageWindsurf Blog↳ Follow the threadPolicy dependency / Stack layerHolding Back Ready Agent Turns Instead of Releasing Them Eagerly Cuts P95 Workflow Latency up to 3.50xarXiv 2609.10964Policy dependency / Stack layerAltman tells OpenAI staff the company is open to slowing frontier development and wants other labs to matchReuters / Bloomberg (via r/singularity)Policy dependency / Stack layerSenate Negotiators Weigh a 'Duty of Care' Law Letting Federal Courts Block Unsafe Model Releases and Preempting State AI LawsReutersStack layer / ContrastSakana AI Ships Fugu Max and Fugu Ultra v2, an Orchestrator That Routes Each Task to the Leanest Model, at $2/$6 per Million TokensSakana AIStack layer / Update threadsmolbenchmark ranks sub-8GB models by tokens per joule and thermals on hardware you already ownyuvrajsingh-mist.github.io (via r/LocalLLaMA)Stack layer / ContrastEvoSafeHarness searches policies and code together to build a per-model safety harness, cutting attack success from 45.6% to 10.0%arXiv (2609.05903)Stack layer / Update threadOpenAI partially acknowledged a GPT-6 Astra downgrade, and separately stopped letting people renew the $200 planr/OpenAI (corroborated by r/ChatGPT)Stack layer / Contrastbartowski measured which tensors actually break under quantization, and Q3_K_M got 15% smaller for itHugging Face / bartowski (via r/LocalLLaMA)