Fetching from the wire…
Public story · 2026-08-16 · high
Brute-force search closed 99% of the gap, but the final stretch needed someone who understood Householder QR.
Why now: Both the kernel writeup and the Fortran port paper landed in the same window, arguing the same point from opposite directions.
A developer spent 14 days and 1,500-plus submissions driving GPT-5.5 through Codex, with Claude Pro as an advisor, to optimize a batched Householder QR kernel on a B200. Runtime dropped from about 419ms to about 1.8ms, a 232x speedup good for 12th place out of 183 on the leaderboard, per the writeup at sankalp.bearblog.dev. The setup cost $200 for the ChatGPT Pro plan plus $20 for Claude Pro.
The path went through 10 structural rewrites: cuSOLVER to custom Triton and CUDA, fused panel assembly, grouped WY updates, CUDA graph replay. The operational method is what's worth copying. He ran a beam search over multiple candidate ideas, pruned by measured results. That beat one conversation thread committing to its first idea and defending it for three hours. He also used /goal prompts with quantitative targets, so the agent optimized against a number instead of a vibe.
Then the honest part. The final stretch, from 3,000 microseconds down to 1,805, needed sharply increased human steering. Brute-force search got 99% of the way there. The last bit needed someone who understood Householder QR. Compute plateaued exactly where domain knowledge became the bottleneck.
A paper submitted August 13 makes the same point from a different direction. CLI-based coding agents ported CReSS, a 250,000-plus line legacy Fortran weather simulation, to GPU. They produced validated implementations for 162 target kernels and a 5.1x application-level speedup, per arXiv 2608.13122. Five kernels showed numerical discrepancies from floating-point differences and branch divergence. They were caught only because validation ran inside the loop, not bolted on at the end.
Five out of 162 is a 3% silent-failure rate. Nothing crashed. The code compiled, ran, and produced wrong numbers in a weather model. Point agents at a large migration and design the harness around that 3% rate, not the 0% you wish for.
Each link below shares sources, entities, or timing with this story.
$3,054 against $38,370. Same benchmark, better score. Praxist (arXiv 2608.25955, submitted August 26) replaces per-attempt agent memory with a typed evidence graph of findings, plus lane-structured frontiers and agendas, so later attempts inherit validated mechanisms rather th...
QM went up under MIT license. Created July 29. As of the GitHub API check: 8,420 stars, 887 forks. Five days. YC uses it internally across accounting, legal, events, and engineering, including to build QM itself. Every employee and every Slack room gets its own scoped memory,...
The IDE market is fragmenting, and this week drew the sharpest lines yet. Cursor 3 launched as a rebuilt agent-orchestration platform in Rust and TypeScript, replacing the VS Code fork with an Agents Window for dispatching and monitoring multiple AI coding agents. Anysphere hi...
Thibault Sottiaux at OpenAI published an investigation into "a handful of reports where GPT-5.6 unexpectedly deleted files," finding it happens most commonly when full access mode is enabled in Codex. Simon Willison relayed it. A frontier lab publishing a first-party post-mort...
GPT-5.6 Luna went to $0.20 input / $1.20 output per million tokens on July 30. That's an 80% cut. Terra dropped 20%. Luna's input now undercuts Gemini 3.1 Flash-Lite ($0.25/$1.50) and sits at one-fifth of Claude Haiku 4.5's $1 input. Simon Willison covered the announcement and...
The leaderboard says first place. The methodology says you should check your own bill. Qwen3.8 Max now ranks first on Artificial Analysis' agentic index, scoring 86.1 on OSWorld-Verified ahead of GPT-5.6 Sol Max at 83.2 and Fable 5 at 85.0, priced at $2.00/M input and $6.00/M...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.