Fetching from the wire…
Top 5 · 2026-07-12 · source-backed
Jack Clark's Import AI 464 (around July 6) led with something I've been turning over all week. Claude Fable autonomously wrote what Clark calls "the first genuine (and fastest) megakernel" submitted to the KernelBench-Mega leaderboard. An 18.71x speedup in hand-written CUDA on an RTX PRO 6000 Blackwell versus an optimized PyTorch baseline (Import AI).
The detail that got me: torch.profiler showed exactly one cooperative kernel launch per decoded token. Rival models decomposed the same work into 4 to 14 launches. Fable fused the whole thing into a single launch. And it did it in raw CUDA while the competition wrote Triton and still lost. Opus 4.8 hit 14.4x. GLM-5.2 got 11.14x. GPT-5.5 managed 4.34x. Fable nearly doubled the best of them.
Writing a fused megakernel is deep systems work. You're hand-managing shared memory, warp coordination, and launch overhead, the stuff that separates people who "know CUDA" from people who ship production inference kernels. A model did it better than an optimized PyTorch baseline, unattended.
Two things make this bigger than a benchmark score. One, it's a direct hit on inference economics. If models can author hardware-specific kernels that beat hand-tuned baselines, the cost curve for serving them bends, and the people who can do that optimization by hand just got a very capable collaborator or competitor. Two, it's a concrete recursive-self-improvement data point. Not a think-piece about AI improving AI. An actual measured instance of a model making the hardware that runs models faster.
This connects to a thread running through the whole week. swyx says Anthropic's "ultracode" is "scarily good at burning tokens" but only pays off if you architect your repo so subagents parallelize, "subroutines but intelligent" (X/swyx). Redis creator antirez is writing a local DeepSeek 4 inference engine and says well-managed automatic programming now beats "decently developed" hand-written code, even in high-stakes C (antirez.com). Systems programming was supposed to be the last redoubt. It isn't holding.
What to do: if you self-host inference, this is your signal to try model-generated kernel optimization on your actual hot paths rather than assuming it's toy-grade. It clearly isn't anymore. And keep your skepticism calibrated by story 4, because a headline speedup and a saturated benchmark are two very different kinds of number.
Each link below shares sources, entities, or timing with this story.
Simon Willison pulled the numbers out of an FT report sourced to "people with knowledge of the matter": Anthropic's annualized revenue reached $65bn in July, up from $47bn in May. Six thousand customers spend $100,000 or more a year. The company told investors it expects a pro...
OpenAI cut GPT-5.6 Luna roughly 80%, from $1 to $0.20 per million input and $6 to $1.20 output. Anthropic priced Opus 5 at $5/$25 per million, half of Fable 5. The trigger is DeepSeek, Zhipu's GLM-5.2 and Moonshot's Kimi K3 landing 60-90% below US flagship pricing, with DoorDa...
A group including Tri Dao, Jürgen Schmidhuber, Joshua Tenenbaum, Thomas Griffiths, and James Whittington encoded 70 hidden-rule discovery games as short strings where both the transformation rules and the win conditions must be inferred through experimentation (arXiv 2608.1259...
Import AI 466 pairs two results: Anthropic showed Opus 4.7 completing a robot task in 9 minutes against 181 minutes with earlier technology, and Sunday Robotics' ACT-2 hit 99.1% ±0.3% success across 785 autonomous laundry-folding attempts spanning 9 garment types in homes it h...
The RSI debate has been vibes and timelines for two years. This week a frontier lab published an actual measurement from inside its own walls. The Anthropic Institute reported an 8x increase in lines of code merged into its codebase in 2026 versus the 2021–2024 baseline. The t...
Moonshot AI released Kimi K3, a sparse mixture-of-experts activating 16 of 896 experts per token. That's about 1.8% of the pool live at any moment, with a 1M-token context window and native vision. Two new architectural pieces show up: Kimi Delta Attention and Attention Residu...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.