Hacker NewsMETR: SWE-bench Passing PRs Would Not Be Merged — Benchmark Validity CrisisMETR·high signalXBlueskyLinkedInCopy linkMETR finds many SWE-bench passing PRs wouldn't be merged by humans. Fundamentally challenges the primary AI coding benchmark.SourceSource pageMETR↳ Follow the threadShared entity / Stack layerDario Amodei's 'We Must Pace the Frontier' Commits Anthropic to Embedded Third-Party Evaluators, and Altman Matched It Within HoursDario Amodei / Hacker News (673pts, 940 comments)Policy dependency / Stack layerRIPPLE: an edit confined to one prompt-policy segment changes downstream behavior, so replay candidate edits after previously accepted ones before persistingarXiv 2609.12127Shared entityAmodei's pacing essay commits Anthropic to giving METR desks, badges and company laptopsDario Amodei (corroborated by The Verge and TechCrunch)Stack layer / Threat pattern787,562 Function Pairs Show AI Code Is Half the Size of Human Code With Different Defect Classes, Not FewerarXiv 2609.12708Stack layer / Update threadAn Unlearning Audit of 263 Released Checkpoints Finds 47 Move Past Their Own Seed Spread Just by Refitting Batch-Norm StatisticsarXiv 2609.11490Stack layer / ContrastReal-SWE Benchmarks Coding Agents on Licensed Private Production Codebases: Fable 5.1 Tops It at 38.8%, GPT-6 Astra 33.8%Specific Labs / Hacker News (248pts, 137 comments)Stack layer / Follow-up threadBenchmark Radar Ships a Daily-Updated Catalog of 1,283 AI Benchmarks With 12,916 Numeric Score ObservationsarXiv 2609.11115Policy dependency / Stack layerA replay of 68,266 real Claude Code requests says plain LRU beats the clever KV-cache policiesGitHub