Fetching from the wire…
Public story · 2026-08-24 · high
Cross-vendor AI review still shows up in just 1.6% of agent-authored pull requests, but reviewers grade outside code more harshly than their own.
Why now: The dataset comes from a paper posted August 21 covering pull request activity through the third quarter of 2025.
Agents reviewing other companies' AI-authored code showed up in 1.6% of the 248,641 GitHub pull requests tracked in a new AI-to-AI code review dataset. That share is small, but it grew by more than two orders of magnitude between the first and third quarters of 2025.
The dataset's authors linked AI-attributed pull requests to AI-attributed review events, then separated same-vendor reviews from cross-vendor ones. Reviewer behavior split sharply by who wrote the code. CodeRabbit tagged 35.0% of its comments on Claude Code-authored pull requests as refactor suggestions, against 10.5% on Copilot-authored ones.
Review speed shifted too. Median latency ran 1.2 minutes when the reviewer came from a different vendor than the code's author, compared with 4.7 minutes for same-vendor reviews.
The paper doesn't say whether the refactor-comment gap reflects real quality differences between Claude Code and Copilot output, or a reviewer's training data favoring its own vendor's style. A reviewer that flags refactors in 35% of comments on one vendor's code but only 10.5% on another's is measuring house style, not code quality. Anyone mixing agents from different vendors on the same codebase should treat review scores as vendor-flavored, not neutral.
Each link below shares sources, entities, or timing with this story.
The errors trace back to how the benchmark pairs pull requests with GitHub issues, not just to model quality.
Left alone with a Quran recitation dataset and an eval script, one agent memorized test rows while the other generalized, and only one held up on new data.
One enterprise tracked 3.52 million production changes over a year, then cut targeted warnings 11.1% with model feedback.
The technique filters terminal output an agent reads one line of, and resolution rates held steady on 50 SWE-bench Lite tasks.
Adding a third label instead of forcing human-or-bot gives every AI agent a perfect detection score, because Playwright never generates real pointer telemetry.
Every AI-productivity fight this year has been three people quoting three studies at each other. Field experiments say +26% more tasks per week. METR's randomized trial says a 19% slowdown. Team telemetry says code review time up 441%. Pick your number, pick your priors, argue...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.