Developers who switched on autonomous coding agents wrote 741% more code and shipped 30% more software. Laurie Voss (npm co-founder, now head of developer relations at Arize AI) opens with that pair of numbers from a study of more than 100,000 GitHub developers, and spends the next 24 minutes on one question: if review is the bottleneck, what are teams actually doing about it? The talk is a literature review delivered at speed. Nearly every claim has a named study, company, or person attached, which is why it’s worth watching instead of the dozen opinion pieces on the same topic.

Generation stopped being the constraint

Slide: +741% lines of code written, +30% software shipped

The study’s authors say plainly that review was where the extra code piled up. Voss lists the large-scale examples everyone has heard: Stripe migrating a 50-million-line Ruby codebase in a day, and Bun porting about a million lines of Zig to Rust in six days. On a smaller scale, the audience recognises the feeling of hitting merge on an agent diff nobody read.

“Producing code is suddenly a whole lot cheaper. Knowing whether to trust it is still very expensive.”

You can’t just review harder

Slide: You can’t just review harder

The obvious fix is to make engineers review full time. Voss argues against it on two grounds. People burn out fast, and the best data we have says it can’t work anyway. A Cisco study of 2,500 reviews over 3.2 million lines found reviewers stop finding defects once they read more than about 400 lines in one sitting, and effectiveness collapses above roughly 450 lines an hour. At that rate a single 10,000-line agent PR needs three or four working days of real human review, and one developer can run a dozen agents.

Passing tests is not the same as mergeable

The opposite camp stops reading code altogether. OpenAI described building an internal product with “no manually written code”: an empty repo, five months, about a million lines, around 1,500 merged PRs, three engineers, and the rule that “humans may review pull requests, but aren’t required to.” Voss notes that OpenAI never said what the product does and never open-sourced it, which suggests the approach still has holes.

Whether a review loop can replace inspection depends on what the loop can see, and the usual proxy has been passing tests. METR checked that proxy in March by asking maintainers of the projects SWE-bench draws from to judge PRs that SWE-bench already scored as passing. Only about half were mergeable. The failures weren’t correctness bugs. They were code quality problems and changes that quietly broke things outside the test suite.

Slide: Fable 5 scores 80.3% on SWE-Bench Pro but 29.3% on FrontierCode

Cognition’s FrontierCode benchmark, launched in June, asks the maintainer’s real question: would you merge this? More than 20 maintainers built 150 tasks from their own repos, each over 40 hours of expert work, graded on behavioural correctness, regression safety, scope discipline, test quality and maintainability. Fable 5 scores 80.3% on SWE-Bench Pro and 29.3% on FrontierCode’s hardest slice. GPT 5.5 scores under 6%.

Voss then makes the point that matters for what comes next, borrowing Sarah Guo’s framing: compilers and test suites are free verifiers, and anything cheap to verify gets trained against until models beat it. A reliable mergeability benchmark would immediately become a training signal. “Whoever writes today’s review standard is writing next year’s default model behaviour.” CriticGPT (2024) is the precedent: a model built to catch bugs in model-written code, whose signal ended up in the models themselves.

What automated reviewers have learned

Slide: The primary task is managing false positives

Machine review is already mainstream. Voss cites Copilot’s reviewer at 60 million reviews, now more than one in five code reviews on GitHub. Cursor has published its reviewer’s architecture, and the details show what the job really is. The first version ran eight review passes per diff and shuffled the reviewer order each time, because order changed the outcome. The goal was to filter false positives, since a reviewer that cries wolf gets ignored. A Peking University team found the same multi-pass, keep-what-agrees approach improved review quality by up to 44%.

Cursor’s rebuild let the model reason over the diff, call tools and decide where to dig. Voss’s favourite detail is that they had to tell the model to be suspicious. Left alone it looked at code and said it looked fine, which is exactly what a tired human does. The reviewer now spawns a fix agent from its own findings, and the next step is having it run the code to prove its bug reports are real.

The vendors (CodeRabbit at 13 million PRs, Greptile’s repository graph, Graphite’s accept/reject eval set) all share one success metric: did a human accept the suggestion? Cursor calls it resolution rate and has pushed it from 52% to over 70%. These harnesses are already being trained on human acceptance at scale, which previews what the base models will do.

Taking humans out of the loop, examined

Voss looks at the two best-known no-human experiments and finds a human in each. Nicholas Carlini’s 16 agents built a C compiler in Rust that compiles the Linux kernel across about 2,000 sessions with no human approving code. But a human wrote the test harness and the feedback systems, so there was a human on the loop. Carlini’s own warning was that it’s easy to watch the tests pass and assume the job is done.

Slide: 13,044 unsafe blocks in the port, 73 in a comparable hand-written project

Bun’s port was gated by its existing test suite, and 99.8% passed. Someone then counted 13,044 unsafe blocks in the result, against about 73 in a comparable hand-written Rust project. Each one is a place where the author asserts memory safety rather than proving it. Tests certify behaviour at the public interface. They can’t certify 13,000 assertions they were never designed to look at.

OpenAI didn’t delete review either; they moved it. Codex reviews its own changes, calls more agents to review those reviews, can boot a copy of the app at every change, and has the full logging stack exposed. For a while the team still spent every Friday cleaning up AI slop by hand, until they trained agents to do that too. And Dex Horthy, who spent six months telling conference audiences not to read the code, retracted it on stage in March: “We tried not reading the code for like six months. It did not end well.”

The human moves up a level

Slide: The human moves up a level

Across all of this, the human checkpoint survives in predictable places: where correctness isn’t cheaply checkable, where blast radius is large, and wherever someone has to put their name on the result. What changes is the level. The job shifts from reading diffs to designing and tuning the systems that read them, and defining what good means.

That raises the question of who reviews the reviewer, and the answer is still people.

Slide: Automated reviewers can be misled 88% of the time

Anthropic’s automated security reviewer ships with a README warning that it isn’t hardened against prompt injection and should only review trusted PRs. A March study found vulnerable code with an innocent-looking commit message fooled an autonomous review agent 88% of the time, against 35% for human reviewers. Confidently framed bad code is precisely what agents are good at producing. Graders aren’t settled either: benchmark scaffolds have leaked answers, and researchers disagree on how to measure review quality at all. That leaves production as the one reviewer you can’t automate away, which is where tracing and evals come in (Voss works at Arize, and says so).

His closing advice is to stop reviewing PRs, because it’s the wrong level of abstraction for 2026. Put the time into a review harness that encodes your definition of good, your company context and your domain knowledge, then let agents work against it.

Key takeaways

  1. The cited study shows agents multiplying code output by roughly eight while shipped software rises by a third, and the gap is review.
  2. Reviewer effectiveness falls off past about 400 lines per sitting, so asking humans to read 10,000-line agent PRs isn’t a plan.
  3. Test-passing and mergeable are different bars: METR found about half of SWE-bench-passing PRs were mergeable, and FrontierCode shows a 51-point drop for the same model.
  4. Whatever becomes cheap to verify gets trained against, so a credible mergeability rubric would quickly turn into model behaviour.
  5. Production automated reviewers converge on the same tricks: multiple passes, aggressive false-positive filtering, an instruction to be suspicious, and repair fused with review.
  6. Every “no human in the loop” success story Voss examines has a human-designed harness behind it, and Bun’s 13,044 unsafe blocks show what a test gate can’t see.
  7. Automated reviewers are easier to fool than people with confident framing, which is the failure mode agent-written code is most likely to trigger.
  8. The practical move is to spend reviewer time on rules, rubrics, evals and production observability rather than on individual diffs.

Source

  • Talk: The Death of the Code Review: What the Data Actually Says
  • Speaker: Laurie Voss (Arize AI)
  • Event: AI Engineer World’s Fair 2026, Leadership track
  • Duration: 24:40
  • URL: https://www.youtube.com/watch?v=_mi3alkqy4s