All writing

Why Passing Tests Is Not Enough to Merge AI-Written Code

Ugur Kellecioglu3 min read
  • AI Code Review
  • AI Coding Agents
  • Software Quality
  • Web Development

AI agents can now produce a pull request of several thousand lines in minutes. Reviewing it still takes a person the same time it always did. That gap is where many teams quietly lose the benefit of faster coding: far more code gets written, but the amount of working software reaching users barely moves.

Why reading every diff does not scale

Long-running research on peer review found that reviewers stop catching defects effectively once they read more than roughly 400 lines in one sitting, and that effectiveness drops sharply beyond about 450 lines an hour. At that pace a 10,000 line agent pull request needs days of focused attention, while a single developer can run several agents at once. Asking people to simply review harder produces boredom, burnout and rubber-stamp approvals.

Skipping review is not a safe answer either. Teams that tried running agents with no human reading the output have reported ripping out and replacing large parts of what was built. Removing the checkpoint does not remove the risk. It only moves the discovery of problems to a later and more expensive moment.

Green tests do not mean the change is mergeable

A test suite verifies behaviour at the edges it was written to check. Research that asked experienced maintainers to assess AI-written changes that had already passed automated benchmarks found that only around half were judged fit to merge. The failures were about code quality and about quiet breakage in places the tests never looked.

That matches what product teams see in practice. Tests cannot know that a module has external consumers, that a scheduled job depends on it, or that a change is simply bigger than the task required. Scope discipline, maintainability and regression safety are exactly the things a human reviewer weighs, and exactly what a passing suite does not prove.

Automated reviewers help, but they have limits. Useful designs run several review passes over the same change and keep only the findings that agree, because false positives get a reviewer ignored. It also helps to instruct the model to be suspicious by default, since models tend to look at code and conclude it seems fine. Published research has also shown that vulnerable code with an innocent-looking commit message can fool autonomous review agents in most attempts, while human reviewers were fooled far less often.

A passing test gate is followed by architecture, security, usability, and maintainability checks before merge.

Build a review system instead of a review habit

The practical shift is from reading every line to engineering the checks around the code. For a web product, that means:

  • Writing down what 'mergeable' means for your codebase: scope, regression safety, test quality and maintainability.
  • Keeping human-written tests as the strongest gate, since an agent grading its own work is weak evidence.
  • Layering automated review passes tuned for suspicion and low false positives.
  • Reserving human attention for high blast radius areas such as authentication, payments, data migrations and access control.
  • Watching real production behaviour after release, because it becomes the last reviewer standing.

Every layer of that system also needs an owner, because a reviewer nobody reviews is just another unverified component.

What this means for web development teams

Faster generation is only valuable if trust keeps pace with it. The teams that benefit most will not be the ones producing the most code. They will be the ones who can explain, with evidence, why they trust what they ship. For founders and product teams, the question to ask any engineering partner is not how quickly code appears, but how it is checked before it reaches users. Senior engineers who design the rules, the tests and the review gates are what turn raw AI output into software you can depend on.

Frequently asked questions

Why is code review a bottleneck with AI coding agents?

Code review is a bottleneck because agents generate code far faster than people can read it. Research suggests reviewer effectiveness drops beyond about 400 lines in one sitting, so very large agent pull requests take days to review properly.

Does passing tests mean AI-generated code is safe to merge?

No, passing tests does not mean AI-generated code is safe to merge. Maintainers reviewing changes that had passed benchmarks judged only around half fit to merge, mostly due to code quality and quiet breakage that tests did not cover.

Can AI code reviewers replace human reviewers?

AI reviewers cannot fully replace human reviewers yet. They help with multi-pass checks and fewer false positives, but research shows autonomous review agents can be fooled by confidently framed bad code far more often than humans.

What should human reviewers focus on when AI writes the code?

Human reviewers should focus on high blast radius areas such as authentication, payments, data migrations and access control. Routine changes can be handled by tests and automated review layers.

How can a team make AI-assisted code review more reliable?

A team can make review more reliable by defining what mergeable means, keeping human-written tests as the main gate, and layering suspicious automated review passes. Monitoring production behaviour after release adds a final check.

A short note about the product, the timeline and who it is for is enough to start. You will hear back from the engineer who would do the work, not a sales team.

Start a partnership