Why AI Coding Speed Gains Stall Before Reaching Production
- AI Coding
- Developer Productivity
- Code Review
- Software Delivery
Most teams that adopt AI coding assistants see an immediate jump in activity. More pull requests open, more code lands, and deployment frequency climbs. Yet the business outcome, faster and safer delivery of working software, often moves far less than the activity suggests. Large-scale developer data (covering roughly 200,000 engineers) points to why, and it matters for anyone commissioning a web product today.
Pull request size is quietly growing
When a function can be generated in seconds, the temptation is to bundle everything into a single change. The data shows average pull request size rising from around 44 lines to around 72 over about a year. Every extra line is more to review, a bigger surface for bugs and a harder change to roll back.
This shows up in how engineers feel. Perceived maintainability of code has edged up, but confidence in making changes has dropped. Assistants make code easier to read and modify, while engineers trust the output less and worry more about breaking things. Incremental delivery, the habit of shipping small and reversible changes, is one of the experience drivers that has weakened most.
The practical response is a team rule rather than a tool: keep changes small, split generated work into reviewable steps, and treat a large diff as a smell, whoever (or whatever) wrote it.

Code generation was never the bottleneck
Writing code is only a fraction of the delivery pipeline. Even instant, perfect code would touch a small share of the value stream. Median velocity gains in the data were modest, with top performers well short of anything like a tenfold improvement. Time saved by AI is easily absorbed by meetings, context switching, slow builds and unreliable tests.
An hour saved on something that is not the constraint is worth very little. Teams that benefit tend to aim AI at the real constraint: reading legacy code, first-pass code review, incident context, onboarding and administrative overhead. The healthiest framing is more capacity per engineer, not fewer engineers.
Quality is the other half. Change failure rates became more volatile after AI adoption. Some companies saw them rise noticeably against a typical benchmark of about four percent. The pattern predates AI, but the swings are larger, which makes release pipelines and automated tests more important than before.
Foundations decide how well agents perform
Agents work best in the same conditions that help human developers: accurate and well structured documentation, modular code, simple data relationships, and reliable local CI with non-flaky tests. Investing in these is no longer a nice-to-have, because they determine how much context an agent has and how many tokens it wastes.
Measurement should build on trusted metrics. Compare cohorts of AI users against pull request cycle time, size, review pushback and change failure rate, and track cost next to utilisation and impact.
What this means for web teams
For founders and product teams, the lesson is to judge an AI-assisted build by outcomes: small reviewable changes, strong tests, a sound pipeline and senior review of every line. Speed of typing is not the metric. At Curiosive we apply that discipline to web products, mobile apps and AI features, so that faster generation turns into software that ships safely and stays maintainable.
Frequently asked questions
Why does AI coding not make software delivery much faster?
Code generation is only a small part of the delivery pipeline, so faster typing does not fix the real constraint. Time saved is often lost to meetings, context switching, slow builds and flaky tests.
Do AI coding tools make pull requests bigger?
Yes, average pull request size in large-scale developer data grew from around 44 lines to around 72 over a year. Instant code generation encourages bundling work into one change, which is harder to review and roll back.
How should teams measure the impact of AI coding assistants?
Compare cohorts of AI users against trusted metrics such as pull request cycle time, size, review pushback and change failure rate. Track utilisation, impact and cost together rather than activity alone.
What makes a codebase ready for AI coding agents?
Agents perform best with accurate documentation, modular code, simple data structures and reliable local CI with non-flaky tests. These are the same foundations that help human developers.
A short note about the product, the timeline and who it is for is enough to start. You will hear back from the engineer who would do the work, not a sales team.
Start a partnership