Why AI Coding Agents Degrade a Codebase One Feature at a Time
- AI Coding Agents
- Code Quality
- Software Maintainability
- Web Development
A model that fixes one well-scoped bug and passes the tests is not the same thing as a model that can extend a product for a year. Most public scores measure the first. Founders pay for the second.
Early benchmarks hand an agent a small, self-contained issue, then check whether the hidden tests pass. Later ones ask it to rebuild a whole system from a specification. Both give the model the entire problem up front. Real products never work that way. Requirements arrive one at a time, each built on top of the last decision.
What iterative work does to code quality
A more realistic test starts with a greenfield build, then asks for feature after feature in the same codebase, with checkpoints along the way. The pattern researchers report is uncomfortable: as the agent keeps working without knowing the whole problem, failing tests trend upward and the structure gets worse. Abstractions stack up instead of being reshaped. Old code is rarely removed.
This matches what many teams feel in practice. Recent models look much stronger at tasks like computer use and visual demos, yet controlling complexity, deleting code and changing abstractions remain weaker areas. Maintainability is a long-horizon skill, and it is not what most training effort rewards.

Do not borrow against a future model
It is tempting to accept sloppy code today on the theory that a better model will clean it up later. That bet has been made through several model generations already, and the cleanup ability has not arrived on schedule. Nobody can say when it will. Treat the quality of the codebase as something the team owns now, not something to defer.
For a product team, that means a few concrete habits:
- Keep human ownership of structure. Agents add code quickly, but someone must decide where it belongs and what should be removed.
- Review for accumulation, not only correctness. A change can pass its tests and still make the next change harder.
- Run less in parallel. Fewer simultaneous agent tasks leave more attention for judging each result.
Plan to remove failure modes before code exists
Asking an agent for a large feature in one shot usually produces a pull request that needs heavy rework. A better approach is progressive narrowing. Start with a one-page description of the problem and the intended end state, which rules out a large share of the wrong approaches. Then zoom in to the specific changes, which rules out most of what is left. By the time code is written, little rework should remain.
Whether this document is called a plan, a spec or a design doc matters less than its content: where the product is going, how the work gets there, and how success will be judged. The code remains the source of truth, so the plan is a filter and not a promise of perfection.
Model choice deserves the same discipline. The largest and most expensive model is not needed for every task, and some organisations restrict it to cases that justify the cost.
What this means for web development
AI coding agents are genuinely useful for web products, but their weakness lies in sustained change, not first drafts. Studios and teams that pair them with upfront design, small reviewed increments and deliberate refactoring keep a codebase healthy. Teams that merge whatever passes the tests find out months later that every new feature costs more than the last.
Frequently asked questions
Do AI coding agents make codebases worse over time?
They can, when they keep adding features without a full picture of the problem. Iterative evaluations show failing tests trending upward and structure degrading, so human ownership of design and refactoring remains necessary.
Will better AI models fix messy code later?
Nobody can promise that. Teams have deferred cleanup to the next model for several generations without the ability arriving, so code quality should be managed now.
How do you reduce rework on large AI-generated features?
Narrow the problem before any code is written. A one-page description of the problem and end state, followed by a detailed set of changes, removes most wrong approaches early.
Should every task use the most expensive AI model?
No. The largest model is not needed for everything, and some organisations reserve it for cases that justify the cost.
What should code review look for in AI-written changes?
Review for accumulation as well as correctness. A change can pass its tests and still make the next change harder, for example by stacking abstractions or leaving dead code behind.
A short note about the product, the timeline and who it is for is enough to start. You will hear back from the engineer who would do the work, not a sales team.
Start a partnership