Why AI Coding Agents Erode Code Maintainability Over Time
- AI Coding Agents
- Code Maintainability
- Code Review
- Software Architecture
An AI coding agent can ship a feature in minutes. Six months later, the same codebase is often harder to change than it was at the start. A small edit in one module breaks three others, and nobody, human or machine, can quickly explain why. This is not simply a prompting failure. It follows from how coding models are taught to behave.
Why passing tests is not the same as good design
Coding models improve through reinforcement learning on tasks with a binary outcome: did the requested fix work, and did the existing tests still pass? Attempts that succeed are reinforced, and attempts that fail are discouraged. That signal is precise about correctness. It says nothing about program design. A model that wraps a risky call in an unnecessary try-catch, or casts a value to another type just to silence an error, earns the same reward as one that restructures the code cleanly.
Maintainability is hard to reward because its cost arrives late. Poor architecture shows itself over months or years, far too late to feed back into a training episode. Newer evaluation approaches try to close the gap with larger multi-step tasks and judge models that check code quality rules. Even so, a model that could reliably recognise good code would probably write it in the first place. Until that improves, agents are strong at one-off changes and weaker at keeping a codebase coherent as it grows. The typical symptom is shotgun surgery, where one logical change forces edits scattered across many files.
Why a fully unattended pipeline is risky
The appealing idea is a pipeline where agents build, other agents review, and nobody reads the code. Automated review can raise the floor, but it inherits the same blind spot as the builder. For a prototype or a throwaway side project, that trade-off may be acceptable. For a product that must be maintained for years, the constraints are different. A team that stops reading its code will eventually meet a failure that the agents cannot diagnose, inside code that nobody on the team recognises.
Plan before you generate
The practical answer is to move effort earlier, where it is cheap. A planning sequence that works well for larger changes:
- Product review: the problem being solved, the intended behaviour, and mock-ups where they help.
- Architecture: component contracts, data models and constraints.
- Program design: types, method signatures, file layout and call paths.
- Vertical slices: an implementation order, with a check at the end of each slice.
Small tasks can still go straight to the agent. For bigger work, agreeing the design first means the resulting pull request matches what the team expected, so review becomes a confirmation rather than an investigation. If reviewers feel they are drowning in pull requests, the cause is usually too many poor ones rather than too many in total. A good pull request is quick and even pleasant to read.
What this means for web projects
For web products with a long life, the discipline stays the same as before AI arrived: people own the design, agents do much of the typing, and someone reads every line. That is the model we follow at Curiosive, where agents speed up delivery while senior engineers own the architecture and review the code.
The time saved by generation should be reinvested in alignment up front. Half an hour spent on types, contracts and slices is usually cheaper than hours spent untangling a pull request that ignores them. Teams that do this keep the speed of AI-assisted development without inheriting a codebase that gets harder to change every month.
Frequently asked questions
Why does AI-generated code get harder to maintain over time?
Coding models are trained mainly on whether tests pass, not on whether the design is clean. Nothing in that signal penalises needless try-catch blocks or type casts, so small design flaws accumulate and changes start to ripple across many files.
Can automated AI code review replace human review?
No, automated review can raise the quality floor but it shares the blind spots of the agent that wrote the code. For long-lived products, people still need to read the code and own the design.
How do you plan work for an AI coding agent?
Start with a product review, then agree architecture, program design and vertical slices before generating code. Small tasks can go straight to the agent, while larger ones benefit from that up-front alignment.
What is shotgun surgery in a codebase?
Shotgun surgery is a code smell where one logical change forces edits in many scattered places. It is a common symptom of AI-generated code that has eroded maintainability.
A short note about the product, the timeline and who it is for is enough to start. You will hear back from the engineer who would do the work, not a sales team.
Start a partnership