All writing

Quality Gates for AI Coding Agents: Rules Are Not Enough

Ugur Kellecioglu3 min read
  • AI Coding Agents
  • Quality Gates
  • Code Review
  • Web Development

Most teams start by writing rules for their AI coding agents: use this pattern, run the tests, match the design. It helps for a while. Then a long task drifts, the agent forgets a rule halfway through, and unfinished work is reported as done. A rule is only advice. A gate is something the work has to pass before anything else can happen.

Why long agent sessions need checkpoints

When one agent holds an entire feature in a single context window, quality tends to fall as the session grows. Early details get crowded out, and the agent starts forgetting things that mattered. The fix is to split a large feature into small checkpoints and give each one a fresh context, usually by delegating it to a separate sub-agent.

Order matters. Checkpoints work best when they run in increasing complexity, with the smallest piece first. Each one builds on the last, so a wrong early decision is caught while it is still small and cheap to fix.

The plan itself deserves a human read before any code is written. Structured files are easy for agents to parse but tiring for people, so it helps to render the plan as a simple page that shows each step, its status and what 'done' means for it. Reviewing that plan thoroughly is the cheapest quality control in the whole process, because it stops agent time being spent on a feature that was misunderstood.

Gates that are enforced, not requested

A gate is different from an instruction because the agent cannot skip it. One practical way to enforce this is a hook that runs whenever the agent tries to stop. If the gates for the current checkpoint are not yet passed, the hook tells the agent to keep working. The agent does not get to declare victory.

A sensible set of gates looks like this:

  • Behaviour gate. Tests are planned before the code, written so they fail first, then the code is written until they pass. Anything that can be verified without a browser, such as a permission rule, should be covered here.
  • Interface gate. The built screen is compared with an approved prototype and a written design file. Separate reviewers can judge appearance and behaviour, and it helps if they do not see the project's own instructions, so they judge only against the prototype.
  • Adversarial review gate. One agent assumes the code is wrong and looks for problems. A second agent fixes what is found. They repeat until the critic approves.
  • Human gate. A person uses the finished feature and decides whether it works the way they wanted.
Four enforced quality gates for AI coding agents: behaviour, interface, adversarial review, and human acceptance.

Turning feedback into the next checkpoint

Human feedback should not be a loose list of requests. Each change becomes a new checkpoint that passes through the same gates, and the feedback is also recorded in a learnings file that every agent reads before starting work. Over time the process absorbs the team's preferences instead of relearning them.

What this means for web development practice

For a web product, this structure maps cleanly onto work we already value: small reviewable changes, tests written ahead of code, design fidelity checked against an agreed reference, and senior review of the final result. The agent may be wrong on any attempt. What matters is that it is not allowed to ship until it is right.

If you are considering AI-assisted delivery for a product, ask how the process stops unfinished work from passing, not only how fast the agent writes code. Enforced checkpoints and gates are what turn a fast tool into a dependable one.

Frequently asked questions

What is a quality gate for an AI coding agent?

A quality gate is a check the work must pass before the agent can move on to the next task. Unlike a written rule, it is enforced, for example by a hook that tells the agent to keep working until the gate passes.

Why split AI coding work into checkpoints?

Splitting work into checkpoints keeps each task in a fresh context and catches wrong decisions early. Long single sessions tend to lose detail, so quality drops as the context grows.

How do you review UI built by an AI agent?

Compare the built screen against an approved prototype and a written design file. Reviewers that judge appearance and behaviour separately, without seeing the project instructions, give a more honest comparison.

What is an adversarial review for AI-written code?

An adversarial review uses one agent that assumes the code is wrong and hunts for problems, plus a second agent that fixes them. The loop repeats until the critic approves the code.

Do AI coding workflows still need a human reviewer?

Yes, a person should use the finished feature and judge whether it works as intended. Their feedback becomes new checkpoints that pass through the same gates.

A short note about the product, the timeline and who it is for is enough to start. You will hear back from the engineer who would do the work, not a sales team.

Start a partnership