Production AI architecture: routing, cost controls and human review
- AI architecture
- Model routing
- LLM evaluations

AI needs architecture. Every model needs a system. For a product team, that system includes the rules for choosing a model, limiting its work, checking its output and deciding who can approve the next action. A convincing demo proves that a model can produce something useful. Production AI architecture determines whether that usefulness survives repeated requests, changing providers and real business constraints.
This guide is for founders and engineering teams moving an AI feature or coding workflow into daily use. It draws on a recent Wagtail engineering experiment, then sets out our recommendations for model routing, AI cost controls, reliability, evaluations and human review. The recommendations are design guidance; they are not results measured in that experiment.
What Wagtail’s GLM 5.3 Flash experiment actually showed
In One month on GLM 5.3 Flash, published on 2 October 2026, Wagtail core team member Thibaud Colas describes trying to use one efficient open model throughout September. He reports meeting that goal for the first half of the month, then using other models because of infrastructure availability and continued experimentation. This is one team member’s experience in a particular engineering workflow, not evidence that one model is best for every product.
The Hacker News discussion adds a useful clarification: Colas says a prototype session used the non-Flash model when he intended to use Flash. The distinction matters. A model name that looks nearly identical can represent a different spending decision. The discussion also contains differing opinions about planning, implementation and review; those comments are anecdotes, not comparative performance data.
We have deliberately avoided turning the reported token counts, spending or energy estimates into product forecasts. Colas clarifies in the discussion that the energy figures cover GPU energy use. They are not a complete lifecycle assessment. The engineering takeaway is narrower and more actionable: make model selection explicit, expect provider failures and measure useful outcomes alongside usage.
Start with the workflow and its failure modes
A production LLM application needs a defined input, an allowed set of actions and a clear stopping condition. Start by writing down what happens when the output is incomplete, wrong, late or unavailable. A summarizer may safely return a draft for review. An agent that edits customer records needs stronger permissions and an approval boundary before it writes anything.
For an illustrative support assistant, separate retrieval, drafting and sending. Retrieval can expose only the documents the current user may read. Drafting can produce a structured proposed answer with supporting references. Sending can remain a deliberate human action. This separation lets a team inspect errors without giving a text generator authority over the whole workflow.
Define success in terms the product can observe: an answer backed by the retrieved policy, a correctly extracted field or a patch that passes the repository’s checks. Fluent prose alone is a weak acceptance criterion. It can hide a missing constraint just as easily as it can communicate a correct answer.
Model routing: match capability to the task
Model routing is a policy that assigns a model and provider to a task under constraints such as quality, latency, data handling and budget. Begin with a small set of task classes. Straightforward extraction and bounded edits can be candidates for an efficient default model. Ambiguous planning or difficult debugging can be candidates for escalation, once evaluations show that escalation helps.
Do not route only on a model’s self-reported confidence. Prefer observable signals: invalid output, missing evidence, failed checks or a task class that requires a different capability. If the cheap path fails, allow a bounded second attempt or a stronger model within a remaining budget. If neither succeeds, return an explicit failure or request human review.
Use exact model identifiers in configuration and logs, including provider and revision where available. Treat a routing-policy change as a release: evaluate it, record it and make it reversible. A friendly label such as “fast” is useful in a user interface, but the system should retain the actual decision behind that label.
Adding a planner, implementer and reviewer can help when their responsibilities are distinct. It also adds calls, coordination and opportunities for repeated work. Start with the simplest workflow that passes your quality checks. Add an agent role only when it solves a demonstrated problem, and give every role a bounded goal.
AI cost controls: budget the whole job
The cost of an AI workflow includes input and output, repeated context, retries, tool calls, infrastructure and human correction. Token price is only one component. Measure cost per accepted result alongside latency and failure rate; a cheaper request that repeatedly produces unusable output can become the more expensive workflow.
Set limits before a job starts. Use a maximum input size, output allowance, tool-call count, retry count, wall-clock deadline and monetary budget. Reserve budget across concurrent requests so several agents cannot each spend the same remaining allowance. Enforce limits in application code or a gateway; a prompt asking a model to be economical is not a spending control.
An illustrative policy might allow one initial attempt and one targeted repair, then escalate to a person when validation still fails. The exact limits depend on your workload. Keep prototype and evaluation spending in a separate budget from customer-facing traffic, so a useful experiment cannot quietly consume the operating allowance.
Cache only where privacy, freshness and correctness permit it. Reusing a public reference summary is different from reusing a customer-specific answer. Track cached responses separately, and invalidate them when the underlying information or permissions change. Reducing context should remove irrelevant material while preserving the evidence needed for a correct decision.
Reliability: design a useful failure path
Provider availability and model quality are separate concerns. A task can fail because the endpoint times out, because capacity is constrained or because the answer violates the application’s contract. Give each failure an explicit response instead of treating all of them as a reason to call the model again.
Use request deadlines, bounded retries with backoff and a circuit breaker when an endpoint repeatedly fails. Choose a fallback provider only after checking its data-handling requirements and evaluating its output on the same tasks. Another endpoint may have different tool support, context limits or output behavior even when its model family sounds familiar.
Preserve a stable internal request and response contract through provider adapters. Validate structured output before it reaches business logic. Treat timeouts around external writes carefully: a retry must not create a second invoice or send a second message. Use idempotency keys and durable job state for side effects, with approval recorded separately from generation.
The failure path should still help the user. Keep their input, explain that the task could not finish and offer a manual next step or a queued retry where appropriate. For a critical workflow, a partial draft that cannot be trusted should remain visibly unapproved.
LLM evaluations: test the work you actually need
Build a representative evaluation set from the intended workflow before selecting a default model. Include routine cases, ambiguous requests, missing information, long inputs and adversarial content. Remove or replace sensitive data. Reserve a holdout set so tuning prompts against known examples does not become your only evidence of quality.
Use deterministic checks when possible: schema validation, unit tests, exact field comparisons, permission checks and verification that cited material exists. Use a documented human rubric for properties such as usefulness, completeness and whether an answer overstates the evidence. A model-based judge can assist triage, but its decisions also need calibration and spot checks.
Compare candidate models on the same inputs and record accepted-result rate, latency distribution and total cost including repairs. Report the evaluation conditions and sample limitations. Do not turn a handful of successful demos into a universal performance claim.
Repeat evaluations when prompts, retrieval, tools, routing or model versions change. Release gradually, keep the previous configuration available and monitor production failures against the evaluation set. When a new failure appears, add a sanitized example so the next release must account for it.
Human review: put approval where consequences begin
Human review works best when a reviewer can see the proposed action, its evidence and what will change. A coding assistant should present a diff and check results. A content assistant should expose source links and unsupported claims. A workflow agent should show the exact recipient or record before a consequential write.
Set review requirements according to impact and reversibility. Low-impact, well-tested outputs may be eligible for automation after evaluation. Access changes, financial actions and difficult-to-reverse writes need explicit authorization and a review step suited to the task. A second model’s agreement does not establish permission to act.
Keep tool credentials scoped to the operation. Retrieved documents and web pages are inputs, not instructions that may redefine tool permissions. Enforce the boundary outside the model, and record who approved an action. This makes accountability a property of the system rather than a hope about the prompt.
A production AI architecture in one view
The infographic below shows a bounded workflow: define the task, route the model, execute within limits, validate the result and obtain approval when the action requires it. Usage and outcome measurements feed the next evaluation and release decision. Its sequence is a reference design, not a claim that every product needs multiple agents.
The same sequence can be implemented with one model and a few ordinary application functions. Complexity should follow the product’s needs. A larger orchestration layer is worthwhile only if it improves a measured bottleneck or makes an important boundary easier to enforce.
A practical path from prototype to production
- Choose one narrow workflow and document its acceptance criteria, allowed tools and manual fallback.
- Assemble a small but representative evaluation set and compare the simplest viable model configurations.
- Put spending limits, deadlines, permissions and output validation around every run.
- Add escalation and provider fallback only after testing their quality, data handling and failure behavior.
- Release gradually, inspect accepted outcomes and corrections, then expand the scope deliberately.
For founders, the decision is about a sustainable product behavior, not winning a model comparison. For engineering teams, it is about keeping that behavior observable and changeable as models evolve. You can replace a model more confidently when the surrounding contract, evaluations and approval boundaries are already clear.
Build the system around your AI feature
Have an AI prototype that now needs to work inside a real product? Explore Curiosive’s engineering services and our approach to product development, or tell us about your workflow. Bring the user problem, the actions the system should take and the constraints it must respect. Those are useful starting points for a conversation about scope and architecture.
Sources and scope
This article is Curiosive’s original analysis, informed by Thibaud Colas’s Wagtail account and the accompanying Hacker News discussion, read on 2 October 2026. The source section reports that experience; the architecture sections present our recommendations and explicitly illustrative examples. No performance benchmarks or customer outcomes are claimed here.
Frequently asked questions
What is production AI architecture?
Production AI architecture is the application design around a model: task boundaries, retrieval, routing, tool permissions, spending limits, validation, monitoring and approval. It makes an AI workflow observable and controllable in everyday use.
How should a team choose between efficient and stronger models?
Evaluate candidate models on representative tasks, then route by demonstrated capability, cost and latency. Use an efficient default where it passes acceptance checks, with bounded escalation for failures or more demanding task classes.
How can an AI agent avoid uncontrolled costs?
Enforce per-job spending limits, input and output allowances, deadlines, retry counts and tool-call limits outside the prompt. Reserve budget across concurrent jobs and measure total cost per accepted result, including repairs.
Does a second AI reviewer replace human approval?
A second model can help identify issues, but agreement between models does not authorize an action. Require human approval at consequential boundaries and show the proposed change, evidence and relevant check results.
A short note about the product, the timeline and who it is for is enough to start. You will hear back from the engineer who would do the work, not a sales team.
Start a partnership