A field report from production agentic systems handling 50,000+ daily customer interactions. The five decisions that separate slideware demos from systems that survive a regulator and a Black Friday.
Dr. Atif Farid Mohammad & Adnan Tahir AI Advisor (PhD, Capitol Tech) · Chief Technology Officer
Every CTO we talk to has the same story. They piloted an AI agent in Q3. It demoed beautifully. It made it to a steering committee deck. Then it died — quietly, in the staging environment, sometime around the third or fourth integration meeting. The cited reasons vary: “compliance concerns,” “model drift,” “integration complexity.” The actual reason is almost always architectural. The agent was built for the demo, not for production.
We’ve shipped agents into FinTech super-apps handling 50,000+ customer interactions a day, into healthcare platforms touching patient records, and into government HR portals where every transaction is auditable. We’ve also watched dozens of pilots — ours and others’ — fail at the deployment line. Below are the five architecture decisions that, in our experience, separate the agents that ship from the ones that don’t.
TL;DR — Production agents fail because of (1) wrong retrieval boundary, (2) no escalation path, (3) no eval harness, (4) shared session state, and (5) no observability. Fix those five, and most other issues become tractable.
The single most common failure pattern: developers point a RAG (Retrieval Augmented Generation) pipeline at “all the company data” and expect intelligence to emerge. It does not. What emerges is a slow, hallucinatory agent that confidently cites the wrong document.
Production-grade retrieval requires a tight, intentional boundary. For a knowledge agent over policy documents, that means scoping retrieval to a curated, versioned corpus — not the entire SharePoint. It means tagging documents with effective dates, jurisdictions, and revocation status. It means having a human in the loop to mark “do not retrieve” on documents that are deprecated but not yet deleted.
In our knowledge agent deployment over 5,283 policy documents, the single biggest accuracy lift came not from a better embedding model, but from cleaning up the corpus and introducing a “freshness score” that biases retrieval toward documents updated in the last 18 months. Phase-1 accuracy went from 67% to 85%.
Demos always show the happy path. Production is 80% edge cases. An agent that handles 70% of conversations beautifully and silently fails the other 30% is worse than no agent at all — because the 30% are the conversations that matter most.
Every production agent we’ve shipped has a triage layer in front of it. Before the LLM ever generates a response, a lightweight classifier (often a fine-tuned smaller model or even a rules engine) routes the conversation: tier-1 self-service, tier-2 LLM-handled, tier-3 human escalation. The LLM only handles tier-2. Everything else routes around it.
The hardest part of agentic AI isn’t getting the agent to answer. It’s knowing when not to.
For our service agent that replaced 180 live support agents, the triage layer routes roughly 15% of conversations to a human within seconds — typically high-value disputes, vulnerable-customer flags, or anything involving threats of legal action. That 15% is sacred. The CSAT for those escalated conversations sits at 4.8/5; if the agent had tried to handle them itself, it would have wrecked the brand.
Most agent projects ship without an evaluation framework. The team’s confidence comes from “it seemed to work in testing” — which is not a metric. Then production traffic hits and behavior drifts and nobody can prove whether it’s getting better or worse.
An eval harness for agents has three layers:
Without these, you’re flying blind. With them, you can ship aggressively because you can detect regressions within hours instead of weeks.
This is the architectural bug that’s hardest to spot in a demo and hardest to debug in production. When agents share context across users — through naive caching, shared embeddings, or sloppy session management — you get cross-contamination. User A’s PII shows up in User B’s response. Conversations bleed. The breach is small, silent, and catastrophic.
Production agents need session isolation by default. Each conversation gets its own context window. Tool calls are scoped to the authenticated user. Caching is keyed by user identity and request signature, not just request signature. RAG retrieval respects row-level access controls in the underlying data store — not just on read, but on indexing.
This is one of the few areas where boring engineering pays massive dividends. If you’ve been an enterprise software engineer for a decade, you know how to design for tenancy isolation. Apply that discipline to your agent stack and you’ll skip an entire class of incidents that the LangChain tutorials don’t warn you about.
An agent in production without observability is a system you cannot operate. You cannot diagnose latency spikes. You cannot trace a hallucination back to the source document. You cannot do FinOps. You cannot prove compliance.
The minimum observability stack for a production agent:
Pull all five together and you get an architecture that looks like this:
User → Triage Classifier → [tier-1: self-serve | tier-2: agent | tier-3: human]
↓
Session Isolator (per-user context)
↓
Scoped Retrieval (boundary + freshness)
↓
LLM + Tools
↓
Response + Citations + Telemetry
↓
Eval Harness (golden + LLM-judge + human sample)
It is not glamorous. It will not win an Awwwards demo. It is the difference between an agent that you can ship to a regulated FinTech and an agent that lives forever in staging.
If we were starting fresh today, we would:
The bigger point — Agentic AI in 2026 is not constrained by model capability. The frontier models are good enough for most use cases. Agentic AI is constrained by operational maturity — eval, observability, escalation, governance. The teams that win are the ones treating agents like the production systems they are, not like demos.
If you’re piloting an agent right now and any of these five problems sound familiar, the fix is mostly architectural — not modelistic. We’ve open-sourced our reference architecture diagrams in our AI services overview, and we’d be happy to walk through your specific situation. Book a 30-minute architecture review with our CTO and AI advisor — no sales loop, just a real conversation.
AUTHOR + RELATED
Dr. Atif Farid Mohammad AI & Quantum Advisor
Dual PhD — Cyber Security & Scientific Computing. Quantum Computing Chair at Capitol Technology University. Adjunct Professor of AI/ML.
Adnan Tahir Chief Technology Officer
Technology architect overseeing all engineering decisions. Owns development of PACT ERP, QuickHCM, and AppsGenii’s agentic platform.
9 min read
HEALTHTECH · COMPLIANCE
10 min read