Artificial intelligence demos can look impressive in a controlled setting. Problems often appear after the technology meets real users, changing data and sustained traffic.
A model that performs well during testing may lose track of a conversation, retrieve irrelevant information or become slower as demand grows. The instinct is often to replace it with a larger model, even when the underlying model is not the source of the problem.
Praveen Asthagiri, a senior technical program manager and senior member of the IEEE, argues that production AI should be treated as a complete system rather than a standalone model.
That idea is central to his book, AI Systems at Scale: Building GenAI, Conversational AI, and Enterprise Platforms That Drive Real Impact.
The book examines how data pipelines, retrieval systems, testing, governance and product goals determine whether an AI system works reliably outside a demonstration.
“The model is rarely the hard part anymore,” Asthagiri said. “The hard part is building everything around it so the model can do its job reliably, on real data, under real load.”
Moving beyond model-first development
Many AI projects begin by selecting a model and measuring its performance against a benchmark. That approach can show whether a model is capable of completing a task, but it does not show whether the full product will perform consistently.
A production system must prepare and retrieve data, manage permissions, measure output quality and recover when something fails. It also needs people who are responsible for each of those functions.
A capable model placed inside a weak system will inherit the system’s weaknesses. Inconsistent data can produce inconsistent answers. Slow retrieval can make the entire product feel unresponsive. Poor monitoring can allow quality to decline without anyone noticing.
Asthagiri calls for a system-first approach in which the model is treated as one component of a larger product.
“A demo only has to work once,” he said. “A product has to work for everyone, every time, and recover gracefully when it does not.”
Memory is an engineering problem
Conversational AI illustrates the difference between a model and a complete system.
A chatbot may answer a single question correctly but struggle to maintain a coherent discussion. It may forget a preference mentioned earlier, repeat a question or bring irrelevant information into a response.
Longer conversations require the system to decide what information to retain, what to discard and how to retrieve the right context without slowing the response.
The model alone does not make those decisions. They depend on the memory and retrieval architecture surrounding it.
Asthagiri separates conversational memory into several categories. Short-term context covers the current exchange. Long-term memory may include stable preferences, while episodic memory allows the system to retrieve relevant details from previous conversations.
Keeping everything is not a practical solution. Large conversation histories increase processing costs and can introduce irrelevant information. Systems need rules that determine which details are useful and how long they should remain available.
“Customers do not experience your model,” Asthagiri said. “They experience whether the system remembered what they said, answered quickly and got it right.”
Results require careful attribution
Asthagiri draws on his experience developing memory and personalization systems for a widely used conversational AI platform.
He reports that one interaction-history project reduced context-loss errors by more than 95%. He also describes a personalization program that achieved a rate above 30% during a four-week measurement period.
Those figures illustrate the potential effect of improving memory architecture, although they are based on Asthagiri’s account and are not presented as independently audited company results.
The broader lesson does not depend on a single metric. Measuring an AI product requires more than tracking whether it produced an acceptable answer during a test.
Teams also need to monitor latency, retrieval quality, consistency, user satisfaction, cost and the system’s ability to recover from errors.
A model can improve on one measure while the overall experience becomes worse. Adding more context, for example, may improve continuity while increasing response time and computing costs.
Governance becomes part of the architecture
Memory also creates privacy and governance questions.
A system capable of remembering a user’s preferences must determine whether it should store that information, how long it should retain it and which other systems may access it.
Incorrect information can also spread. If one component records a preference inaccurately and shares it with other agents, the error may influence several future responses.
Asthagiri argues that teams should establish boundaries around how memory moves between systems. Users should also be able to understand what information is stored and correct or remove it when necessary.
These controls become more important as companies deploy AI agents that can take actions instead of simply answering questions.
An inaccurate chatbot response may frustrate a user. An autonomous agent working with inaccurate data could complete the wrong transaction, change a business process or expose protected information.
A larger model cannot fix every problem
A systems-first approach does not eliminate limitations within AI models. Companies still need to evaluate accuracy, security, bias and whether a model is appropriate for the intended task.
It does, however, discourage teams from treating model selection as the entire AI strategy.
Reliable production systems require clear ownership, realistic testing and continuous measurement. They also need an architecture that can handle changing traffic and data without quietly degrading.
The book’s central argument is that AI success depends on the engineering surrounding the model.
“Everyone is racing to build agents,” Asthagiri said. “The teams that win will be the ones who built the system underneath them first. The model was never the whole story. The system is.”
©2026 Cox Media Group








