Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale explores a new way to test AI customer support agents before they interact with real people. In regulated industries like banking, updating AI agents is risky because errors can erode customer trust. While manual testing is slow and live A/B testing can expose customers to faulty AI, this paper proposes a "simulation-first" approach. By using synthetic customers to interact with AI agents in a controlled environment, developers can catch mistakes and refine agent behavior without any risk to actual users.
How the Simulation Works
The researchers use a tool called Snowglobe to create a "sandbox" for AI agents. Instead of connecting to real backends, the agent interacts with synthetic personas that mimic human customers. These personas are designed to follow specific scenarios, such as requesting a card reissue or checking a delivery status. The simulator acts as an orchestrator, managing the conversation flow and providing mock tool responses. This allows developers to test how an agent handles complex, multi-step tasks—like verifying user data or following internal policies—before the agent is ever deployed to a live customer. The ai agents story also surfaces in OpenAI Unveils GPT-Red an Automated Model..., adding another angle.
Testing and Validation
To ensure the simulator is reliable, the team compared it against real-world data from Nubank’s existing customer support agents. They used several diagnostic methods, including comparing conversation lengths, analyzing the "distance" between simulated and real conversations using text embeddings, and having human experts try to distinguish between real and simulated chat logs. The results showed that the simulator’s performance closely mirrored real-world trends, providing a trustworthy signal for whether an agent version was ready for production.
Real-World Impact
The simulation-guided approach significantly improved the development lifecycle at Nubank. By screening agents through simulation, the team was able to iterate 4.8 times faster than they could without the simulator. In live A/B tests, this method led to measurable business improvements:
The Card Management agent saw a 36.69-point increase in transactional net promoter score (tNPS).
A new model selected via simulation increased the self-service rate (SSR) by 8.82 percentage points, reaching the highest level ever observed at the company.
The selected model also reduced latency by 25%, proving that simulation can help optimize both quality and performance. To see openai in practice, Note-Taking is Dead walks through a concrete example.
Key Takeaways
The study demonstrates that even if a simulator is not perfect, it can still provide "directionally correct" guidance that leads to better outcomes in the real world. By using simulation as a screening layer, companies can explore a wider range of prompts, models, and reasoning settings than would be possible through live testing alone. This workflow allows developers to move faster and with more confidence, ensuring that only the most reliable and effective AI agents reach the customer. The ai agents story also surfaces in Google opens early access to AI..., adding another angle. as detailed in the full paper on Arxiv
Comments (0)
to join the discussion
No comments yet
Be the first to share your thoughts!