Back to AI Research

AI Research

Screen Before You Serve: Simulation for Production... | AI Research

Key Takeaways

  • Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale explores a new way to test AI customer support agents before t...
  • Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products.
  • Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably.
  • Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust.
  • We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment.
Paper AbstractExpand

Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale explores a new way to test AI customer support agents before they interact with real people. In regulated industries like banking, updating AI agents is risky because errors can erode customer trust. While manual testing is slow and live A/B testing can expose customers to faulty AI, this paper proposes a "simulation-first" approach. By using synthetic customers to interact with AI agents in a controlled environment, developers can catch mistakes and refine agent behavior without any risk to actual users.

How the Simulation Works

The researchers use a tool called Snowglobe to create a "sandbox" for AI agents. Instead of connecting to real backends, the agent interacts with synthetic personas that mimic human customers. These personas are designed to follow specific scenarios, such as requesting a card reissue or checking a delivery status. The simulator acts as an orchestrator, managing the conversation flow and providing mock tool responses. This allows developers to test how an agent handles complex, multi-step tasks—like verifying user data or following internal policies—before the agent is ever deployed to a live customer. The ai agents story also surfaces in OpenAI Unveils GPT-Red an Automated Model..., adding another angle.

Testing and Validation

To ensure the simulator is reliable, the team compared it against real-world data from Nubank’s existing customer support agents. They used several diagnostic methods, including comparing conversation lengths, analyzing the "distance" between simulated and real conversations using text embeddings, and having human experts try to distinguish between real and simulated chat logs. The results showed that the simulator’s performance closely mirrored real-world trends, providing a trustworthy signal for whether an agent version was ready for production.

Real-World Impact

The simulation-guided approach significantly improved the development lifecycle at Nubank. By screening agents through simulation, the team was able to iterate 4.8 times faster than they could without the simulator. In live A/B tests, this method led to measurable business improvements:

  • The Card Management agent saw a 36.69-point increase in transactional net promoter score (tNPS).

  • A new model selected via simulation increased the self-service rate (SSR) by 8.82 percentage points, reaching the highest level ever observed at the company.

  • The selected model also reduced latency by 25%, proving that simulation can help optimize both quality and performance. To see openai in practice, Note-Taking is Dead walks through a concrete example.

Key Takeaways

The study demonstrates that even if a simulator is not perfect, it can still provide "directionally correct" guidance that leads to better outcomes in the real world. By using simulation as a screening layer, companies can explore a wider range of prompts, models, and reasoning settings than would be possible through live testing alone. This workflow allows developers to move faster and with more confidence, ensuring that only the most reliable and effective AI agents reach the customer. The ai agents story also surfaces in Google opens early access to AI..., adding another angle. as detailed in the full paper on Arxiv

Comments (0)

No comments yet

Be the first to share your thoughts!