Observing And Testing CX Agents | Interrupt 26

LangChain · Beginner ·🤖 AI Agents & Automation ·1mo ago

Key Takeaways

Tests and observes CX agents using feedback and code integration

Original Description

At Interrupt, the agent conference by LangChain, Carlos Pereira from Cisco's Customer Experience team showed how you close the loop between production feedback and code at 16 million customer interactions per year. The core problem with scaling agents: once adoption climbs, feedback volume outpaces your team. Every thumbs down, every confused user, every low-confidence routing decision is a signal. If you treat it as noise, or if your team becomes the bottleneck processing it, adoption drops. This talk walks through the system Cisco built to take a user's thumbs down all the way to a merged PR, with AI handling triage and diagnostics and humans in the loop only where decisions matter. Observing And Testing CX Agents | Interrupt 26 0:00 Introduction and Day 2 context 0:49 What today covers: observability, testing, and support 1:11 Cisco support at scale: 16M interactions per year 1:34 When network outages hit: why this matters 1:54 CX methodology and team structure 2:16 Evolution: from chatbot to autonomous teammate 2:51 Today's question: closing the production feedback loop 3:23 The approach: continuous feedback loop, not a ticket queue 4:11 Every signal matters: thumbs down, errors, and confusion 4:41 Why humans become the bottleneck at scale 5:57 Signal capture via LangSmith traces 6:12 Triage agent: LangSmith MCP and Jira MCP in practice 6:55 Code agent: clustering and diagnosing issues 7:35 Human in the loop: only on writes, not reads 7:56 Proactive and reactive feedback pipeline 9:20 Treat evals like tests, not experiments 10:56 MCP as the integration layer 11:21 Human oversight only on writes 12:47 Lessons learned 13:53 Observability is your new bottleneck 14:05 Close the feedback loop with agents 14:36 Evals are infrastructure, not a side project 15:34 Support use case: Cisco technical support 15:48 Live example: enterprise network assessment 17:04 2,176 security findings: where do you start? 18:13 Semantic routing for ambiguous prompts 19:33 Parallel pipel
Watch on YouTube ↗ (saves to browser)
Sign in to unlock AI tutor explanation · ⚡30

Related Reads

📰
Evaluating AI Agents: A production blueprint with Strands and AgentCore
Learn to build an end-to-end evaluation pipeline for AI agents using Strands and AgentCore, reducing incorrect results and issue detection time
AWS Machine Learning
📰
AMD's $5B Anthropic Bet Signals a Shift in AI Chip Wars
AMD invests $5B in Anthropic, challenging NVIDIA's dominance in AI chip market, and shifting focus from building best chips to controlling access to them
Dev.to AI
📰
Fine-Tuning Is Often the Wrong First Move: Introducing Harneloop
Learn why fine-tuning an AI model might not be the best first approach and how Harneloop can help, and discover an alternative method to improve AI agent performance
Dev.to AI
📰
Detecting silent agent failures with Amazon Bedrock AgentCore optimization
Learn to detect silent agent failures in production AI agents using Amazon Bedrock AgentCore optimization and fix high-impact issues first
AWS Machine Learning

Chapters (27)

Introduction and Day 2 context
0:49 What today covers: observability, testing, and support
1:11 Cisco support at scale: 16M interactions per year
1:34 When network outages hit: why this matters
1:54 CX methodology and team structure
2:16 Evolution: from chatbot to autonomous teammate
2:51 Today's question: closing the production feedback loop
3:23 The approach: continuous feedback loop, not a ticket queue
4:11 Every signal matters: thumbs down, errors, and confusion
4:41 Why humans become the bottleneck at scale
5:57 Signal capture via LangSmith traces
6:12 Triage agent: LangSmith MCP and Jira MCP in practice
6:55 Code agent: clustering and diagnosing issues
7:35 Human in the loop: only on writes, not reads
7:56 Proactive and reactive feedback pipeline
9:20 Treat evals like tests, not experiments
10:56 MCP as the integration layer
11:21 Human oversight only on writes
12:47 Lessons learned
13:53 Observability is your new bottleneck
14:05 Close the feedback loop with agents
14:36 Evals are infrastructure, not a side project
15:34 Support use case: Cisco technical support
15:48 Live example: enterprise network assessment
17:04 2,176 security findings: where do you start?
18:13 Semantic routing for ambiguous prompts
19:33 Parallel pipel
Up next
Agentic Workflow Series Ep. 1 | Build Your First AI Agent Workflow from Scratch
Pavithra’s Podcast
Watch →