Using LLMs to Evaluate Agents Is Harder Than Building the Bot
Using LLMs to Evaluate Agents Is Harder Than Building the Bot
The fashionable AI conversation in CCaaS right now is about replacing agents with bots. The work I'm actually doing at Cisco is the opposite — using LLMs to evaluate agents, both the human ones and the AI ones. After spending the last year building this, I can tell you the headline lesson up front: building the agent is the easy half. Evaluating it is where the real product work lives.
The QA problem nobody fixed
In a typical contact centre, around 2% of conversations get reviewed. The other 98% are invisible. QA managers know this. Compliance teams know this. Nobody has a credible plan to fix it because reviewing a conversation costs 10–15 minutes of someone's day, and the maths just doesn't close.
Our bet is straightforward: if an LLM can read a conversation transcript, it can evaluate it against a rubric. Not perfectly, but consistently, and at 100% coverage. The target we set ourselves was an 80% reduction in QA time while raising the coverage from 2% to 100%.
The bet is right. The execution is what's hard.
What "evaluation" actually has to do
A real evaluation pipeline isn't one prompt. It's at least four distinct jobs running in sequence:
The hidden product question
The deeper, less obvious lesson is that an LLM-based evaluator is itself a coaching surface. The output isn't really "a score." The output is *a story the manager and the agent will read together on Monday*.
That changes the design entirely. Suddenly the priorities are:
What I'd tell another PM building this
If you're starting a similar journey, three things:

Shivani skipped presentations and built real AI products.
Shivani Kolala was part of the March 2026 cohort at Curious PM, alongside 17 other talented participants.
