Back

Using LLMs to Evaluate Agents Is Harder Than Building the Bot

5 MINS

Using LLMs to Evaluate Agents Is Harder Than Building the Bot

The fashionable AI conversation in CCaaS right now is about replacing agents with bots. The work I'm actually doing at Cisco is the opposite — using LLMs to evaluate agents, both the human ones and the AI ones. After spending the last year building this, I can tell you the headline lesson up front: building the agent is the easy half. Evaluating it is where the real product work lives.

The QA problem nobody fixed

In a typical contact centre, around 2% of conversations get reviewed. The other 98% are invisible. QA managers know this. Compliance teams know this. Nobody has a credible plan to fix it because reviewing a conversation costs 10–15 minutes of someone's day, and the maths just doesn't close.

Our bet is straightforward: if an LLM can read a conversation transcript, it can evaluate it against a rubric. Not perfectly, but consistently, and at 100% coverage. The target we set ourselves was an 80% reduction in QA time while raising the coverage from 2% to 100%.

The bet is right. The execution is what's hard.

What "evaluation" actually has to do

A real evaluation pipeline isn't one prompt. It's at least four distinct jobs running in sequence:

Compliance check — did the agent say the disclosures the regulator requires?
Soft-skills scoring — did the agent acknowledge the customer's emotion, set expectations, close the loop?
Outcome detection — was the customer's actual problem resolved, or did the call just end?
Coaching summary — what's the one thing this agent should do differently next call? Each of those has a different prompt shape, a different failure mode, and a different cost profile. Treating them as one "evaluator" is a recipe for an unreliable product.

The hidden product question

The deeper, less obvious lesson is that an LLM-based evaluator is itself a coaching surface. The output isn't really "a score." The output is *a story the manager and the agent will read together on Monday*.

That changes the design entirely. Suddenly the priorities are:

Will the agent trust this feedback, or feel ambushed?
Will the manager be able to act on it without a 30-minute deep-dive?
Will the language survive translation into 10+ markets without losing nuance? Those are not LLM questions. They're product and design questions wearing an AI costume.

What I'd tell another PM building this

If you're starting a similar journey, three things:

Pick the rubric before the model. A well-defined rubric is 80% of the work and is model-agnostic. The model is swappable. The rubric is your IP.
Ship the manager view before the agent view. Managers tolerate rough edges in the name of saving time. Agents won't. Earn manager trust first, then build the agent-facing layer.
Treat hallucinations as a UX problem. The honest answer is "the model will sometimes get it wrong." Your UI either makes that easy to challenge, or it makes it dangerous. There is no middle option. The fun part of CCaaS used to be optimising routing and dashboards. The fun part now is figuring out what *quality* even means at 100% coverage. That's a far better problem.
Background

Shivani skipped presentations and built real AI products.

Shivani Kolala was part of the March 2026 cohort at Curious PM, alongside 17 other talented participants.