How I run an AI quality review
The same five moves every time: read the transcripts before the rubric, write the rubric down, score a sample twice, keep the failures, hand it over so the team can run it without me.
Every quality review I run has the same shape, whether the thing being reviewed is a support bot, an agent that files tickets, or a model writing in a brand's voice.
Read before you grade. The first day is transcripts, fifty or a hundred of them, with no scoring sheet open. The point is to see what the system does when nobody is watching it, and to notice the failures the team has stopped seeing because they see them every day.
Write the criteria down. A rubric is a list of things a good answer must do and must never do, each one specific enough that two people would score the same transcript the same way. "Helpful" is too loose to score. "Answers the question asked before offering anything else" can be scored.
Score a sample twice. Once by me, once by someone on the team, without comparing. Where we disagree, the rubric is unclear, and the rubric gets fixed before anything else does.
Keep the failures. Every transcript that fails becomes a test case with a note on why. That set is worth more than the report, because it catches the same problem when it comes back in the next release.
Hand it over. The deliverable is a rubric and a failure set the team can run without me in the room. If the process only works while I am there, it is not finished.
If your problem looks like one of the four on the home page, book the call. If it does not, you will hear that in the first reply, along with who to talk to instead.