Skip to content

Calibration

Calibration answers one question: when a person looked at a call, did they agree with the QA Bot?

For every review a person confirmed or changed in the period, it keeps the bot’s original answer to each question next to the person’s current answer.

  • Reviews that nobody looked at are left out.
  • A review that was only disputed, with no confirm or change, is left out.
  • A question where both the bot and the person said Not applicable is not counted.

Pick the period at the top: 7, 30 or 90 days. If nobody looked at a call in the period, the page says “No reviewer has looked at a call in this period”.

Figure Meaning
Reviews compared Reviews a person confirmed or changed
Confirmed as scored The person changed no answer
Changed by a person At least one answer was changed
Column Meaning
Compared How many reviews had that question applying
Agreement Share of those where bot and person gave the same answer
Bot passed, person failed The bot was too lenient
Bot failed, person passed The bot was too strict
  1. Wait until a question has a fair number of comparisons. A rate from 3 reviews tells you little.
  2. Look for the lowest agreement first.
  3. Check which way the mistakes go. Many “Bot passed, person failed” means the bot misses real problems. Many “Bot failed, person passed” means it flags things that are fine.
  4. Fix the question, not the people. Make “What good looks like” more exact, add a “When does it not apply” line, or change a judgement item into a rule. See scorecards.
  5. Come back after the new version has scored a few weeks of calls.

High agreement on a question is a reason to trust the bot on that question. It is not proof. Reviewers mostly open the calls that already needed a look, so the sample leans toward unusual calls.

Score again only replaces a review nobody has worked on. Once a person confirms, changes or disputes it, the QA Bot refuses to overwrite it. This keeps the comparison honest.