Calibration
Calibration answers one question: when a person looked at a call, did they agree with the QA Bot?
What it compares
Section titled “What it compares”For every review a person confirmed or changed in the period, it keeps the bot’s original answer to each question next to the person’s current answer.
- Reviews that nobody looked at are left out.
- A review that was only disputed, with no confirm or change, is left out.
- A question where both the bot and the person said Not applicable is not counted.
Pick the period at the top: 7, 30 or 90 days. If nobody looked at a call in the period, the page says “No reviewer has looked at a call in this period”.
The totals
Section titled “The totals”| Figure | Meaning |
|---|---|
| Reviews compared | Reviews a person confirmed or changed |
| Confirmed as scored | The person changed no answer |
| Changed by a person | At least one answer was changed |
The table by question
Section titled “The table by question”| Column | Meaning |
|---|---|
| Compared | How many reviews had that question applying |
| Agreement | Share of those where bot and person gave the same answer |
| Bot passed, person failed | The bot was too lenient |
| Bot failed, person passed | The bot was too strict |
How to use it
Section titled “How to use it”- Wait until a question has a fair number of comparisons. A rate from 3 reviews tells you little.
- Look for the lowest agreement first.
- Check which way the mistakes go. Many “Bot passed, person failed” means the bot misses real problems. Many “Bot failed, person passed” means it flags things that are fine.
- Fix the question, not the people. Make “What good looks like” more exact, add a “When does it not apply” line, or change a judgement item into a rule. See scorecards.
- Come back after the new version has scored a few weeks of calls.
High agreement on a question is a reason to trust the bot on that question. It is not proof. Reviewers mostly open the calls that already needed a look, so the sample leans toward unusual calls.
Scoring again
Section titled “Scoring again”Score again only replaces a review nobody has worked on. Once a person confirms, changes or disputes it, the QA Bot refuses to overwrite it. This keeps the comparison honest.