Skip to content

How the QA Bot scores

The QA Bot reads the transcript of a finished call and answers each question on a scorecard with Passed, Failed or Not applicable. It adds up the weights into a score out of 100. Nobody has to listen to every call.

Every question on the scorecard is decided one of two ways. The review screen labels each one “Rule check” or “Judgement check”.

Rule check Judgement check
Decided by A fixed, exact test on the transcript An AI judge that reads the call
Same call, same answer Always Not guaranteed
Evidence The line it found, quoted A line number and a quote

The starter scorecard uses rule checks for the recording notice, saying it is an AI when asked, not re-asking, honouring a stop request the first time, not repeating sensitive numbers, ending after goodbye, and long silences.

The judge answers the rest, such as “Never claimed to be human” and “No pressure or guarantees”.

A verdict with evidence shows the line number and the exact words. Tap the quote and the transcript jumps to that line.

The judge must cite a line and a quote for every Failed answer. If the quote is not in that line, the QA Bot throws the answer away and marks it Not applicable. The review then carries the flag “Judgement checks were not trustworthy”. A Passed answer about something that did not happen, such as “did not pressure”, has nothing to quote and is accepted without a line.

  • Each question has a weight. The weights on a scorecard add up to 100.
  • Questions marked Not applicable leave the total. Only the rest count.
  • The score is the weight earned divided by the weight that applied.
  • A call passes at the scorecard’s pass mark or higher. The starter scorecard uses 80.
  • Any Failed critical question fails the call, whatever the score.

Each review is either “Scored automatically” or “Needs a look”. A call needs a look when any one of these is true:

  • It did not pass.
  • It carries a risk flag: “Possible vulnerable person”, “Complaint language”, “Pressure” or “Claimed to be human”. The QA Bot sets these from the words used.
  • The judge could not be trusted: “Judgement checks were not trustworthy”, “Judgement checks failed to run” or “Judgement checks skipped today (daily limit)”.
  • The judge said it had low confidence.

Everything else is scored automatically. Reviewing a call says what to do with the ones that need a look.

With no AI judge switched on, the QA Bot runs the rule checks only. Judgement questions show “Waiting for AI” and count as Not applicable. The review carries the flag “Judgement checks waiting for AI”.

That flag describes your setup, not the call. A call that passes its rule checks and has no risk flag is still scored automatically. A call that fails a rule check, or shows a risk flag, still needs a look.

Each workspace has a limit on judge calls per day. The default is 500, and a long transcript uses one call for each section the QA Bot splits it into. When the limit is used up, new calls get rule checks only and the flag “Judgement checks skipped today (daily limit)”. There is no screen for this limit in this build. It is changed through the API.