How to Create a Customer Support QA Scorecard

How to choose categories, set weighting, and calibrate reviewers so a QA scorecard measures something real rather than a reviewer's mood.

A QA scorecard is only useful if two different reviewers scoring the same conversation land close to the same number. Most scorecards fail that test, usually because the categories are too vague to score consistently.

Choosing categories

Keep it to six to nine categories. More than that and reviews take too long to run consistently; fewer and you lose the ability to give specific feedback.

CategoryTypical weightWhat it actually measures
Greeting 5% Professional, on-brand opening
Verification 10% Correct identity or account checks performed
Accuracy 25-30% Information given was correct
Communication 15% Clarity, tone and pace
Resolution 20% The actual problem was addressed
Documentation 10% Notes and disposition are usable by the next person
Escalation handling 10% Correctly identified and routed when needed
Closing 5% Confirmed next steps and ended appropriately

Writing scorable criteria

"Was the agent helpful?" cannot be scored consistently — two reviewers will disagree based on their own standards. "Did the agent confirm the customer's issue before proposing a solution?" can be answered yes or no by anyone watching the same conversation. Write every category as a specific, checkable behaviour.

Weighting for what matters to you

A technical support queue should weight accuracy heavily — wrong information is a worse outcome than a slightly slow reply. An appointment-setting campaign might weight booking technique and qualification accuracy more heavily than a support queue would. There is no universal correct weighting; there is a weighting that reflects what failure actually costs your business.

Calibrating reviewers

  1. Have two reviewers independently score the same five conversations
  2. Compare scores category by category, not just the total
  3. Where scores diverge significantly, discuss why — usually a category description was ambiguous
  4. Rewrite the ambiguous category and repeat until scores converge

Using the results

Deliver feedback with the specific conversation attached, category by category, not as a single aggregate number. An agent told "your score was 78" learns nothing actionable. An agent shown exactly where accuracy or documentation fell short can actually improve.


Where this comes from

This guide reflects how we actually run campaigns and what we see go wrong. We have tried to be useful whether or not you ever work with us — including where that means recommending you do something other than outsource.