How to Create a Customer Support QA Scorecard
How to choose categories, set weighting, and calibrate reviewers so a QA scorecard measures something real rather than a reviewer's mood.
A QA scorecard is only useful if two different reviewers scoring the same conversation land close to the same number. Most scorecards fail that test, usually because the categories are too vague to score consistently.
Choosing categories
Keep it to six to nine categories. More than that and reviews take too long to run consistently; fewer and you lose the ability to give specific feedback.
| Category | Typical weight | What it actually measures |
|---|---|---|
| Greeting | 5% | Professional, on-brand opening |
| Verification | 10% | Correct identity or account checks performed |
| Accuracy | 25-30% | Information given was correct |
| Communication | 15% | Clarity, tone and pace |
| Resolution | 20% | The actual problem was addressed |
| Documentation | 10% | Notes and disposition are usable by the next person |
| Escalation handling | 10% | Correctly identified and routed when needed |
| Closing | 5% | Confirmed next steps and ended appropriately |
Writing scorable criteria
"Was the agent helpful?" cannot be scored consistently — two reviewers will disagree based on their own standards. "Did the agent confirm the customer's issue before proposing a solution?" can be answered yes or no by anyone watching the same conversation. Write every category as a specific, checkable behaviour.
Weighting for what matters to you
A technical support queue should weight accuracy heavily — wrong information is a worse outcome than a slightly slow reply. An appointment-setting campaign might weight booking technique and qualification accuracy more heavily than a support queue would. There is no universal correct weighting; there is a weighting that reflects what failure actually costs your business.
Calibrating reviewers
- Have two reviewers independently score the same five conversations
- Compare scores category by category, not just the total
- Where scores diverge significantly, discuss why — usually a category description was ambiguous
- Rewrite the ambiguous category and repeat until scores converge
Using the results
Deliver feedback with the specific conversation attached, category by category, not as a single aggregate number. An agent told "your score was 78" learns nothing actionable. An agent shown exactly where accuracy or documentation fell short can actually improve.
This guide reflects how we actually run campaigns and what we see go wrong. We have tried to be useful whether or not you ever work with us — including where that means recommending you do something other than outsource.