Scorecard best practices
Last updated: September 7, 2026
A scorecard is a prompt: it tells the AI exactly what "good" looks like for a section, then it scores against it. Most scoring complaints, too harsh, too lenient, inconsistent, come from ambiguity in that prompt rather than an AI bug. This article covers how to write criteria the AI can score consistently.
How Solidroad reads a scorecard
A scorecard is made of sections, each one thing you want to measure. For every conversation, the AI reads the transcript against your scorecard, scores each section with a quoted example, and rolls everything up into one 0-100% score.
Section type | Use it for | How it scores |
|---|---|---|
Graded | Soft skills: empathy, tone, rapport, clarity | A number on a 0-5 scale, based on the poor / average / strong descriptions you write |
Pass/fail | Objective steps: verifying an account, giving correct info | Full marks or zero, nothing in between |
Process-linked | Multi-step SOPs | Scored against an attached process document, can auto-detect at runtime |
Three advanced controls sit on top of any section: Exclusion/N/A marks a section not applicable and drops it from the score entirely, Section fail forces just that section to zero, Scorecard fail forces the whole scorecard to zero regardless of other scores. Sections marked N/A or excluded don't count on either side of the percentage.
Writing criteria the AI can score
The rule behind all of these: a human reviewer can read "rep was professional" and apply judgement, the AI needs the judgement written down. Replace anything subjective or implied with a specific, observable behaviour.
One behaviour per line. "Rep confirms the details, explains the product, and offers a discount" should be three separate lines, not one.
Write statements, not questions. "Rep confirms the customer's problem before proposing a fix," not "Did the rep confirm the problem?"
Describe the behaviour instead of naming a trait. "Professional" means nothing to the AI. "Rep uses a polite greeting, avoids slang, and maintains a respectful tone" does.
Match your wording to your intent. Generic phrasing ("explains the warranty terms") for when any reasonable version counts, specific phrasing ("states the warranty covers parts and labour for 12 months") for when the exact content matters.
Use "If X, then Y" for anything conditional, so a rep isn't penalised for a situation that never came up.
Use AND / OR / AND-OR to combine requirements when more than one step, or more than one acceptable path, is required.
Anchor every graded item with poor, average and strong transcript-style examples. This is the single biggest fix for "the AI is too harsh" or "I don't know what a 3 vs a 4 means."
Name the keywords you expect, and list acceptable synonyms, so a correct answer in different words isn't marked wrong.
Order items in the sequence they'd occur in a real interaction, greeting, verification, resolution, close.
Keep structure consistent. A fixed stem like "Rep must…" and one requirement per line parses far more reliably than a dense paragraph.
Only score what's in the transcript. A CRM field, a system screen, or an internal action the AI can't see can't be scored, attach it as a process document instead of referencing it.
Choosing graded vs. pass/fail
If the behaviour is… | Use | Because |
|---|---|---|
A soft skill (empathy, tone, rapport) | Graded | You want how well, not just whether |
An objective step (verify identity, give correct policy) | Pass/fail | It either happened correctly or it didn't |
A compliance must-do | Pass/fail + scorecard-fail | One miss should fail the interaction |
A documented multi-step procedure | Process-linked | The AI scores against the real SOP instead of you re-typing it |
Default graded sections to 0-5, not 0-10. A wide scale is where AI scoring drifts, an 8 and a 9 are hard to tell apart, so scores bunch up in the middle and produce exactly the "too harsh, inconsistent" pattern customers report. Reserve pass/fail for anything objective, it's the most consistent call an AI can make, with no partial credit to drift on.
Structuring the whole scorecard
Group into single-topic sections. Two skills sharing a section (say, active listening and open-ended questions) get blended into muddy feedback.
Start lean, 3-6 sections is a healthy starting point, prove it scores well, then expand.
Build one scorecard per topic and reuse it across simulations, rather than rebuilding it each time.
Set weights deliberately. Points are your priorities made visible, weight compliance and core-outcome sections above nice-to-haves.
Use exclusion/N/A criteria generously so a rep isn't penalised for a step that was never relevant to that conversation.
Before you publish
Every item tests one behaviour, no subjective words left undefined
Soft skills are graded on 0-5 with poor/average/strong anchor examples
Objective and compliance steps are pass/fail, compliance also carries scorecard-fail
Conditional behaviour uses If/Then, exclusion rules cover anything not always relevant
Sections are distinct, nothing tests the same thing twice
Nothing asks the AI to judge something outside the transcript
Weights reflect real priorities
Run in Testing mode against already-reviewed conversations, aim for roughly 90%+ agreement with your human reviewers before going live
Editing a scorecard doesn't retroactively change past scores, re-run the evaluation to apply the update. Keep calibrating monthly once live, and when scoring drifts on the same section repeatedly, fix that criterion rather than rewriting the scorecard over a single odd result.
Templates
Graded soft-skill section (Empathy)
Type: Graded (0-5)
Criteria: Rep acknowledges the customer's emotion before moving to a solution.
Poor (0-1): Ignores or dismisses the emotion, e.g. "That's not my problem."
Average (2-3): Names the emotion, e.g. "I understand that's frustrating."
Strong (4-5): Names the emotion and commits to act, e.g. "I understand
that's frustrating, and I'll fix it now."
Pass/fail compliance section (Identity Verification)
Type: Pass/Fail
Criteria: Rep must verify the customer before discussing account details.
- Rep verifies via registered email
OR
- Rep verifies via account ID and phone number
Scorecard-fail: If account details are shared before verification, fail
the scorecard.
Conditional section (De-escalation)
Type: Graded (0-5)
Exclusion (N/A): Mark N/A if the customer never expresses frustration.
Criteria: If the customer expresses frustration, then the rep
acknowledges it, apologises where appropriate, and slows the pace
before proposing next steps.
Related articles
If you have any further questions, contact the Solidroad team via the Get Help tab in the platform.