Building an interview scorecard your team will actually use
ยท 7 min read
Every hiring team has a scorecard. Almost none of them change a decision. The interviewer forms a view in the first ten minutes, fills the form in afterwards from memory, and the numbers rationalise a conclusion that was already reached.
The failure is structural, not a discipline problem. Scorecards are usually designed as documentation and then expected to function as evaluation, which are different jobs.
Three things that break scorecards
Rating traits instead of evidence
"Communication: 4/5" is not a measurement, it is an impression with a number attached. Two interviewers scoring the same candidate a 4 may have watched entirely different things and meant entirely different things.
The fix is to score demonstrated behaviour against a specific definition, and to require the evidence alongside the number. If an interviewer cannot cite the moment that produced the score, the score is not real.
Too many dimensions
A scorecard with twelve criteria gets filled in with a column of 4s. Attention is finite, and the twelfth criterion always gets whatever the eleventh got. Four or five criteria that genuinely predict performance in this specific role beat twelve generic ones every time.
Numbers with no shared meaning
Unanchored 1โ5 scales drift toward 3 and 4 because nobody wants to be the outlier. If your rubric does not say what a 2 looks like versus a 4, you are collecting sentiment, not assessment.
A structure that survives contact with real interviews
For each role, pick four or five competencies. For each competency, write one sentence defining it in role-specific terms, and anchor the scale at three points only.
Anchoring three points rather than five is deliberate โ people can hold three definitions in their head during a live conversation. Here is what a single competency looks like fully specified:
- Competency: Debugging under uncertainty โ can narrow down a failure they have never seen before, without the answer being available.
- 1 โ Describes following a runbook or escalating. No independent narrowing.
- 3 โ Describes forming a hypothesis, testing it, and revising when wrong. Names the tools used.
- 5 โ Describes hypotheses they rejected and why, what the misleading signal was, and what they changed afterwards so it would not recur.
- Evidence required: the specific incident, in their words.
The evidence line is the part that does the work. It converts the scorecard from an opinion form into a record, and it is the reason the number becomes defensible later.
Score during, not after
A score written twenty minutes after the interview is a memory of a feeling. The single highest-leverage change most teams can make is scoring each competency the moment it has been covered, while the candidate is still talking about something else.
This is uncomfortable at first and it is worth pushing through. It also surfaces coverage gaps while you can still fix them: if you are twenty-five minutes into a forty-minute interview with two competencies untouched, you can redirect. Discovering that afterwards is just a wasted interview.
Kill the average
Averaging competency scores into a single number destroys exactly the information you need. A candidate scoring 5, 5, 5, 1 and a candidate scoring 4, 4, 4, 4 both average 4, and they are completely different hires โ the first is a specialist with one real gap, the second is uniformly unremarkable.
Keep the vector. Decide against a threshold per competency, and make the trade-off explicitly: "we can live with a 2 on stakeholder management for this role, we cannot live with a 2 on debugging." That conversation is the actual hiring decision, and averaging hides it.
Debrief in writing before anyone speaks
Verbal debriefs are dominated by whoever speaks first and whoever is most senior. Both are anchoring effects and both are well documented.
- Every interviewer submits scores and evidence before the debrief, with no visibility of anyone else's.
- Open the meeting on the competency with the widest disagreement, not the overall verdict.
- Make people cite evidence, not restate scores. "What did you hear that I did not?" is the whole meeting.
- Record the reason for the final decision, not just the decision. This is what makes the next hire better, and what protects you if the decision is ever challenged.
Disagreement is a feature. Two interviewers scoring a competency 4 and 2 means one of them heard something the other missed โ that is the most valuable ten minutes of your hiring process, and consensus-first debriefs throw it away.
The honest constraint
Everything above asks interviewers to hold a rubric in their head, score against it live, capture verbatim evidence, and still run a conversation that feels human to the person on the other side. That is a lot, and it is why scorecards degrade back into after-the-fact rationalisation within about a quarter.
This is the part worth automating. Hiriso scores competencies against role criteria as the interview happens and attaches the candidate's actual words as the evidence, so the scorecard is complete before the call ends โ without the interviewer splitting their attention between the person and the form.