Most interview feedback is a paragraph written four days late that says the candidate was strong technically but they are not sure about communication. Nobody can act on that, and two interviewers writing the same sentence usually mean different things by it.
A scorecard is not bureaucracy added to interviewing. It is the interview having a defined subject. Four to six criteria, a scale with words attached to each number, and evidence written down before anyone picks a score.
Why scorecards, briefly
Structured interviews predict job performance better than unstructured ones. Researchers have argued about the size of the gap for thirty years and rather less about the direction of it. The mechanism is not mysterious: asking every candidate about the same things, against the same standard, produces comparisons that mean something, and asking whatever comes to mind produces a series of pleasant conversations.
The second benefit is more immediate. A scorecard makes it possible to disagree usefully. Two people who both say "I liked them" cannot resolve anything. Two people looking at a criterion where one scored a 2 and the other a 4 have located the disagreement, and it usually turns out they heard different answers or were applying different standards, both of which are fixable.
A scorecard does not make the decision. It makes the disagreement specific enough to be worth having.
Start with four to six criteria, not a wish list
Criteria come from what the person will actually do in their first year, which is not the same as the job advertisement. Write that list first, in plain language, then pick the four to six things that will most determine whether they succeed.
The cap matters. Beyond about six, interviewers stop discriminating between criteria and start giving everyone roughly the same score across the board, which is the halo effect wearing a spreadsheet. Three or four is fine for a volume role.
Each criterion needs a definition specific enough that two people would recognise the same thing. "Communication" is not a criterion. "Can explain a technical trade-off to a non-technical stakeholder and get a decision out of it" is one, and you can immediately see what question would test it.
Be careful with anything that is really a proxy for similarity. Culture fit is the most common offender, and in practice it often measures whether the candidate reminds the interviewer of themselves. If you want to assess values or ways of working, name the specific behaviour you expect and test for that instead.
The rating scale, and why four points
Four points, with words against each. An even number removes the neutral middle, which on a five-point scale absorbs every rating an interviewer does not want to defend.
| Rating | Meaning | What it sounds like in the evidence box |
|---|---|---|
| 1. Well below | Would struggle with this part of the role | Could not describe an example, or the example showed the opposite behaviour |
| 2. Below | Some capability, would need support | Relevant example, but at smaller scope or with the difficult part handled by someone else |
| 3. Meets | Can do this part of the role independently | Clear example at the right scope, owned the outcome, could explain the reasoning |
| 4. Exceeds | Stronger than the role requires, could raise the bar for others | Example at greater scope or complexity, plus evidence of teaching or improving how others do it |
Write role-specific versions of those anchors for each criterion. Generic anchors get interpreted differently by everyone, which defeats the purpose. It takes twenty minutes per role and you reuse it for every hire in that family.
One rule worth stating in the template: a 3 is a good outcome. Teams that treat "meets" as faint praise end up with scorecards where everything is a 4, and a scale that only uses its top half is not a scale.
The scorecard template
Copy this structure into whatever your ATS or document tool uses. The order is deliberate, because evidence comes before the number.
- Header: role, candidate, interviewer, date, round, and the criteria this round is responsible for
- For each criterion: the definition, the evidence box, then the 1 to 4 rating with its anchors visible
- Evidence box prompt: what did the candidate actually say or do, in their words rather than your summary
- Red flags or concerns, kept separate from the ratings so a single worry does not silently depress every score
- What the next interviewer should probe, which is the field that makes a multi-round loop worth having
- Overall recommendation: strong no, no, yes, strong yes
- Confidence in that recommendation: low, medium or high
No maybe. The recommendation field exists to force a position, and an interviewer who genuinely cannot choose should say "no, low confidence" and explain what would change their mind. That is more useful than a maybe, because it tells the hiring manager exactly what the next round needs to establish.
The "what to probe next" field is the one most templates leave out and the one that changes how a loop feels. It converts four separate interviews into one investigation where each round starts from what the last one found.
Who scores what, and why it matters
Assign every criterion to exactly one round. Without that, the predictable thing happens: three interviewers all ask about the same strength because it is the interesting one, and nobody assesses the criterion that would have been decisive.
| Criterion | Owned by | Assessed how |
|---|---|---|
| Core craft for the role | Technical or functional round | Work sample or a walkthrough of something they built or ran |
| Judgement under ambiguity | Hiring manager round | A situation from their history, probed for the reasoning rather than the outcome |
| Working with others across functions | Cross-functional interviewer | A specific disagreement they had to resolve, and what they did |
| Ownership and follow-through | Hiring manager or peer | Something that went wrong and what they did about it afterwards |
| Motivation and role fit | Recruiter screen | Why this role, what they want next, and what they are leaving behind |
Interviewers should see their own criteria before the interview and the full scorecard only afterwards. Knowing what everyone else is covering is useful for the loop design and unhelpful in the room, where it encourages people to stray into someone else's territory.
Evidence first, then the rating
The single most effective rule in this whole document. The evidence box gets filled in before the score is chosen, and the score has to be defensible from what is written in it.
This works because it interrupts the order in which people actually form judgements. Interviewers decide in the first few minutes and then assemble reasons, which is well documented and applies to everyone including the people who are certain it does not apply to them. Writing the evidence first does not eliminate that, but it makes the gap visible when the note says very little and the rating says 4.
Write what the candidate said, close to their words. "Said they rewrote the reconciliation job after the third production incident, walked through why the retry logic was the cause, mentioned they wrote the runbook afterwards" is evidence. "Strong ownership" is a conclusion, and conclusions are what the rating is for.
A practical consequence: fill the scorecard within an hour of the interview. Not the same evening, and certainly not the morning of the debrief. What you have at that point is a summary of your impression rather than a record of the conversation.
The debrief: independent first
Everyone submits their scorecard before the debrief opens. This is a hard rule, and it is the one most often broken because someone was busy and wants to fill it in during the meeting.
Ratings shared before they are recorded converge, and they converge on whoever is most senior, most confident or speaks first. Independent ratings first, then discussion, preserves the disagreement that is the entire point of having several interviewers.
The meeting itself is short. Each interviewer states their rating on the criteria they owned and the evidence behind it. The discussion goes to criteria where the ratings differ, not to a general conversation about the candidate. The hiring manager decides and the decision is recorded with the reason. Fifteen or twenty minutes is usually enough, and a debrief that runs to an hour is generally one where nobody wrote anything down beforehand.
If someone changes their rating during the discussion, that is legitimate, and it should be recorded as a change with the reason. Silently editing a scorecard afterwards to match the group is how a structured process quietly becomes an unstructured one.
The decision rule
Agree this before the first candidate is interviewed, because a rule set afterwards is a rule set to fit the person you already like.
- Which criteria are must-haves, where a 1 or 2 ends the conversation regardless of everything else
- Which are trainable, where a 2 is acceptable with a named development plan
- What a single strong no from any interviewer means: a veto, or a discussion
- Who decides when the loop is split, which should be one named person rather than a vote
- What happens with an incomplete loop, where a criterion was never actually assessed
That last one is worth its own sentence. A loop where the decisive criterion was never covered is not a close call. It is an incomplete assessment, and the right response is another conversation rather than a decision made on what happens to be available.
Where scorecards go wrong
- Filled in during or after the debrief, which produces a record of the group decision rather than four independent views
- Criteria that are proxies for similarity, culture fit being the usual one
- A five-point scale where almost everything lands on 3
- Everyone scoring everything, so the specialist criterion gets rated by four people who cannot assess it
- Evidence boxes containing conclusions instead of quotes, which makes the rating unreviewable
- The same scorecard reused across very different roles, so the criteria stop matching the work
- Scores treated as arithmetic, where a total is computed and the highest number wins regardless of which criterion was weak
That last one deserves care. Adding up scores across criteria implies they are equally weighted and interchangeable, and they are not. A candidate who is a 4 on three things and a 1 on the must-have is not a strong candidate with an average of 3.25.
Adapting it for different roles
Volume roles: three criteria, one round, and a scorecard short enough to complete in three minutes. The value there is consistency across dozens of interviews rather than depth on any one.
Technical roles: the work sample or exercise gets its own rubric, scored against the same 1 to 4 scale so it sits alongside the interview ratings rather than in a separate document nobody reads at the debrief.
Leadership roles: keep the same structure and expect a longer loop and more evidence per criterion. The temptation at senior levels is to drop the scorecard because everyone in the room is experienced, which is precisely where unstructured judgement performs worst and where a bad decision costs the most.
One more practical benefit. Consistent criteria applied to every candidate, with the evidence recorded, is also what lets you show that a decision was made on the requirements of the role. That matters for fairness on its own terms, and it matters if anyone ever asks you to explain a hiring decision.
Frequently asked questions
What should an interview scorecard include?
Role, candidate, interviewer and round in the header, four to six criteria with definitions, a four-point rating scale with written anchors, an evidence box completed before the rating, a note on what the next interviewer should probe, and an overall recommendation of strong no, no, yes or strong yes.
How many criteria should an interview scorecard have?
Four to six for most roles, and three or four for high volume roles. Beyond six, interviewers stop distinguishing between criteria and give similar scores across all of them.
Should an interview rating scale have four or five points?
Four. An even number removes the neutral middle, which on a five-point scale collects every rating an interviewer does not want to defend. Write role-specific anchors for each of the four points.
When should interviewers fill in the scorecard?
Within an hour of the interview, and always before the debrief starts. Anything later is a record of an impression rather than of the conversation, and anything shared before submission causes ratings to converge on the loudest view.
Should interview scores be added up to make a decision?
No. Totalling scores treats criteria as equally weighted and interchangeable. Agree in advance which criteria are must-haves where a low score ends the conversation, and which are trainable.
Is culture fit a valid scorecard criterion?
Not as usually written, because it frequently measures whether the candidate resembles the interviewer. If you want to assess values or ways of working, name the specific behaviour you expect and test for that instead.
How long should a hiring debrief take?
Fifteen to twenty minutes when scorecards were submitted beforehand. Each interviewer gives their rating and evidence on the criteria they owned, discussion focuses on where ratings differ, and one named person decides.
What is the difference between a scorecard and an interview evaluation form?
In practice they are the same document. What separates a useful one from a form is anchored ratings, criteria owned by specific rounds, and evidence recorded before the score is chosen.
If you build one thing from this page, build the criteria list and the anchors for a single role you hire for repeatedly. That is the part that takes judgement and gets reused. The form around it is a header and seven fields, and it matters far less than whether the four to six things you chose are the four to six things that actually determine whether someone succeeds in the job.