Hiring Scorecard and Interview Kit
Hiring goes wrong in two predictable ways, and both happen before the first interview. The role was never defined beyond a title, so five interviewers assess five different jobs and the debrief becomes an argument about what the role even is. And the interviews were unstructured, so what got measured was rapport: how much the candidate resembled the interviewer, how comfortable the conversation felt. Both failures feel like good judgement while they are happening, and the confidence of the people conducting an unstructured interview is unrelated to its accuracy.
The cost is not the recruiting fee. A mis-hire in a role that matters costs the salary, the months before anyone admits it, the manager's attention through a performance process, the work that did not happen, and the damage to a team that watched someone struggle in a job nobody had defined. In a small organisation, one mis-hire into a leadership role can consume a quarter of the chief executive's year.
This skill front-loads the work: the role defined as outcomes before anyone is contacted, every interviewer given specific competencies and a written standard for a good answer, and the decision made against evidence recorded before anyone spoke in the room.
When to use this, and when not to
Use it when a role is being opened, whether or not a job description already exists. Use it when a process is already running and going badly, which usually shows up as interviewers disagreeing without being able to say what they disagree about. Use it when a hire has to be defensible to someone who was not in the room, and when the same role has already failed once, where the scorecard should be written from what went wrong rather than from the old description. Use it for internal promotions too, where the bar is most often applied loosely.
Do not use it to decide whether the role should exist. That question is about capacity, budget and sequencing and belongs to annual-planning-and-headcount; a scorecard for a role that should not be filled is a well-built path to the wrong destination. Do not use it for an engagement scoped by deliverable rather than by capability, which belongs to sow-and-scope. Do not use it to assess people already in post; that is a performance question, and reusing hiring language there reads as a threat.
onboarding-plan takes over the day the offer is accepted, and takes the outcomes straight from this scorecard. sales-team-competency-assessment covers evaluating an existing team against a capability model, which is the same vocabulary applied to a different decision.
What you need before starting
The reason the role exists now. What is not getting done, or will not get done next quarter. This drives the outcomes, and a role that cannot answer it produces a scorecard of duties. Missing: ask the hiring manager one question, what breaks if we do not fill this, and write the answer down verbatim.
The compensation range and level. Missing: get it before advertising. Discovering the gap at offer stage costs the candidate, the pipeline, and usually the runner-up who has already accepted elsewhere.
Who the role works with, and who among them gets a say. Interviewers, and separately, anyone whose objection would block an offer. Missing: ask the hiring manager to name both lists explicitly. The person who is not on the panel and can still veto is the most common cause of a process that collapses after the offer is drafted.
The panel's availability. Missing: book their slots before the first candidate is contacted, and tell the hiring manager what a slow panel costs. Strong candidates have other processes running, and the process that moves in two weeks beats the one that moves in six even when the second offer is better.
Constraints. Location, time zone, work authorisation, travel, start date. Missing: state them as assumptions in the scorecard and confirm before advertising, since each can invalidate a whole pipeline late.
The method
Write the scorecard first, and get it agreed before anything else exists. One page, agreed in writing by the hiring manager and at least one person the role serves. Four parts.
The mission, one sentence, from the organisation's point of view rather than the candidate's: "own the operating layer around the chief executive so that decisions are made once and executed". Not a summary of tasks.
The outcomes, three to five, each stating what will be true at twelve months and how it will be known. "Forecast accuracy within ten percent by the second quarter." "The weekly leadership meeting running on written status by month two, with decisions logged." The test for an outcome is whether two people could disagree at twelve months about whether it happened. If they could, it is a duty wearing an outcome's clothes.
The competencies, five to eight, each defined by an observable behaviour and each traceable to an outcome. If a competency does not serve an outcome, cut it; if an outcome has no competency, the scorecard is incomplete. Split into must-have and strong-plus, honestly, because the must-have list is what a debrief resolves against. Include the working style the team genuinely needs, stated as behaviour: works without a queue of assigned tasks, writes before meeting, holds a position under pressure from someone senior. Style mismatch is the most common cause of a hire that fails despite obvious competence.
The constraints: level, range, location, start date.
Write the advertisement from the scorecard, never the reverse. Lead with the mission and the outcomes, because that is what a strong candidate self-selects on. State the competencies honestly, including the hard parts of the job. Say what the person gets: scope, exposure, what the role becomes. State the range. Cut the list of twenty responsibilities; a role defined by outcomes does not need it. Any requirement you would waive for the right person is not a requirement, and its only effect is to deter the candidates who read it literally.
Design the stages, each with a purpose, an owner and a length. The default shape, adjusted for level:
A screen of thirty minutes on motivation, constraints, compensation alignment and the two most essential competencies, whose job is to kill mismatches cheaply and which should reject a majority. A hiring manager interview of sixty minutes on the outcomes: evidence of having done the nearest equivalent, at what scale, with what result. A bounded work sample, the highest signal stage and the one most often cut for schedule reasons, which is the wrong trade. Two or three panel interviews of forty-five minutes, each interviewer assigned specific competencies so that coverage is complete and no competency is assessed by everyone while another is assessed by nobody. References with at least one former manager. Then a decision meeting within two days of the last interview.
A stage earns its place only if it assesses a competency no earlier stage covers. A fourth conversation over the same ground adds delay and the appearance of thoroughness, and degrades the candidate's view of the organisation.
Set an elapsed target, under three weeks from first contact to offer, and hold it publicly.
Build the interview guides. For each interviewer: the competencies they own, three to five questions per competency in past-behaviour form, and the probes that get past the rehearsed answer. What was the situation, what did you personally do as distinct from the team, what happened, what would you do differently. The probes matter more than the questions; a polished candidate has an answer ready for every opening question in general use and has rarely rehearsed the third follow-up.
Beside each question, write what a strong answer contains and what a weak one sounds like. This is the part usually skipped and the part that does the work: it converts an impression into a judgement against a standard, and lets an interviewer who has never hired for this role score consistently. A strong answer on prioritisation names what was dropped and who was told; a weak one describes a system for organising work without naming a single trade-off.
Reserve five minutes for the candidate's questions and record what they asked, which is evidence about what they were assessing you on.
Design the work sample around a real outcome from the scorecard. Bounded, honest about time, and never work the organisation would otherwise pay for, which candidates notice and resent. Three hours is the ceiling for a take-home; beyond that it selects for who has free time rather than who is capable. Write the rubric before any candidate sees the task, listing the three or four things a strong submission does and the traps a weak one falls into. Give every candidate identical materials and identical time, and where someone cannot do a take-home for accessibility or scheduling reasons, offer the live version rather than dropping the stage.
Score in writing, before the room. Each interviewer submits per-competency scores on an anchored one to four scale, with the evidence behind each, before the decision meeting starts. Use four points, not five, because a midpoint gets used to avoid taking a position. Written first is not a formality: the first opinion voiced in a debrief anchors everyone else, and the effect is largest on the most junior person present, who is often the one who spent the most time with the candidate.
Run the decision meeting on evidence, competency by competency. Not on overall impressions and not on who liked whom. Where two interviewers disagree, resolve it by comparing what each actually observed rather than by averaging, because a disagreement usually means one of them saw something the other did not ask about.
The decision rule is against the scorecard: does the evidence support that this person achieves these outcomes here. A candidate who fails a must-have is a no regardless of how strong the rest is. A likeable candidate with no evidence on the outcomes is a no. A candidate strong on outcomes and unproven on one strong-plus is a yes with a development plan that carries into onboarding. The hiring manager decides; the kit makes that decision defensible rather than replacing it with an arithmetic mean.
Take references as evidence, not as ceremony. With the candidate's consent, and with at least one former manager. Ask about the context, the outcomes achieved, a specific strength with an example, where the person needed the most support, how they handled a setback, and whether the reference would hire them again and into what. Listen for the praise that avoids the competency you asked about. A reference who will not answer the last question directly has answered it.
Close the loop. Check what the process predicted against the day ninety review in
onboarding-plan: which competency was over-weighted, which stage produced the signal that mattered, which question produced nothing. Three of these reviews make the next scorecard materially better, and almost no organisation does them, which is why hiring processes do not improve.
Fairness, consistency and legal boundaries
Every candidate for a role gets the same stages, questions, work sample, time and rubric. Consistency is what makes comparison meaningful, and it is the strongest available defence if a rejection is challenged.
Questions concern the competencies only. Nothing about family plans, health, age, religion, nationality beyond confirmed work authorisation, or anything that indirectly reveals a protected characteristic, such as the year someone finished school. Compensation history is prohibited to ask in several jurisdictions and is a poor predictor everywhere; ask about expectations against your stated range.
Employment law varies by jurisdiction. This skill designs the assessment; local legal or people-team review covers compliance. Where none is available, keep every question on observable job behaviour, which is the safe default nearly everywhere.
Worked example
Situation. A fifty-person analytics company had failed twice to hire a head of customer success. The first hire left after seven months by mutual agreement; the second offer was declined after a six-week process. The chief executive described the role as "someone senior to own our customers". The existing job description ran to twenty-two bullet points and included both "build the customer success function from scratch" and "five years in a SaaS environment", which had between them produced a pipeline of candidates from much larger companies who expected a team already in place.
Task. Fill the role within eight weeks, with a definition the chief executive and the head of sales both agreed on, since their disagreement had been visible to candidates in the previous process.
Action. The scorecard took two sessions and one argument. The reason the role existed turned out to be specific: renewals were being handled by whoever had time, three of the twelve largest accounts had no named contact on the company side, and net revenue retention had fallen from 104 to 96 percent over three quarters.
That produced four outcomes. Net revenue retention back above 102 percent by month twelve. Every account above 40,000 of annual value with a named owner and a documented account plan by month four. A renewal forecast delivered monthly with accuracy within ten percent from month six. And two customer success managers hired and productive by month nine.
The argument was over the fourth. The head of sales wanted a hands-on operator to hold the top accounts personally; the chief executive wanted a builder who would hire a team. Those are different people, and the previous process had advertised for both, which explained the first departure: the person hired was a builder, given no headcount for five months, and spent that time doing account work she was not especially good at. The resolution was explicit sequencing written into the scorecard, hands-on through month six and hiring from month seven, with the headcount committed in writing at offer stage.
Competencies came to six, including the style item that mattered most: comfortable being the only person in the function for six months, stated as behaviour and probed for directly.
The wrong turn: the first work sample asked candidates to build a customer health scoring model, the intellectually interesting part of the job. Three were sent it before it was withdrawn. The submissions were all competent, all similar, and distinguished nothing, and the task took six hours despite being scoped at three. It was replaced with a narrower sample: an anonymised account with eighteen months of usage decline, three support tickets and a renewal in ninety days, and a request for a one-page account plan and the opening email. That version separated the field immediately.
Scores went in writing. Two candidates reached the panel. Candidate A scored 4 on renewal ownership and 2 on building from nothing, with the evidence being that every process she described had been inherited and refined. Candidate B scored 3 and 4, with a weaker revenue record at a smaller company but a documented case of building an account management function of three people from nothing.
The debrief nearly went the other way, because Candidate A interviewed better and the chief executive said so first, before the written scores were read out. Reading the scores changed the discussion: the must-have was building from nothing, and A had no evidence for it in any of four attempts to find some.
Result. Candidate B accepted at nine weeks, one week over target, delayed by a reference that took eleven days to schedule. Account plans were done by month five, one month late. Net revenue retention reached 101 percent at month twelve, short of the 102 target, and the review concluded the target had been set on optimism rather than on the renewal base. The retrospective also found one clear error in the kit: forecasting discipline had been assigned to an interviewer who left the process and was assessed by nobody, and the forecast was the thing that ran late.
A second scenario, where it goes differently
The same company, six months later, hiring two customer success managers into the function that now exists, with a manager in post who has done the job.
Almost everything compresses. The scorecard is written in an hour because the outcomes are known and the competencies were validated by the previous process. The stages drop from six to four: screen, manager interview, work sample, one peer conversation. The work sample is the same account plan exercise, reused deliberately, because a rubric applied to seven candidates is far better calibrated than a fresh one. The elapsed target drops to twelve days.
What does not compress is the written scoring and the identical treatment of every candidate, and here it matters more rather than less: two roles filled from one pipeline invites the temptation to move the bar between the first and second offer, which is how a second hire gets made at a standard the first would have failed. The rule is that the rubric is set once for the pipeline and the second offer is made against the same bar or not at all.
What changed is the amount of definition needed, not the discipline. Volume hiring makes the structure cheaper per hire, not less necessary.
Output
The kit is one folder with seven artefacts.
Scorecard, one page:
ROLE: [title] Level: [ ] Range: [ ]
Hiring manager: [name] Decision by: [date]
Mission: [one sentence]
OUTCOMES AT TWELVE MONTHS
| # | Outcome | How it will be measured | Baseline today |
COMPETENCIES
| Competency | Observable behaviour | Must-have / strong-plus | Serves outcome |
CONSTRAINTS
Location, time zone, start date, work authorisation, travel.
Job advertisement, written from the scorecard, leading with mission and outcomes, stating the range.
Stage plan:
| Stage | Purpose | Owner | Length | Competencies assessed | Target day |
Interview guide, one per interviewer:
| Competency | Question | Probes | Strong answer contains | Weak answer sounds like |
Work sample: the candidate brief, the time limit, the materials, and separately the rubric, dated before the first candidate received it.
Scoring template, submitted before the decision meeting:
| Competency | Score 1 to 4 | Evidence observed | Not assessed |
Plus one line: would you hire this person into this role, yes or no, and the strongest reason.
Reference script and decision meeting agenda: read the scores, competency by competency, resolve disagreements on evidence, decide against the must-haves, then agree offer terms and the development items that carry into onboarding.
Failure modes
The scorecard written as duties. Recognise it because entries begin with manage, support or oversee, and because nobody could tell at twelve months whether they happened. Rewrite each as a state of the world with a number and a date.
The unstructured senior interview. Recognise it when the most senior interviewer runs a conversation rather than a guide, on the grounds that they can tell. Their instinct is no better than anyone else's and their opinion carries more weight, which makes this the most damaging exception. Give them the shortest guide, two competencies, and the same written score.
The process that drifts. Recognise it when elapsed time passes four weeks or a candidate has waited five days for a response. The best candidates leave first, so the pipeline degrades in quality faster than in quantity, which is invisible until the final two are both compromises.
A competency assessed by nobody. Recognise it by building the coverage map before the first panel interview: every must-have competency should appear on at least two interviewers' guides, and none should appear on all of them. A blank row is the thing that runs late after the hire, which is exactly what happened to forecasting discipline in the worked example. Fix it while the guides are being written, not at the debrief when the evidence no longer exists.
The work sample that outgrew its cap. Recognise it by asking the first three candidates how long the task actually took them. Where the median runs more than half an hour past the stated cap, the sample has stopped measuring capability and started measuring who had a free weekend. Cut its scope rather than cutting the stage, and reissue the same cut version to everyone still in the process.
Scores written after the room. Recognise it when any interviewer's scores arrive during the decision meeting or in the thread afterwards. Anything submitted once the first opinion has been voiced is not independent evidence, whatever it says. Postpone the meeting by a day, once, and state why; it does not happen twice.
References taken as ceremony. Recognise it when a reference call runs under fifteen minutes and produces no example with a date attached. Require at least one former manager and one specific instance behind every strength claimed, and treat a reference who will not answer whether they would hire the person again as having answered it.
Edge cases
A role nobody has done before. No equivalent experience exists to look for. Build competencies from first principles, weight the work sample heavily since it is the only direct evidence, and accept a wider band on the outcomes with an explicit review at month six.
One candidate, internal, and the decision is effectively made. Run the scorecard anyway, honestly. Its value shifts from selection to onboarding: the gaps it finds become the development plan and the support the person gets, which determines whether an inevitable promotion works.
Hiring above your own level of expertise. Where the hiring manager cannot assess the craft, bring an external assessor into one stage: an advisor, a board member, a peer from another organisation who has done the job. Give them one or two competencies and the same written rubric, and never let the whole assessment rest on someone with no accountability for the outcome.
An urgent hire with a week to decide. Compress stages, never the scorecard or the written scores. One combined interview and a ninety minute live work sample, competencies still assigned, scores still written before the discussion. Record what evidence you are choosing not to gather, so the risk is a decision rather than an accident.
The role changes mid-process. Stop, rewrite the scorecard, and tell everyone in the pipeline what changed. Continuing on a stale scorecard produces an offer against a job that no longer exists.
Quality bar
- Every outcome is a state of the world at a date, measurable enough that two people could not disagree about whether it happened.
- Every competency traces to an outcome, and every outcome to at least one competency.
- Every interviewer has assigned competencies and a written standard for what a strong and a weak answer contain.
- The work sample mirrors real work, is capped at three hours, and its rubric is dated before the first candidate saw the task.
- Scores and evidence are submitted in writing before the decision meeting opens.
- The process, the questions and the rubric are identical for every candidate for the role, with any adjustment recorded.
- The decision is stated against the must-have competencies, naming the evidence for each.
- Compensation range, level and constraints were fixed before the first candidate was contacted.
Adapting this to your context
The defaults come from hiring into commercial companies of fifty to two hundred people, where the employer sets its own process. Local law and sector rules override anything here.
- Under three weeks from first contact to offer. Realistic where a hiring manager books their own panel. Public bodies, universities and unionised employers often mandate an advert period, a panel composition rule and a fixed scoring form; build the kit inside their process rather than alongside it.
- The three-hour work sample cap. Set so the task selects for capability rather than free time. For clinical, teaching or safety-critical roles the equivalent is an observed session or a simulation, with the same rubric written before anyone is assessed.
- Compensation history and the range. Asking about history is prohibited in several jurisdictions, and pay transparency rules increasingly require the range in the advertisement. Check both before writing it, not before the offer.
- The one to four anchored scale. Chosen to remove the midpoint. Where a mandated form uses five points or weighted criteria, score on their scale and keep the written evidence per competency, which is what does the work.
- What not to change. Every candidate for a role meets the same stages, questions and rubric, and scores with their evidence are written before the decision meeting opens.
Related skills
annual-planning-and-headcount decides whether the role should exist and hands the approved role to this skill. onboarding-plan takes the scorecard's outcomes into the first ninety days, and its day ninety review is where this kit learns whether it was right. decision-memo is the format for taking a contested hire or a headcount exception to a decider. sales-team-competency-assessment applies the same competency vocabulary to people already in post. process-documentation-sop documents the hiring process itself once its shape is stable enough to repeat.