Performance Calibration Pack
Produces the pack a People team runs a calibration cycle from: session design, manager pre-work brief, facilitator run sheet, bias interrupters, consistency checks, and the follow-through. The test of the output is that someone who has never facilitated a calibration session can run one on Tuesday and produce ratings that mean the same thing across every manager in the room.
Calibration is the mechanism that makes a rating scale mean anything. Run well it is the highest-integrity moment in the people calendar — the one point where a company checks its own judgement against evidence. Run badly it is a room where the most senior or most articulate manager's ratings survive intact and everyone else's get moved to fit, which is worse than no calibration at all: it launders inconsistency as rigour and gives the resulting pay and promotion decisions a legitimacy they have not earned.
What you need to start
Minimum viable input: the rating scale in use, roughly how many people are being rated, and when the sessions are. That is enough to design the whole thing.
If the user cannot answer even that, do not stall. Assume a four-point scale, assume ratings feed pay, design the pack, and label the assumptions at the top. A pack built on stated assumptions that the user corrects in five minutes beats a perfect pack they never got to because the conversation turned into a form.
Ask in small batches, two or three at a time, reflecting back what you have before asking more. Where the environment supports structured multiple-choice questions, use them for the scale, the consequences and the session level — those are picks, not essays.
Batch one — the scale and its consequences. What is the rating scale and do the points have written definitions? Do ratings drive pay, promotion, both, or neither? This is the first question because it sets the stakes of everything that follows: a rating that only informs a development conversation can tolerate looseness, a rating that sets a bonus multiplier cannot.
Batch two — the population and the calendar. How many people are rated, across how many managers, and how are they distributed by level and function? When do managers submit and when do the sessions run? Population size drives session count and time per person; the calendar drives whether the pre-work deadline is real or aspirational.
Batch three — the room. How many sessions, at what level do they run (team, function, company), and who chairs them? Has this population been calibrated before, and what went wrong? The failure story is the most useful answer in the whole intake — it usually names the exact intervention the run sheet needs to carry.
Useful but never required: the current rating distribution, last cycle's ratings by manager, the career framework or level definitions, the promotion criteria, and the pay matrix.
Process
1. Fix the purpose of the session before designing it
Say plainly what calibration is for, because most rooms drift without it: to check that the same evidence produces the same rating in different managers' hands. It is not a meeting to decide ratings from scratch, not a talent review, not a succession discussion, and not a budget meeting. Each of those is legitimate and each destroys calibration when merged into it — the moment the room starts allocating a bonus pool, the discussion stops being about evidence and becomes about who gets what. Schedule those separately, afterwards, using the calibrated ratings as input.
2. Design the room
Who is in it. The managers whose people are being discussed, a facilitator who owns no ratings in the session, and a People partner who holds the data and the checks. Nobody else — every additional observer changes what managers are willing to say about their own people, and honesty about a weak performer is the thing calibration most needs.
Who should not be. The rated person's skip-level, where their presence would silence the direct manager. Anyone attending to protect a favourite. Executives dropping in for one name — a senior leader who wants to intervene on one person does it through the facilitator beforehand, on the record, and the room evaluates it like any other input.
The facilitator owns no ratings in the room they run, because a facilitator defending their own team cannot credibly challenge anyone else's. Where a company is too small for that, calibrate the facilitator's own team in a different session chaired by someone else.
3. Group the population by level and function, never by manager
Group people who are doing comparable work at a comparable level. Calibration is a comparison exercise and the comparison is only meaningful between like and like — a senior engineer and a junior marketer share a rating scale but nothing else.
Manager-by-manager walkthroughs are the default in most tools and they defeat the purpose. Each manager presents their own internally ordered list, so the room evaluates that manager's consistency with themselves — the one thing that was never in doubt — while the inconsistency between managers, which is what you convened to find, stays invisible. Interleaving by level forces the actual question: is this manager's "exceeds" the same as that manager's "exceeds"?
Practical grouping rules: one session per level band per function where the population supports it, and where it does not, group adjacent levels while holding the level distinction explicitly in the discussion. Cap a session at what fits the time — two focused sessions beat one that runs long and starts nodding people through at the ninety-minute mark. Where someone spans functions or changed manager mid-cycle, place them where the work was done and brief both managers to attend that session.
4. Set the time budget honestly, and decide what is discussed
Work from the total time available, not from an aspiration: budget a few minutes per person on average, and spend it unevenly on purpose. Not everyone needs discussion, so sort before the session.
- Discuss — the top and bottom of the scale, every rating change proposed since the last cycle, anyone whose rating carries a promotion or a performance-management consequence, anyone whose manager is new, anyone flagged by the pre-session checks.
- Confirm quickly — mid-scale ratings with a clear rationale and no flags. Read the name, state the rating, pause for challenge, move on. Ten seconds each is honest.
- Never rubber-stamp silently. Say aloud that a name is being confirmed without discussion so anyone can stop it. An unspoken name is an unchallenged name.
Tell the room the budget at the start and hold it. Sessions that overrun do not distribute the shortfall evenly — they discuss the first third properly and rush the rest, which means rating quality ends up depending on the running order.
5. Settle the distribution question — with the argument, not an assertion
This is the most contested design decision and the user will be challenged on it, so give them the reasoning rather than a position to defend without one.
The case for distribution guidance. Left alone, ratings inflate. Managers rate generously because it is the path of least conflict, because a high rating is a cheap way to reward someone when pay is constrained, and because nobody wants to be the manager whose team scored lowest. Once most of the population sits in the top two points, the scale stops carrying information: pay cannot be differentiated, promotion signals nothing, and the genuinely exceptional are indistinguishable from the merely fine. Guidance counteracts that drift and gives managers cover — "the expectation is that most people are performing well, and the top rating is rare" is a much easier conversation than a manager alone deciding to be the strict one.
The case against hard quotas. A quota applied to a small population produces injustice mechanically. A genuinely strong team of six does not contain a low performer, and forcing one out of it means telling someone their rating reflects their team's size rather than their work. Managers work this out quickly, and the rational response is to game it — importing a weak hire to absorb the low slot, trading names across teams, rating strategically for next cycle. Quotas also collide badly with small samples: the smaller the group, the more its distribution is noise. And the reputational cost is durable — a quota teaches people that their rating is a function of the distribution rather than of their work, and once that reading takes hold they discount every part of the performance process, including the parts that were sound.
The recommendation: guidance, not quota, with a requirement to justify departures. Publish an expected shape for the population as a whole, state that it describes the company and not any individual team, and require any manager whose team departs materially from it to explain why in the session. The justification requirement is what makes guidance bite — without it guidance is a suggestion; with it, a manager who rates six of eight people at the top must produce six evidenced cases, which either holds up or collapses under one round of questions. Two implementation notes:
- Apply the shape at the level the sample is meaningful — usually function or company, not team. A shape enforced on a team of six is a quota by another name.
- Never move an individual's rating to fix a shape. The shape is a diagnostic that says "look here". The only legitimate reason to change a rating is that the evidence does not support it. If a manager's distribution looks wrong and every individual case holds up under questioning, the distribution was right and the guidance was wrong for that team.
6. Build the manager pre-work brief
The session succeeds or fails on what managers bring. Specify it precisely and give it a deadline that is genuinely before the session, not the night before.
What each manager brings for each person: the proposed rating, a written rationale referencing the framework or level expectations, two or three specific pieces of evidence from across the whole period, and their view on the two questions the room will ask — compared to who, and what would have made this a higher rating.
The rating rationale standard, which is the part most managers get wrong:
- Evidence, not adjectives. "Consistently excellent" is not a rationale. "Led the billing migration, unblocked it when the vendor slipped, delivered three weeks late against a plan that had assumed no slippage" is.
- Referenced to the level, not the person's own history. The question is whether they met the expectations of the level they are at. Improvement is a development conversation; the rating is a standard.
- Spanning the whole period. Require at least one substantial piece of evidence from the first half. This single requirement does more against recency bias than any amount of in-room facilitation.
- Written as if someone else will read it. They may — an appeal, a future manager, a court. Professional, factual, about work.
Read references/bias-interrupters.md for the language patterns to warn managers about
before they write, particularly how personality-based and achievement-based language gets
distributed unevenly across a population.
The deadline discipline: submissions close far enough ahead that the People partner can run the checks and the facilitator can build the agenda from them. State the consequence plainly — a person whose rationale is not submitted is discussed last, from whatever their manager can say live, and the facilitator records it. That is not a punishment, it is what is physically possible.
7. Run the pre-session consistency checks
Cut the proposed ratings before the session and bring the cuts into the room. Their purpose is to build the agenda: they tell the facilitator which names to spend time on. Cut by manager, level, function, tenure band, full-time versus part-time and any reduced or flexible arrangement, people who had leave during the period, people who changed manager mid-cycle, and new hires who joined part-way through. Then, separately and carefully, by any demographic dimension the organisation lawfully monitors.
How to frame the demographic cut, and this matters. Its purpose is to surface a pattern for investigation — a group rated systematically lower is a signal that something in the process is not working, and the investigation looks at the evidence, the rationales and the managers, not at the individuals. It is never a basis for changing an individual's rating. Adjusting any individual's rating because of their sex, race, age, disability, or any other protected characteristic is unlawful discrimination, regardless of the direction of the adjustment or the good intention behind it. Say this in the pack, in those terms, so that nobody in the room can misread the chart as an instruction to rebalance.
The legitimate responses are: re-examine the rationales in the affected group against the
standard, check whether the work allocation that preceded the ratings was equitable, and
identify whether particular managers account for the pattern. Where a pattern is material or
recurring, run the analysis through legal counsel — in several jurisdictions analysis
conducted at counsel's direction attracts privilege, and analysis run casually in a
spreadsheet does not. references/bias-interrupters.md gives the specific cut for each bias
pattern and what a concerning result looks like.
8. Prepare the facilitator run sheet
Read references/facilitator-script.md in full before producing this section. It carries
the opening framing, the per-person protocol, the evidence questions, and the handling for
the moments that decide whether a session is worth anything: the manager who cannot
evidence a rating, the dominant voice, the trade, the reopened decision, and the room that
has gone quiet because it learned that challenge is expensive. The run sheet in the pack
should be usable standing up — timings, opening words, the question set per person, the
interventions in escalation order, and the closing.
9. Design the follow-through before the session, not after
The session is the middle of the process, not the end. Specify:
- Who tells whom, and when. The rating is delivered by the person's own manager, in a conversation, before it appears in any system. A rating that arrives by notification before a manager has explained it is the single most reliable way to turn a fair rating into a grievance.
- What managers may say about calibration. Give them the line: the rating was reviewed with other managers to make sure the standard is applied consistently across the company. Managers may not attribute a rating to the room ("I wanted to give you a 4 but they knocked it down") — it abandons their own accountability and tells the person their manager does not stand behind the decision. A manager who cannot defend a rating is a signal the session did not finish its job. They may not disclose other people's ratings or evidence, who said what, or their team's distribution.
- People whose rating moved in the room. Every change gets a written reason recorded against it, and the manager is briefed on how to deliver it beforehand. A manager who first learns the rating changed when they open the system will communicate it badly, and will be right to be annoyed.
- The appeals route. Who hears an appeal, on what grounds, in what window, and what outcomes are possible. Ground appeals in process and evidence — material evidence not considered, or the process not followed — rather than disagreement with the judgement, or every appeal becomes a re-litigation of the rating. The appeal is heard by someone who was not the deciding manager.
- What feeds back into next cycle. The facilitator's notes on which managers rated consistently, where the framework was ambiguous, and which parts of the run sheet did not work. Calibration quality compounds across cycles only if someone writes this down.
10. Produce the pack
Fill assets/calibration-pack-template.md and write it to disk as a Markdown file. Then
offer, without building unprompted: a slide version of the manager briefing for a kick-off
session, a spreadsheet for the consistency-check worksheet with the cuts pre-built, a
one-page facilitator card for the room, or a shareable page for the manager population.
Where the pack rests on assumptions — an assumed scale, an assumed population shape — carry
an assumptions block at the top naming each and what would replace it.
The degraded case: no framework to rate against
If the user is designing their performance approach rather than running a cycle, say so and produce the calibration design anyway — it is easier to design the framework when you know what the session will need from it. Then name the dependencies that have to exist first, flagging which the user has:
- A rating scale with written definitions. Points named with adjectives calibrate to nothing; each point needs a description of observable performance.
- A framework to rate against — levels and expectations, so "met expectations" has a
referent. Without it, calibration compares managers' private standards and the most
articulate private standard wins. Use
career-framework-builderto build this; it is the prerequisite, not an optional companion. - A decision on what ratings drive. Pay, promotion, both, neither. This sets how much rigour the process must carry and how much appeal exposure it creates.
- A cycle calendar with manager submission, checks, sessions and communication dated backwards from the pay effective date.
- Say plainly that the pack runs at reduced value until the framework exists.
Output
A Markdown file with these sections, in this order:
- Cycle summary — scale, population, sessions, dates, what ratings drive, assumptions block.
- Session plan — grouping, attendees, time budget, discuss-versus-confirm sort.
- Distribution approach — the guidance, the level it applies at, the justification requirement, and the reasoning to use when challenged.
- Manager pre-work brief — a standalone document, sendable as is.
- Facilitator run sheet — timings, opening, per-person protocol, interventions, close.
- Bias interrupters — the in-room card.
- Consistency-check worksheet — the cuts to run before and after, and what a concerning result looks like.
- Follow-through checklist — communication, manager script, changed ratings, appeals, records.
- Next-cycle notes — what to capture during the session for the next one.
Records, evidence and legal exposure
Calibration records are discoverable. Ratings drive pay, promotion and sometimes exit, so the written rationale for a rating can end up in front of a tribunal, a court, a regulator or an equal-pay claim. Written well, that record is the organisation's best defence. Written casually, it is the claimant's best evidence.
- Write rationales as evidence about work. Factual, specific, referenced to the level standard, free of speculation about personality, health, family circumstances or motivation. Anything a manager would not want read aloud should not be written.
- Record every decision and reason, including changes. A rating that changed with no recorded reason looks arbitrary years later, and arbitrary is the finding you least want.
- Apply the process consistently. Inconsistent application is where discrimination claims find their footing. Where exceptions are made — people on leave, late joiners, a team in reorganisation — write the rule and apply it to everyone in that circumstance.
- Never adjust an individual rating on the basis of a protected characteristic. It is unlawful when done to correct an imbalance as well as when done to create one. Patterns are investigated at process level; individuals are rated on evidence.
- People on leave, on reduced hours, or with adjustments. Rate the work done against the expectation that applied, not against a full-time colleague's volume — pro-rate the expectation, not the person. A disability-related adjustment is part of how the role is performed, not a performance deficit; where a rating turns on one, get it reviewed before it lands.
- Anything approaching dismissal, performance management with an exit path, or redundancy selection goes to legal and local employment law review before it is actioned. A session may legitimately identify that someone is not meeting the standard. It is not the forum that decides what happens next, and using calibration ratings as redundancy selection criteria without legal review is a well-worn route to a claim.
- Personal data. Use identifiers rather than names in any analysis file shared outside the session, and keep special-category data out of the rating record entirely.
This is structural guidance, not legal advice. Ask which jurisdictions are in scope — consultation requirements, discrimination frameworks and data rules differ materially — and have the cycle design reviewed locally where ratings drive pay or exit.
Reference files
references/facilitator-script.md— read at step 8, in full, before running a session. Opening framing, per-person protocol, the evidence questions in escalation order, and scripted handling for the manager who cannot evidence a rating, the dominant voice, the silent room, the trade, and the decision that reopens.references/bias-interrupters.md— read at steps 6, 7 and 8. Each bias pattern: how it presents, the words that signal it, the facilitator's intervention, and the pre- and post-session check that detects it.assets/calibration-pack-template.md— the output skeleton. Fill it; do not restructure it.
Part of the Claude Skills for TA and People Teams collection — open-source skills for in-house talent and people teams. Built and maintained by the team at MOVE.