# Performance Calibration Pack

> Designs and prepares a performance calibration cycle end to end — session structure and grouping, the manager pre-work brief, a full facilitator run sheet with the questions that move a discussion from impression to evidence, in-the-room bias interrupters, the distribution and consistency checks to run before and after, and the follow-through and appeals route. Use when someone says they are "running calibration", "prepping for the performance cycle", "our ratings mean different things in different teams", "one manager rates everyone a 4", "we need a moderation session", "how do we stop the loudest manager winning the room", "should we force a distribution", "designing our performance review process", "the ratings feed into pay and I do not trust them", or is facilitating a talent review or ratings moderation meeting. For building the levels and competencies that ratings are made against, use career-framework-builder.

- Skill: `we-are-move/performance-calibration-pack` (Agent Skill, multi-file: 4 files)
- Install (CLI): `npx skillmds@latest add we-are-move/performance-calibration-pack`
- Raw SKILL.md: https://api.skillmd.com/api/skills/we-are-move/performance-calibration-pack/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Productivity
- Author: we-are-move (https://skillmd.com/u/we-are-move)
- Updated: 2026-09-22
- Page: https://skillmd.com/skills/we-are-move/performance-calibration-pack

---


# Performance Calibration Pack

Produces the pack a People team runs a calibration cycle from: session design, manager
pre-work brief, facilitator run sheet, bias interrupters, consistency checks, and the
follow-through. The test of the output is that someone who has never facilitated a
calibration session can run one on Tuesday and produce ratings that mean the same thing
across every manager in the room.

Calibration is the mechanism that makes a rating scale mean anything. Run well it is the
highest-integrity moment in the people calendar — the one point where a company checks its
own judgement against evidence. Run badly it is a room where the most senior or most
articulate manager's ratings survive intact and everyone else's get moved to fit, which is
worse than no calibration at all: it launders inconsistency as rigour and gives the
resulting pay and promotion decisions a legitimacy they have not earned.

## What you need to start

Minimum viable input: **the rating scale in use, roughly how many people are being rated,
and when the sessions are.** That is enough to design the whole thing.

If the user cannot answer even that, do not stall. Assume a four-point scale, assume ratings
feed pay, design the pack, and label the assumptions at the top. A pack built on stated
assumptions that the user corrects in five minutes beats a perfect pack they never got to
because the conversation turned into a form.

Ask in small batches, two or three at a time, reflecting back what you have before asking
more. Where the environment supports structured multiple-choice questions, use them for the
scale, the consequences and the session level — those are picks, not essays.

**Batch one — the scale and its consequences.** What is the rating scale and do the points
have written definitions? Do ratings drive pay, promotion, both, or neither? This is the
first question because it sets the stakes of everything that follows: a rating that only
informs a development conversation can tolerate looseness, a rating that sets a bonus
multiplier cannot.

**Batch two — the population and the calendar.** How many people are rated, across how
many managers, and how are they distributed by level and function? When do managers submit
and when do the sessions run? Population size drives session count and time per person;
the calendar drives whether the pre-work deadline is real or aspirational.

**Batch three — the room.** How many sessions, at what level do they run (team, function,
company), and who chairs them? Has this population been calibrated before, and what went
wrong? The failure story is the most useful answer in the whole intake — it usually names
the exact intervention the run sheet needs to carry.

Useful but never required: the current rating distribution, last cycle's ratings by manager,
the career framework or level definitions, the promotion criteria, and the pay matrix.

## Process

### 1. Fix the purpose of the session before designing it

Say plainly what calibration is for, because most rooms drift without it: **to check that
the same evidence produces the same rating in different managers' hands.** It is not a
meeting to decide ratings from scratch, not a talent review, not a succession discussion,
and not a budget meeting. Each of those is legitimate and each destroys calibration when
merged into it — the moment the room starts allocating a bonus pool, the discussion stops
being about evidence and becomes about who gets what. Schedule those separately, afterwards,
using the calibrated ratings as input.

### 2. Design the room

**Who is in it.** The managers whose people are being discussed, a facilitator who owns no
ratings in the session, and a People partner who holds the data and the checks. Nobody else
— every additional observer changes what managers are willing to say about their own people,
and honesty about a weak performer is the thing calibration most needs.

**Who should not be.** The rated person's skip-level, where their presence would silence
the direct manager. Anyone attending to protect a favourite. Executives dropping in for one
name — a senior leader who wants to intervene on one person does it through the facilitator
beforehand, on the record, and the room evaluates it like any other input.

**The facilitator owns no ratings in the room they run**, because a facilitator defending
their own team cannot credibly challenge anyone else's. Where a company is too small for
that, calibrate the facilitator's own team in a different session chaired by someone else.

### 3. Group the population by level and function, never by manager

Group people who are doing comparable work at a comparable level. Calibration is a
comparison exercise and the comparison is only meaningful between like and like — a senior
engineer and a junior marketer share a rating scale but nothing else.

Manager-by-manager walkthroughs are the default in most tools and they defeat the purpose.
Each manager presents their own internally ordered list, so the room evaluates that
manager's consistency with themselves — the one thing that was never in doubt — while the
inconsistency between managers, which is what you convened to find, stays invisible.
Interleaving by level forces the actual question: is this manager's "exceeds" the same as
that manager's "exceeds"?

Practical grouping rules: one session per level band per function where the population
supports it, and where it does not, group adjacent levels while holding the level
distinction explicitly in the discussion. Cap a session at what fits the time — two focused
sessions beat one that runs long and starts nodding people through at the ninety-minute
mark. Where someone spans functions or changed manager mid-cycle, place them where the work
was done and brief both managers to attend that session.

### 4. Set the time budget honestly, and decide what is discussed

Work from the total time available, not from an aspiration: **budget a few minutes per
person on average, and spend it unevenly on purpose.** Not everyone needs discussion, so
sort before the session.

- **Discuss** — the top and bottom of the scale, every rating change proposed since the
  last cycle, anyone whose rating carries a promotion or a performance-management
  consequence, anyone whose manager is new, anyone flagged by the pre-session checks.
- **Confirm quickly** — mid-scale ratings with a clear rationale and no flags. Read the
  name, state the rating, pause for challenge, move on. Ten seconds each is honest.
- **Never rubber-stamp silently.** Say aloud that a name is being confirmed without
  discussion so anyone can stop it. An unspoken name is an unchallenged name.

Tell the room the budget at the start and hold it. Sessions that overrun do not distribute
the shortfall evenly — they discuss the first third properly and rush the rest, which means
rating quality ends up depending on the running order.

### 5. Settle the distribution question — with the argument, not an assertion

This is the most contested design decision and the user will be challenged on it, so give
them the reasoning rather than a position to defend without one.

**The case for distribution guidance.** Left alone, ratings inflate. Managers rate
generously because it is the path of least conflict, because a high rating is a cheap way
to reward someone when pay is constrained, and because nobody wants to be the manager
whose team scored lowest. Once most of the population sits in the top two points, the
scale stops carrying information: pay cannot be differentiated, promotion signals nothing,
and the genuinely exceptional are indistinguishable from the merely fine. Guidance
counteracts that drift and gives managers cover — "the expectation is that most people are
performing well, and the top rating is rare" is a much easier conversation than a manager
alone deciding to be the strict one.

**The case against hard quotas.** A quota applied to a small population produces injustice
mechanically. A genuinely strong team of six does not contain a low performer, and forcing
one out of it means telling someone their rating reflects their team's size rather than
their work. Managers work this out quickly, and the rational response is to game it —
importing a weak hire to absorb the low slot, trading names across teams, rating
strategically for next cycle. Quotas also collide badly with small samples: the smaller the
group, the more its distribution is noise. And the reputational cost is durable — a quota
teaches people that their rating is a function of the distribution rather than of their
work, and once that reading takes hold they discount every part of the performance process,
including the parts that were sound.

**The recommendation: guidance, not quota, with a requirement to justify departures.**
Publish an expected shape for the population as a whole, state that it describes the company
and not any individual team, and require any manager whose team departs materially from it
to explain why in the session. The justification requirement is what makes guidance bite —
without it guidance is a suggestion; with it, a manager who rates six of eight people at the
top must produce six evidenced cases, which either holds up or collapses under one round of
questions. Two implementation notes:

- **Apply the shape at the level the sample is meaningful** — usually function or company,
  not team. A shape enforced on a team of six is a quota by another name.
- **Never move an individual's rating to fix a shape.** The shape is a diagnostic that says
  "look here". The only legitimate reason to change a rating is that the evidence does not
  support it. If a manager's distribution looks wrong and every individual case holds up
  under questioning, the distribution was right and the guidance was wrong for that team.

### 6. Build the manager pre-work brief

The session succeeds or fails on what managers bring. Specify it precisely and give it a
deadline that is genuinely before the session, not the night before.

What each manager brings for each person: the proposed rating, a written rationale
referencing the framework or level expectations, two or three specific pieces of evidence
from across the whole period, and their view on the two questions the room will ask —
compared to who, and what would have made this a higher rating.

The rating rationale standard, which is the part most managers get wrong:

- **Evidence, not adjectives.** "Consistently excellent" is not a rationale. "Led the
  billing migration, unblocked it when the vendor slipped, delivered three weeks late
  against a plan that had assumed no slippage" is.
- **Referenced to the level, not the person's own history.** The question is whether they
  met the expectations of the level they are at. Improvement is a development conversation;
  the rating is a standard.
- **Spanning the whole period.** Require at least one substantial piece of evidence from the
  first half. This single requirement does more against recency bias than any amount of
  in-room facilitation.
- **Written as if someone else will read it.** They may — an appeal, a future manager, a
  court. Professional, factual, about work.

Read `references/bias-interrupters.md` for the language patterns to warn managers about
before they write, particularly how personality-based and achievement-based language gets
distributed unevenly across a population.

The deadline discipline: submissions close far enough ahead that the People partner can run
the checks and the facilitator can build the agenda from them. State the consequence
plainly — a person whose rationale is not submitted is discussed last, from whatever their
manager can say live, and the facilitator records it. That is not a punishment, it is what
is physically possible.

### 7. Run the pre-session consistency checks

Cut the proposed ratings before the session and bring the cuts into the room. Their purpose
is to build the agenda: they tell the facilitator which names to spend time on. Cut by
manager, level, function, tenure band, full-time versus part-time and any reduced or
flexible arrangement, people who had leave during the period, people who changed manager
mid-cycle, and new hires who joined part-way through. Then, separately and carefully, by any
demographic dimension the organisation lawfully monitors.

**How to frame the demographic cut, and this matters.** Its purpose is to surface a pattern
for investigation — a group rated systematically lower is a signal that something in the
process is not working, and the investigation looks at the evidence, the rationales and the
managers, not at the individuals. It is never a basis for changing an individual's rating.
Adjusting any individual's rating because of their sex, race, age, disability, or any other
protected characteristic is unlawful discrimination, regardless of the direction of the
adjustment or the good intention behind it. Say this in the pack, in those terms, so that
nobody in the room can misread the chart as an instruction to rebalance.

The legitimate responses are: re-examine the rationales in the affected group against the
standard, check whether the work allocation that preceded the ratings was equitable, and
identify whether particular managers account for the pattern. Where a pattern is material or
recurring, run the analysis through legal counsel — in several jurisdictions analysis
conducted at counsel's direction attracts privilege, and analysis run casually in a
spreadsheet does not. `references/bias-interrupters.md` gives the specific cut for each bias
pattern and what a concerning result looks like.

### 8. Prepare the facilitator run sheet

Read `references/facilitator-script.md` in full before producing this section. It carries
the opening framing, the per-person protocol, the evidence questions, and the handling for
the moments that decide whether a session is worth anything: the manager who cannot
evidence a rating, the dominant voice, the trade, the reopened decision, and the room that
has gone quiet because it learned that challenge is expensive. The run sheet in the pack
should be usable standing up — timings, opening words, the question set per person, the
interventions in escalation order, and the closing.

### 9. Design the follow-through before the session, not after

The session is the middle of the process, not the end. Specify:

- **Who tells whom, and when.** The rating is delivered by the person's own manager, in a
  conversation, before it appears in any system. A rating that arrives by notification
  before a manager has explained it is the single most reliable way to turn a fair rating
  into a grievance.
- **What managers may say about calibration.** Give them the line: the rating was reviewed
  with other managers to make sure the standard is applied consistently across the company.
  Managers may not attribute a rating to the room ("I wanted to give you a 4 but they knocked
  it down") — it abandons their own accountability and tells the person their manager does
  not stand behind the decision. A manager who cannot defend a rating is a signal the session
  did not finish its job. They may not disclose other people's ratings or evidence, who said
  what, or their team's distribution.
- **People whose rating moved in the room.** Every change gets a written reason recorded
  against it, and the manager is briefed on how to deliver it beforehand. A manager who
  first learns the rating changed when they open the system will communicate it badly, and
  will be right to be annoyed.
- **The appeals route.** Who hears an appeal, on what grounds, in what window, and what
  outcomes are possible. Ground appeals in process and evidence — material evidence not
  considered, or the process not followed — rather than disagreement with the judgement, or
  every appeal becomes a re-litigation of the rating. The appeal is heard by someone who was
  not the deciding manager.
- **What feeds back into next cycle.** The facilitator's notes on which managers rated
  consistently, where the framework was ambiguous, and which parts of the run sheet did not
  work. Calibration quality compounds across cycles only if someone writes this down.

### 10. Produce the pack

Fill `assets/calibration-pack-template.md` and write it to disk as a Markdown file. Then
offer, without building unprompted: a slide version of the manager briefing for a kick-off
session, a spreadsheet for the consistency-check worksheet with the cuts pre-built, a
one-page facilitator card for the room, or a shareable page for the manager population.
Where the pack rests on assumptions — an assumed scale, an assumed population shape — carry
an assumptions block at the top naming each and what would replace it.

### The degraded case: no framework to rate against

If the user is designing their performance approach rather than running a cycle, say so and
produce the calibration design anyway — it is easier to design the framework when you know
what the session will need from it. Then name the dependencies that have to exist first,
flagging which the user has:

- **A rating scale with written definitions.** Points named with adjectives calibrate to
  nothing; each point needs a description of observable performance.
- **A framework to rate against** — levels and expectations, so "met expectations" has a
  referent. Without it, calibration compares managers' private standards and the most
  articulate private standard wins. Use `career-framework-builder` to build this; it is the
  prerequisite, not an optional companion.
- **A decision on what ratings drive.** Pay, promotion, both, neither. This sets how much
  rigour the process must carry and how much appeal exposure it creates.
- **A cycle calendar** with manager submission, checks, sessions and communication dated
  backwards from the pay effective date.
- Say plainly that the pack runs at reduced value until the framework exists.

## Output

A Markdown file with these sections, in this order:

1. **Cycle summary** — scale, population, sessions, dates, what ratings drive, assumptions block.
2. **Session plan** — grouping, attendees, time budget, discuss-versus-confirm sort.
3. **Distribution approach** — the guidance, the level it applies at, the justification requirement, and the reasoning to use when challenged.
4. **Manager pre-work brief** — a standalone document, sendable as is.
5. **Facilitator run sheet** — timings, opening, per-person protocol, interventions, close.
6. **Bias interrupters** — the in-room card.
7. **Consistency-check worksheet** — the cuts to run before and after, and what a concerning result looks like.
8. **Follow-through checklist** — communication, manager script, changed ratings, appeals, records.
9. **Next-cycle notes** — what to capture during the session for the next one.

## Records, evidence and legal exposure

Calibration records are discoverable. Ratings drive pay, promotion and sometimes exit, so
the written rationale for a rating can end up in front of a tribunal, a court, a regulator
or an equal-pay claim. Written well, that record is the organisation's best defence.
Written casually, it is the claimant's best evidence.

- **Write rationales as evidence about work.** Factual, specific, referenced to the level
  standard, free of speculation about personality, health, family circumstances or
  motivation. Anything a manager would not want read aloud should not be written.
- **Record every decision and reason, including changes.** A rating that changed with no
  recorded reason looks arbitrary years later, and arbitrary is the finding you least want.
- **Apply the process consistently.** Inconsistent application is where discrimination
  claims find their footing. Where exceptions are made — people on leave, late joiners, a
  team in reorganisation — write the rule and apply it to everyone in that circumstance.
- **Never adjust an individual rating on the basis of a protected characteristic.** It is
  unlawful when done to correct an imbalance as well as when done to create one. Patterns
  are investigated at process level; individuals are rated on evidence.
- **People on leave, on reduced hours, or with adjustments.** Rate the work done against the
  expectation that applied, not against a full-time colleague's volume — pro-rate the
  expectation, not the person. A disability-related adjustment is part of how the role is
  performed, not a performance deficit; where a rating turns on one, get it reviewed before
  it lands.
- **Anything approaching dismissal, performance management with an exit path, or redundancy
  selection goes to legal and local employment law review before it is actioned.** A session
  may legitimately identify that someone is not meeting the standard. It is not the forum
  that decides what happens next, and using calibration ratings as redundancy selection
  criteria without legal review is a well-worn route to a claim.
- **Personal data.** Use identifiers rather than names in any analysis file shared outside
  the session, and keep special-category data out of the rating record entirely.

This is structural guidance, not legal advice. Ask which jurisdictions are in scope —
consultation requirements, discrimination frameworks and data rules differ materially — and
have the cycle design reviewed locally where ratings drive pay or exit.

## Reference files

- **`references/facilitator-script.md`** — read at step 8, in full, before running a session.
  Opening framing, per-person protocol, the evidence questions in escalation order, and
  scripted handling for the manager who cannot evidence a rating, the dominant voice, the
  silent room, the trade, and the decision that reopens.
- **`references/bias-interrupters.md`** — read at steps 6, 7 and 8. Each bias pattern: how it
  presents, the words that signal it, the facilitator's intervention, and the pre- and
  post-session check that detects it.
- **`assets/calibration-pack-template.md`** — the output skeleton. Fill it; do not
  restructure it.

---

*Part of the [Claude Skills for TA and People Teams](https://github.com/we-are-move/claude-skills-for-ta-and-people-teams) collection — open-source skills
for in-house talent and people teams. Built and maintained by the team at MOVE.*

