How to Run Performance Calibration

If your ratings swing wildly depending on who the manager is, you don’t have a rating problem — you have a calibration problem. This guide walks People leaders and team leads through how to run performance calibration that produces fair, defensible results: how to prep managers, structure the meeting, surface bias, and turn the outcome into growth instead of resentment.

What’s inside

  • What performance calibration actually is

  • Why calibration matters more than the rating scale

  • Who should be in the room

  • How to prepare managers before the meeting

  • How to run the calibration meeting itself

  • Surfacing and correcting bias in real time

  • Turning ratings into fair, defensible decisions

  • Common calibration mistakes to avoid

  • How Blomma helps you run fairer calibration at scale

What performance calibration actually is

Calibration is the step between individual managers scoring their people and the company committing to final ratings. Each manager arrives with proposed ratings; the group reviews them side by side and adjusts so that a “meets expectations” from one team means the same thing as a “meets expectations” from another.

That’s the whole point. Left alone, one manager grades on a curve, another hands out top marks to everyone, and a third quietly protects a struggling report. None of them are acting in bad faith — they just have different bars. Calibration meetings exist to align those bars against a shared standard, using evidence rather than gut feel.

Done well, calibration protects your people from the lottery of who they happen to report to. Done poorly, it becomes horse-trading where the loudest manager wins. The difference is almost entirely in the preparation and facilitation, which is what the rest of this guide covers.

Why calibration matters more than the rating scale

Teams spend weeks debating whether to use a 3-point or 5-point scale, then skip the step that actually determines fairness. A perfect scale applied inconsistently still produces unfair performance ratings. A rough scale applied consistently is defensible.

Calibration is where consistency gets manufactured. It’s also where you catch the patterns that quietly drive attrition: the high performer who’s been “meets expectations” for three cycles because their manager is stingy, or the whole team rated a notch high because their lead confuses being liked with being effective.

Those patterns have real downstream cost. Ratings feed promotions, compensation, and who gets stretch work. When they’re miscalibrated, your strongest people notice they’re being under-recognized relative to peers on other teams — and they leave. If you want to understand how far the ripple travels, it’s worth reading how executives actually make promotion decisions: calibration is the raw material those decisions are built on. Fix calibration and you fix the input to nearly every talent decision that follows.

Who should be in the room

Keep the group small enough to move but broad enough to see across teams. The right default is the managers who wrote the ratings for a given level or function, plus their skip-level leader, plus a facilitator from People who owns the process but not the outcomes.

The facilitator’s job is neutrality: keep time, enforce the standard, and make sure evidence — not volume — drives changes. The skip-level’s job is cross-team perspective, spotting when one group’s “exceeds” looks like another’s “meets.” The managers’ job is to bring specifics and stay open to being wrong.

Group people by level or role, not by team. Calibrating all the senior individual contributors together, across teams, is what makes the standard real — a mid-level engineer on one team gets compared to mid-level engineers everywhere, not just their own pod. Avoid stacking the room with too many layers of hierarchy; when a manager’s boss’s boss is watching, honesty evaporates and everyone rates to protect themselves.

How to prepare managers before the meeting

Most bad calibration meetings were lost in prep. Send managers the standard — what each rating actually means in observable behavior — well before the session, and ask them to write a short evidence note for each rating: what the person did, not how the manager feels about them.

Ask for specifics tied to the level’s expectations. “Strong communicator” is an opinion; “rewrote the onboarding runbook that cut new-hire ramp from six weeks to three” is evidence. If a manager can’t point to evidence, that’s a signal to discuss, not a rating to rubber-stamp.

Prime managers to expect challenge and to see it as normal, not personal. The managers who take calibration best are usually the ones who’ve already learned to let their people struggle productively rather than shielding them — they show up with a clear-eyed read instead of a defensive one. Give them the questions you’ll ask in advance: Where might you be too generous? Too harsh? Who surprised you this cycle? Managers who’ve thought about their own bias before the meeting change their ratings for the right reasons once they’re in it.

How to run the calibration meeting itself

Start by restating the standard out loud so the room shares one definition. Then work level by level, not manager by manager — this keeps the focus on comparing people at the same bar rather than defending one manager’s whole list.

A simple flow: display proposed ratings for a level, look at the distribution, and spot the outliers. Discuss the outliers first — the surprisingly high and surprisingly low. For each, the manager gives their evidence in 60 seconds, the group asks clarifying questions, and the facilitator asks whether the evidence matches the standard. Adjust, or don’t, and move on. Don’t debate every rating; the middle of the distribution rarely needs it.

Timebox ruthlessly. A calibration meeting that runs long turns into fatigue-driven concessions, which is the opposite of fair. If a case can’t be resolved with the evidence present, park it, assign someone to gather what’s missing, and decide offline. The goal is aligned ratings backed by evidence — not unanimous comfort in the room.

Surfacing and correcting bias in real time

The facilitator’s sharpest tool is a set of standing questions that expose bias without accusing anyone. “What’s the evidence for that?” catches halo effects. “Would you rate them the same if they’d joined last month?” catches tenure bias. “Are we describing impact or personality?” catches likeability creep.

Watch for recency bias — a great or terrible final month coloring a whole year — and for the “strong team” trap, where being on a high-performing team inflates an average contributor. Watch the distribution too: if one manager’s ratings cluster far higher or lower than everyone else’s, name the pattern, not the person. “Your team is rated a notch above the room. Walk us through the top two.”

Bias toward underrepresented groups often hides in vague language. When feedback for one group skews to personality (“abrasive,” “not a culture fit”) while another gets concrete achievements, that gap is the finding. Make it a rule: no rating changes on adjective alone. Every adjustment needs a behavior attached. That single rule removes most of the bias a calibration meeting is meant to catch.

Turning ratings into fair, defensible decisions

A rating you can’t explain to the person is a rating that will cost you their trust. So the test for every final rating is simple: can the manager give the report two or three concrete examples that justify it and a clear path to the next level? If not, the calibration isn’t finished.

Capture the evidence and the reasoning as you go, not afterward. When a rating changes in the room, write down why. That record is what protects the decision later — in a promotion case, a compensation review, or, if it comes to it, a disputed exit. Defensible means documented.

Then close the loop with managers so the message to each report is consistent with what the group agreed. The fastest way to destroy calibration’s credibility is for a manager to walk out and tell their report “I fought for you but they overruled me.” Align on the narrative, not just the number. Fair performance ratings that come with clear, evidence-based reasoning turn a review from a verdict into a conversation about growth.

Common calibration mistakes to avoid

A few patterns sink calibration meetings again and again. Forcing a distribution — mandating that exactly 10% land in the bottom bucket — replaces one bias with another and punishes strong teams. Rank rather than distribute: identify genuine outliers and let the shape fall where the evidence puts it.

Don’t let seniority win arguments. If the most senior person in the room can override evidence with a raised eyebrow, you’ve built theater, not calibration. The facilitator has to protect the standard even against the highest-paid opinion.

Don’t calibrate once a year and call it done. Bar drift happens continuously; a lightweight mid-cycle check keeps managers honest and makes the annual meeting far less contentious. And don’t skip the debrief — after each cycle, ask managers what felt unfair or unclear, and fix the process, not just the ratings. Finally, don’t treat calibration as purely defensive. The best outcome isn’t just accurate scores; it’s managers who leave with a sharper, shared sense of what great work looks like.

How Blomma helps you run fairer calibration at scale

Calibration is only as good as the managers feeding into it — and most managers were never taught how to write evidence-based ratings, name their own bias, or deliver a hard rating without damaging the relationship. That’s usually where People teams get stuck: you can standardize the process, but you can’t sit in every manager’s prep.

Blomma gives every manager an always-on coach for exactly those moments. Before calibration, a lead can talk through where they might be too generous or too harsh and pressure-test their evidence. After it, they can rehearse the delivery conversation so a tough rating lands as a growth path, not a gut punch. It’s the kind of in-the-moment coaching that used to be reserved for executives, available to your whole team — see how Blomma coaching works and what it looks like across a team.

The payoff compounds: more consistent ratings, managers who can defend their calls, and review outcomes your best people actually trust.

Bring Blomma to your team →

Related reading

  • How execs make promo decisions

  • Let them fail

  • Sure, AI can wipe out middle management

Start your growth journey with Blomma

Start your growth journey with Blomma

Growth looks good on you

AI powered coaching, accountability and insights to help you grow

©2026 Blomma. All rights reserved.

Growth looks good on you

AI powered coaching, accountability and insights to help you grow

©2026 Blomma. All rights reserved.

Growth looks good on you. AI powered coaching, accountability and insights to help you grow.

©2026 Blomma. All rights reserved.