DECISION BRIEF

Should we grade performance on a bell curve?

Question

Should we grade performance on a bell curve?

The "bell curve" in performance management means forced distribution — requiring ratings to fit a fixed shape, classically Jack Welch's 20/70/10: a top 20% rewarded, a middle 70% developed, a bottom 10% managed out. It travels under several names — stack ranking, forced ranking, the vitality curve, "rank and yank." The promise is discipline: stop everyone being rated "great," force managers to differentiate, surface your best and worst.

The real decision a leader faces is three-way: mandate a forced curve, use a distribution as a loose calibration check (a sanity test, not a quota), or neither. The distinction between the second and the first is where this whole debate lives.

Evidence

The curve's founding assumption is empirically wrong — performance isn't normally distributed. The largest study on this question, O'Boyle & Aguinis's "The Best and the Rest" (Personnel Psychology, 2012; cited finding), examined 198 samples covering more than 633,000 people and found individual performance consistently follows a power-law (Paretian) distribution, not a normal one — a small number of "stars" produce a hugely disproportionate share of output. A bell curve assumes a symmetrical spread that mostly doesn't exist; forcing ratings into it therefore manufactures a bottom tier and caps a top tier that the real distribution doesn't support. The model is mis-specified before a single rating is entered.

When you simulate forced ranking, the outcomes are close to random. A 2025 agent-based simulation, Tournament-Based Performance Evaluation and Systematic Misallocation (McEntire, arXiv, CC BY 4.0), found forced-ranking mechanisms generate classification error rates of 32% under idealised conditions and 53% under realistic ones — i.e. results "essentially indistinguishable from random allocation." It also finds that cross-team calibration "transforms evaluation into influence contests where persuasive managers secure promotions independent of merit," and concludes these systems persist mainly to look like rigorous procedure rather than to differentiate performance accurately.

Ranking people against each other entrenches luck and erodes meritocracy. Don't follow the leader: how ranking performance reduces meritocracy (Livan, Royal Society Open Science, 2019, CC BY 4.0) models what happens when people chase rank by imitating top performers: it is "a largely self-defeating endeavour" that concentrates advantage among early, often lucky, winners, widens inequality, and weakens the actual correlation between talent and success. Ranking, in other words, doesn't just measure a hierarchy — it manufactures a sticky one.

The companies that invented it abandoned it. Microsoft dropped stack ranking in 2013 (Harvard Business Review, "Don't Rate Your Employees on a Curve"; cited finding) after internal reviews blamed it for hoarding of information, killed collaboration, and talent flight; GE — the home of Welch's vitality curve — and Amazon moved away around 2016, with Adobe and Accenture among many others. The broader retreat from rigid annual ratings is now mainstream, as The Conversation's "Judgement day: farewell but not goodbye to performance reviews" (CC BY-ND) documents — farewell to the forced ritual, not to evaluation itself.

Disagreement

ViewThe claimWhere it holds — and breaks
"Force the curve — it drives performance"A fixed distribution stops rating inflation, forces managers to differentiate, and reliably surfaces top and bottom performers.Holds on the need: rating inflation is real and differentiation matters. Breaks on the method: the distribution is imposed, not discovered; simulation shows it misclassifies a third to half of people; and the firms that pioneered it dropped it for the collateral damage. You can get differentiation without forcing a shape.
"Any look at the distribution is wrong"Rating spreads, calibration, comparing teams — all of it is corrosive ranking; judge each person on their own.Holds against rigid quotas and peer-vs-peer ranking. Breaks if it means never checking the spread: looking at distribution to catch inflation or a lenient/harsh manager — calibration, not a quota — is a legitimate sanity check. The harm is in forcing the curve, not in glancing at it.

The real split isn't differentiate vs. don't. It's forced quota vs. calibration: a mandatory shape every team must hit (manufactures a fake bottom, pits colleagues against each other) versus managers comparing notes to keep ratings honest and consistent (keeps judgement, drops the quota).

Peoplense Verdict

Don't force the curve. Calibrate instead.

  • What to rely on: the forced bell curve rests on a distribution that doesn't match reality (performance is power-law, not normal), simulation shows it misclassifies 32–53% of people, and the companies that built it abandoned it. That's about as clear as workplace evidence gets.
  • What to avoid: a mandatory quota (every team must produce a bottom 10%); ranking employees against teammates (turns colleagues into competitors and punishes strong people on strong teams); and — highest risk of all — tying a forced bottom tier to firing.
  • The point that matters: the goal you actually want — honest differentiation and no grade inflation — is delivered by calibration: managers discussing ratings together for consistency and bias, using the distribution as a sanity check, never a rule. Keep the discipline; drop the quota.

What to do today

  1. Remove any mandatory quota. Stop requiring a fixed percentage in each rating band. That single change converts a forced curve into honest assessment.
  2. Replace forced ranking with a calibration session. Bring managers together to compare ratings for consistency and bias — but explicitly without a target shape. The output is fairer ratings, not a filled-in curve.
  3. Treat inflation as its own problem. If everyone is "exceeds," the fix is clearer standards and manager discipline (and spot-checking against real outcomes) — not forcing a distribution that hides the inflation instead of correcting it.
  4. Never wire a distribution to dismissal. The "bottom 10% out" rule is the least evidence-backed and most legally and culturally hazardous piece — and the fastest way to destroy trust and collaboration.
  5. Differentiate where it's genuinely real. Power-law performance means a few authentic stars exist — recognise and reward them specifically. Don't manufacture a fake bottom tier just to "balance" the top.

GCC Relevance

There is no Gulf-specific study on forced distribution that we can cite, so read the following as informed inference from non-Gulf evidence applied to the regional context, not as Gulf findings.

The collaboration and "relational fairness" costs likely bite harder in high-power-distance workplaces. In Gulf organisations where hierarchy is strong and feedback is indirect (the same dynamic flagged in our performance-ratings brief), a forced ranking that publicly sorts people into winners and losers is more likely to be experienced as a relationship rupture than a neutral process — and, as with ratings generally, a forced curve can quietly launder relationship-based decisions as if they were objective measurement.

A forced "bottom 10% out" collides directly with Saudization obligations. A rule that mechanically pushes out a fixed share each cycle runs into Nitaqat / Saudization headcount commitments — cutting Saudi nationals to satisfy a quota carries localization consequences that a US-style "rank and yank" never had to weigh. In the KSA context this isn't just a culture question; it's a workforce-planning and compliance one.

Honest scope: the regional power-distance and Saudization points are context and inference; the core evidence against forced distribution is robust but non-Gulf. We'll add Gulf-specific evidence as it emerges.

Sources

Library sources (analysed on Peoplense under the editorial-summary contract — our summary on-site, full text + author voice at the original publisher):

Cited findings (named and linked, not republished — these do not carry an open licence):

  • O'Boyle, E. & Aguinis, H. (2012), The Best and the Rest: Revisiting the Norm of Normality of Individual Performance, Personnel PsychologyWiley. 198 samples, 633,000+ individuals: individual performance is power-law (Paretian), not normal. Cite-only / paywalled — the load-bearing "the bell curve's assumption is wrong" finding.
  • Don't Rate Your Employees on a Curve, Harvard Business Review (2013) — article. Microsoft's abandonment of stack ranking and why. Cite-only.
  • Vitality curve (Jack Welch's 20/70/10; GE and Amazon's later retreat) — overview. Background reference for the model's origin and history.

Scope of the evidence: the three open-licensed sources above are held in our library, and each licence was verified on the publisher's own page (McEntire CC BY 4.0 · Livan CC BY 4.0 · The Conversation CC BY-ND 4.0). O'Boyle & Aguinis and the HBR/Microsoft piece are paywalled, so they are named and linked as cited findings and are not reproduced here. Complements Should we remove performance ratings?.

Get the Monday Brief

Evidence-based people development research, summarized weekly. Free. No ads. Every article links to its source.

Email used only to deliver the brief. Unsubscribe anytime.