DECISION BRIEF

You have argued about the rating scale, the curve and the cadence. Has anyone asked whether the person being rated believes the process was just?

Question

You have argued about the rating scale, the curve and the cadence. Has anyone asked whether the person being rated believes the process was just?

Six of our own briefs address performance management, and every one of them is about the machinery: whether to grade on a curve, whether to keep ratings at all, whether to move to continuous feedback, whether to let AI draft the review, whether there are too many KPIs, what makes the whole system work. Not one asks what the person on the other side of the table concluded about the fairness of it.

That is not a soft question. In the research it is the variable that carries the outcome.

Evidence

Fairness works through satisfaction with the appraisal, not through the score. A structural-equation study of 161 Italian schoolteachers (154 in the model) found that appraisal satisfaction fully mediated the relationship between perceived appraisal justice and two work outcomes — job performance and job satisfaction. Justice did not act on those outcomes directly. It acted by changing how people felt about the process they had just been through, and that feeling carried everything downstream. (Read in library)

Fairness perceptions predict who leaves. In 775 employees of foreign-invested enterprises in Seoul, distributive, procedural and interactional justice all reduced turnover intention, directly and through lower organisational conflict and higher job satisfaction. (Read in library)

And it holds at national scale. The 6th Korean Working Conditions Survey — 25,285 employees — modelled fairness alongside recognition, involvement and transformational leadership as workplace resources shaping engagement, burnout and job satisfaction, with effects varying by sector. This is the largest sample in this brief by two orders of magnitude. (Read in library)

Reviewing more often is not automatically fairer, or better. A study of performance-appraisal interval found an inverted-U relationship with proactive work behaviour. Longer gaps raise delay of gratification, which helps, and simultaneously raise perceived uncertainty, which hurts. There is an optimum, and both ends of the range sit below it. (Read in library)

How the conversation is framed changes whether it lands. Across a survey of 382 managers describing real feedback incidents and two role-play studies with 370 and 162 executives and MBA students, diagnostic feedback — analysing the causes of past performance — was found to amplify the recipient's self-serving attributional bias. Future-focused advice did not. (Read in library)

Evaluation instruments carry demographic bias at scale. Across 523,703 student evaluations covering 3,123 teachers at a large Australian university over seven years, scores showed systematic bias against female instructors and instructors from non-English-speaking backgrounds, after controlling for course and teacher variation. (Read in library)

Disagreement

The bias finding is an analogy, and we are marking it as one. It measures students evaluating teaching, not managers appraising staff. The mechanism — an evaluator rating a person against a loose standard, at volume — is close enough to be worth your attention and not close enough to be evidence about your appraisal. Treat it as a reason to check your own distributions, not as a finding about them.

The interval study cuts against a fashion we have written about approvingly elsewhere. Our own brief on continuous feedback treats more frequent contact as the direction of travel. The inverted U says the relationship is not monotonic. Both can be true — continuous coaching is not the same act as a more frequent formal appraisal — but anyone reading both briefs deserves to see the tension rather than have it smoothed over.

Two of the six studies use convenience or single-country samples, and the appraisal-justice study rests on 161 teachers in one profession in one country. The direction is consistent across all six; the effect sizes are not transferable.

Peoplense Verdict

Fairness is not the soft half of performance management. It is the variable that decides whether any of the machinery works. Across six studies the same pattern holds: perceived justice predicts acceptance, wellbeing and turnover intention, and it does so through how people experience the process rather than through the number they receive. An organisation can fix its scale, drop its curve, shorten its cycle and still change nothing, because none of those touch the perception that carries the outcome.

Two qualifications keep this honest. More frequent is not automatically fairer — the interval evidence describes an optimum, not a slope. And the instrument itself can be the problem: at 523,703 observations, evaluation scores carried demographic bias, which is a reason to look at your own rating distributions by group before concluding your process is neutral.

⚠️ None of these six studies was conducted in the Gulf. Italy, South Korea twice, China, Australia, and one multi-sample study run with international executives and MBA students. The direction of the finding is consistent enough to act on; the numbers are not local, and we are not going to present them as though they were.

What to do today

  1. Ask one fairness question, separately from the rating. After the next review cycle, ask the person rated whether they felt the process was fair — not whether they agreed with the outcome. Those are different questions and only the first predicts anything.
  2. Pull your rating distribution by gender and by nationality. You are not looking for proof of bias; you are looking at whether anyone has ever checked. If the answer is no, that is the finding.
  3. Count your interval honestly. Not what the policy says — what actually happened. If the gap between real conversations is twelve months, the inverted U says you are on the wrong side of the optimum, and so does everyone who has been waiting.
  4. Move one conversation from diagnosis to advice. In the next review you run, spend the time on what the person should do next rather than on reconstructing why last quarter went the way it did.
  5. Before you trust the answers, read the Gulf note below.

GCC Relevance

The instrument you would use to check fairness is subject to the fear it is measuring.

Reviewing our Arabic edition of a related brief in July, one of our editors made the point plainly: in the Gulf, responses to engagement surveys and interviews are frequently exaggerated for fear of the consequences of being candid. She had raised the same issue on our own performance-review template, noting that peer and 360 input can be guarded or over-positive, particularly at leadership levels.

That has a specific consequence for this brief. Every study above measures perceived justice by asking people. In a setting where candour carries perceived risk, a high fairness score is ambiguous: it may mean the process is fair, or it may mean the person answering did not feel safe saying otherwise. The reading that flatters you is not the only one available.

Two practical implications follow. Do not treat a good fairness score as a result until you know whether dissent is sayable in that team — our brief on whether speaking up gets punished is the companion question. And look at behaviour alongside the survey: turnover intention, rating distributions by group, and whether anyone has ever appealed an outcome. Those are harder to answer politely.

We hold no Gulf-specific study on appraisal justice. That is an evidence gap, not a settled matter, and it is the honest state of the question here.

Sources

Get the Monday Brief

Evidence-based people development research, summarized weekly. Free. No ads.

Email used only to deliver the brief. Unsubscribe anytime.