DECISION BRIEF

Should we put AI in our performance reviews?

Question

Should we put AI in our performance reviews?

Every HR-tech vendor in 2026 is pitching an AI feature on top of performance management — review drafting, peer-feedback summarization, capability scoring, manager coaching, AI-usage tracking. Adopting something is the easy decision; where to place it in the review process is the hard one.

Two large platforms went in opposite directions in the same period. Meta added employee AI usage as a performance signal; Duolingo dropped it after public backlash. Neither has published outcome data, so what we have is two reported company policies, not a controlled comparison. This page is for people-development leaders deciding which side to land on — and which uses of AI carry the least risk.

Evidence

These findings frame the decision; the last two add legal and implementation depth.

Duolingo reversed course publicly — after about a year, not overnight. Duolingo announced its "AI-first" policy, with AI use tracked in performance reviews, in April 2025. Entrepreneur covered the about-face in April 2026 — roughly twelve months later — reporting that the CEO dropped AI usage as a formal performance metric after employees questioned whether leadership wanted them to use "AI for AI's sake." Storyboard18 reported the implementation detail: AI usage is no longer a performance metric there. What this is and isn't: a reported policy reversal at one company, driven by internal pushback. The reporting does not measure trust, morale or performance, so it cannot be read as evidence of an outcome — only that the company withdrew the metric after employees raised concerns.

An expert commentary against rewarding tool usage. In HR Executive, Peter Cappelli argues against AI-usage incentives — bonuses, ratings, promotion criteria — on the grounds that tying compensation to a tool's usage rewards activity over outcome, distorts adoption patterns, and breaks the link between performance and impact. This is expert commentary, not an empirical study: it carries the weight of informed argument and supports caution about rewarding usage instead of results — it does not measure what happens when organisations do it.

Meta is running the opposite experiment. Fortune reported in January 2026 that Meta is reworking performance reviews to reward output over effort, enforcing a forced-distribution model (roughly 20% Outstanding / 70% Excellent / 7% Needs Improvement / 3% Not Meeting Expectations). Alongside that, HR Grapevine reported that Meta now ties employee AI usage to reviews and rewards through its "Checkpoint" system — adding a new "Meta Award" tier with a bonus multiplier up to 300%, and, for software engineers, tracking AI-generated code lines among 200+ data points. Both items above are reported company policies, not outcome evaluations — neither Fortune nor HR Grapevine reports measured results, and Meta has published none. The honest read: two large platforms adopted opposite policies within the same six-month window, and there is no outcome evidence on either side to call a winner.

Cornell research: narrative reviews are perceived as fairest — and it is not a study about AI. Cornell researchers ran four experiments with 1,600 participants, comparing numerical-only, narrative-only, and combined feedback on identical evaluations. Narrative-only feedback was generally perceived as the fairest and helped recipients understand how to improve. The researchers also cautioned against removing numerical ratings entirely — numbers give necessary clarity for consequential decisions such as pay and promotion. ⚠️ This study did not involve AI in any form — no AI-written reviews, no AI disclosure, no human-versus-AI comparison. It tells you which format people find fairest. It says nothing about who or what should write that format, and it cannot be used to argue that AI drafting raises perceived fairness.

People Management cataloged the risk surface. People Management's "what HR needs to know" brief flags the audit, transparency, and consent issues: a model trained on historical ratings can carry the patterns in those ratings forward, and disclosure matters because employees respond differently when they cannot tell whether a human or a model wrote what they are reading. This is a practitioner overview rather than primary research — read it as a list of what to watch, not as measured effect sizes. JD Supra's legal-risk guidance makes the parallel point that AI-assisted performance decisions create disclosure and discrimination-claim exposure that most review policies don't yet cover.

Disagreement

The corpus disagrees in three predictable ways.

Vendors and consultants read the technology shift as inevitable and pro-productivity — the bulk of the 2026 AI-in-PM market-growth coverage falls here. Practitioner press (HR Executive, People Management) is cautious, focused on second-order effects on trust and bias. The companies themselves are running opposing experiments — Meta and Duolingo same quarter, opposite verdicts. Academic research (Cornell) speaks only to review format — it finds narrative feedback is perceived as fairest, while cautioning against dropping numerical ratings altogether. It does not address AI on either side of the question.

So the live disagreement isn't "AI yes or no." It's three sequential questions, and the evidence pattern is different for each:

  1. Should AI generate review content (drafts, summaries) — where it operates as a writing aid? (No study in this corpus tests this. Treat it as a controlled operational hypothesis: plausibly the lowest-risk placement, but it needs local testing and human verification of every output.)
  2. Should AI score employees against criteria — where it operates as evaluator? (No outcome evidence either way; the identified risks — opacity, historical patterns in training data, contestability — are unresolved.)
  3. Should AI usage itself be a performance signal — where the tool becomes the outcome? (One expert commentary against it, and one company that adopted the metric and dropped it about a year later. That is an argument plus a precedent, not measured evidence.)

Most vendor pitches blur questions 1, 2, and 3 into a single product story. The HR leader's job is to keep them separate.

Peoplense Verdict

This is a cautious decision made under limited evidence, not a proven verdict. No study in this corpus tests AI in performance reviews directly. What follows is our judgement about where the risk is lowest, and it should be treated as a hypothesis to test locally rather than a settled answer.

Do: consider a limited, disclosed pilot of AI for low-stakes drafting or summarisation — first-pass review text, peer-feedback summarisation, evidence aggregation for the manager. A human must check every output against documented evidence, and a human remains accountable for the final review and the decision. Disclose AI involvement as a matter of policy, not on request. We are not claiming this is safe — we are saying it is the lower-risk test case, and it needs verification on your own data before it becomes routine.

Don't: use AI as the final scorer. And don't tie AI usage to ratings or compensation without validated evidence that the metric reflects outcomes you actually value — the expert argument against it is strong, and the one company that tried it publicly abandoned it.

On bias — read the inference carefully. The Harvard Kennedy School study examined human self-ratings and manager anchoring. It did not study AI and does not show that AI models inherit bias. Our inference, labelled as such: if historical performance labels contain human bias, then training or prompting a model on those labels creates a risk of reproducing it — in a form that is harder to inspect and challenge. That is a reason to validate before deploying, not a demonstrated finding.

Watch out: disclosure is the part most likely to be skipped and hardest to repair. We have no measured figure for how undisclosed AI involvement affects trust, and we are not going to invent one — but the practitioner sources converge on disclosure and a contestation route as the baseline conditions, and both are cheap to put in place before a pilot starts.

What to do today

Three concrete actions for HR leaders this week.

  1. Audit your current AI-in-reviews exposure. List every tool in your review process that uses AI — Workday/SAP-built features, Lattice/15Five summaries, embedded ChatGPT, vendor-trained capability scoring. For each, identify which step it touches (draft / summary / score / recommendation). Most teams have more AI in the review process than they realize, often via vendor features quietly enabled in the last update.

  2. Write a one-page AI-in-Reviews disclosure policy. Three lines: what AI does, what the human does, how employees can see and contest the AI portion. Communicate it before the next review cycle opens. The practitioner sources agree that transparency and a contestation route are the baseline conditions, whatever the tool can do.

  3. Drop any AI-usage performance metric you have or are considering, unless you can validate it. This is the highest-leverage move on the page. The expert argument against it is strong and one company adopted it then abandoned it — an argument and a precedent, not proof. Measure outcomes — project impact, customer outcomes, peer trust — not tool usage. If you genuinely need to encourage adoption, do it with training time and tool access, not with ratings or compensation.

Data Privacy — The Legal Dimension

Any conversation about AI in performance reviews eventually hits one question that can't be sidestepped: what does the law say about employee data?

This isn't an ethical concern alone — data-protection regimes place real obligations on automated decision-making about people.

This section is an editorial summary, not legal advice. It describes what the official instruments say at a general level. Whether any of it applies to a particular system depends on that system's purpose, legal basis, design and use. Organisations should obtain qualified Saudi legal review before deploying automated employment-decision systems.

Globally, major data-protection frameworks — GDPR in the EU being the most consequential — set explicit rights for employees regarding automated decision-making. The most important among them:

  • GDPR Article 22 concerns decisions based solely on automated processing that produce legal or similarly significant effects — and it sets out exceptions and required safeguards rather than a flat prohibition.
  • Where it applies, the associated rights include meaningful information about the logic involved and the ability to contest the decision and obtain human review.

An important limit on that: it does not follow that every AI-assisted review falls under Article 22. A review that a manager writes, edits and owns — with the model contributing a draft — is not obviously a decision made solely by automated processing. Whether a given setup crosses that line is a legal question about your specific design, not something this brief can answer for you.

Practical steps that are sensible regardless of how that question resolves:

  • Identify a clear lawful basis for processing with advice, rather than assuming one.
  • Up-front disclosure in the employee privacy notice, before the tool goes live.
  • A documented contestation channel with a named human owner.
  • A Data Protection Impact Assessment (DPIA) before deployment, documenting risks and mitigations.

In Saudi Arabia, the Personal Data Protection Law and its implementing regulations include transparency requirements relating to decisions made solely through automated processing, and identify explicit consent in that context. Which requirements apply to a particular employment system depends on its purpose, legal basis, design and use — consent is identified in that context, and this brief does not assert it is the only possible legal basis.

Two local points worth flagging:

  • Cross-border data transfer is regulated, not categorically prohibited. Most major AI vendors host data outside the Kingdom, so the transfer regulations make this a legal-review question alongside the procurement one.
  • Performance information is not automatically sensitive data. Enhanced requirements may apply where the data actually includes sensitive categories — health information being the clearest example. Whether your review data does is a question about your own system.

Bottom line on this dimension:

Technical capability and legal permissibility are different tests, and the second one is specific to your design. Get it reviewed before deployment, not after.

GCC Relevance

The honest position first: no cited Gulf outcome study in our current corpus establishes whether AI-assisted performance reviews improve fairness or accuracy. We previously carried several confident claims here about regional mandates, cultural dynamics and inevitable adoption. They were not supported by sources we can point to, so we have removed them rather than relabel them as opinion.

What we can say is narrower, and it is about how to test rather than what to expect.

Treat Gulf-specific application as a controlled local test, not a rollout. Because there is no regional outcome evidence either way, the responsible path is a small, bounded pilot with four conditions in place before it starts:

  • Disclosure — employees are told what the model does and what the manager does, before the cycle opens.
  • Human accountability — a named person owns the final review and the decision, and cannot defer to the tool.
  • Validation on your own data — you check the outputs against documented evidence rather than assuming they are sound.
  • A contestation route — a real channel to challenge an output, with a human on the other end.

On drafting specifically: we are not claiming AI is safe for drafting. Drafting may be a lower-risk test case than scoring, because a human writes, edits and owns the result — but that is a hypothesis about relative risk, and it holds only where every output is verified.

Practical implication for Gulf people-development leaders: the useful question is not "should we adopt AI in reviews" but "what would we need to observe, on our own data, before we let a model near a consequential decision?" Answer that before the vendor conversation, not after.

Sources

All corpus articles below open in our admin reader with the editorial-summary contract banner — our text summary on Peoplense, full text and figures at the original publisher.

Get the Monday Brief

Evidence-based people development research, summarized weekly. Free. No ads. Every article links to its source.

Email used only to deliver the brief. Unsubscribe anytime.