The Library
PLOS · Talent Management

Disparities in ratings of internal and external applicants: A case for model-based inter-rater reliability

Peer-reviewedby Patrícia Martinková, Dan Goldhaber, Elena EroshevaOctober 5, 2018 23 min read
inter-rater reliability hiring bias internal vs external applicants teacher hiring mixed-effect models selection tools measurement error applicant screening

Editorial summary. This is our text summary of an article published by plos. Charts, figures, and the author’s full voice are at the original — read it there .

Licence. The original is licensed under CC BY 4.0. This page adapts it: the summary and analysis are ours, not the authors’.

Editorial verdict

Methodologically sound empirical study. The bias finding against external applicants is statistically robust across multiple model specifications, but the single-district sample and limited quality controls constrain generalizability — treat the IRR methodology as transferable, the magnitude of bias estimates as district-specific.

Executive summary

This study examines rating disparities between internal and external applicants for teaching positions in Spokane Public Schools (SPS), focusing on both bias in scores and inter-rater reliability (IRR). The authors argue that external applicants are systematically rated lower than internal applicants, and that this disadvantage cannot be fully explained by observable qualifications or subsequent teacher effectiveness measures. Using a dataset of 3,474 ratings across 1,090 applicants, 137 raters, and 54 schools collected between 2008–2013, the study employs mixed-effect models to decompose variance sources and introduces a model-based IRR estimation approach that accommodates hierarchical data structures and group-specific variance terms. Key findings show internal applicants averaged approximately 39 points on the 54-point rubric versus roughly 36 points for external applicants, with the gap persisting after controlling for experience, licensure scores, and teacher value-added estimates. IRR was significantly higher for internal applicants (0.51) than external applicants (0.42). The authors propose that familiarity effects, recommendation letter quality, and rater social proximity contribute to these disparities, and demonstrate that increasing the number of raters improves reliability, though not to the threshold of 0.70 for external applicants even with three raters.

researchRelevance: 7/10United States

Key insights

  • 1External applicants received significantly lower ratings across all nine subcomponents of a 54-point screening rubric, with the gap persisting even after controlling for teaching experience, state licensure scores, and estimated teacher value-added.
  • 2Inter-rater reliability was significantly lower for external applicants (IRR = 0.42) compared to internal applicants (IRR = 0.51), with the difference (0.09) confirmed as statistically significant, suggesting raters are less consistent when evaluating candidates they are less familiar with.
  • 3Increasing the number of raters improves reliability and predictive validity for both groups, but the commonly cited 0.70 reliability threshold remains unattainable for external applicants even with three raters, indicating a structural disadvantage in the rating process for outsiders.

Practical takeaways

  • The model-based IRR framework developed in this study — using mixed-effect models with group-specific variance terms — offers a more flexible and statistically powerful alternative to stratified IRR calculation, applicable to rating contexts beyond teacher hiring such as grant peer review, journal review, and university admissions.
  • Rating standard errors for a single rater on the 54-point scale exceed 5.0 points, meaning scores could vary by 10 points from a single rater alone, illustrating the measurement imprecision inherent in single-rater screening processes.

Frameworks mentioned

Generalizability Theory

A statistical framework used to decompose sources of measurement error in ratings and estimate reliability across various scoring designs, applied here to model IRR under different numbers of raters.

References

  1. Various (references 6–8: Wennerås 1997, Sandström 2008, Van den Besselaar 2012). Affiliation bias in grant proposal peer reviews.
  2. Prior SPS-linked publication (reference 17) (2017).Spokane Public Schools four-stage hiring process and selection tool predictive validity.
  3. Journal of the Royal Statistical Society (implied standard citation) (1995).Benjamini–Hochberg correction for multiple comparisons.

Source & Provenance

Verified
Publisher / Source

plos

Author

Patrícia Martinková, Dan Goldhaber, Elena Erosheva

Publication Date

October 5, 2018

Article Type

Research Study

Geography

United States

Content Type
Unknown Source Type
Original Source

Original source metadata is preserved. AI analysis is generated separately.

Like this? Get the Monday Decision Brief — free, every week.

No spam, unsubscribe anytime.

Rate this article

Want the full article? Read it at the original source — free, no paywall.

Read original article
All content belongs to original publishers. AI analysis is for research purposes only. View original source

Where Peoplense used this source

This source is cited in the Decision Brief below — our own answer to the question it bears on, with the evidence weighed and a verdict.