Skip to main content
Business Writing

Performance Review Bias: What the Research Shows and What to Do About It

Meta-analyses of performance review data consistently show that review outcomes correlate more strongly with reviewer characteristics than employee performance. The writing is a major reason.

BellerDocs · August 7, 2026 · 10 min read

Filed under Publish & Promote

← Back to Blog

A 1998 meta-analysis by Scullen, Mount, and Goff — published in the Journal of Applied Psychology and widely cited in organizational psychology research — analyzed performance rating data and found that approximately 62% of the variance in performance ratings was attributable to the rater, not to the ratee. Only about 21% of the variance reflected actual differences in performance. The remaining variance came from the interaction between rater and ratee and from random error.

These numbers have been reproduced and refined in subsequent research, and they are not uniformly accepted — other researchers have found higher performance-signal content in ratings under different conditions. But the direction of the finding is robust: performance reviews, as typically conducted, capture rater characteristics as much as or more than they capture employee performance. The review is a subjective document, and the subjectivity operates at a scale that most organizations do not appreciate when they rely on review documentation for promotion, compensation, and termination decisions.

What makes this particularly relevant to professional writing is that the bias in performance reviews is not primarily in the numerical ratings. It is primarily in the narrative text. Numerical ratings are visible, comparable, and subject to calibration. Narrative sections are read, filed, and occasionally reviewed by HR or legal counsel — at which point the writing quality and bias content of the narrative becomes consequential in ways that were not apparent when the manager wrote it.

Documented Bias Types in Performance Review Writing

The bias types that appear most consistently in performance review research are well-characterized. Understanding them helps managers avoid inadvertent bias in their own writing and helps HR professionals identify it in reviews they are reading.

The halo and horn effect. The halo effect — allowing an overall positive impression of an employee to elevate ratings across all dimensions — and its reverse, the horn effect, are among the most thoroughly documented biases in performance evaluation. In written narratives, the halo effect appears as a consistent positive framing where specific behavioral examples are selected from a full year to support a pre-formed positive impression, while performance instances that would complicate the picture are omitted or minimized. The horn effect produces the reverse: a consistent negative framing where examples are selected to support a negative impression rather than to represent performance comprehensively.

Recency bias. Performance reviews are typically written annually, but the events that are most accessible in memory are recent ones. Research on recall patterns in performance evaluation, including work by Bretz, Milkovich, and Read in the 1992 Journal of Management review, documents that performance events from the final months of a review period are significantly overrepresented in review narratives relative to their occurrence in the full year. In practice, this means that an employee who performed exceptionally through Q1-Q3 and had a difficult Q4 will receive a review that underrepresents their performance, while an employee with the reverse trajectory will receive a review that overrepresents it.

Idiosyncratic rater effect. The idiosyncratic rater effect, identified and named by Mount and colleagues, refers to the tendency of individual raters to apply their own performance standards, which may differ significantly from organizational standards or from other raters' standards. A manager who values assertive communication styles will rate assertive employees more favorably on "communication" regardless of whether assertiveness is actually the communication behavior the role requires. This effect explains why calibration processes — where managers discuss their ratings collectively — change the distribution of ratings significantly.

The behavioral specificity test: For each evaluative claim in a performance review narrative ("strong communicator," "takes initiative," "sometimes struggles with detail"), ask whether you could substitute a specific behavioral example that a third party could verify. "Presented the Q3 roadmap to the board of directors with minimal preparation and fielded technical questions without deferring to the engineering lead" is specific and verifiable. "Strong communicator" is an evaluative conclusion that may or may not be grounded in observable behavior. Reviews made of conclusions rather than behaviors are both less accurate and more legally vulnerable.

Why Narrative Sections Carry More Legal Risk Than Ratings

The EEOC's guidance on performance evaluation documentation — published under its authority to enforce Title VII of the Civil Rights Act, the Age Discrimination in Employment Act, and the Americans with Disabilities Act — emphasizes that employment decisions made on the basis of performance reviews can be challenged if the review process or documentation reflects discriminatory bias. The legal standard for a performance review used in a termination or demotion decision is that it must be based on legitimate, nondiscriminatory criteria applied consistently.

Numerical ratings are subject to statistical analysis in employment litigation: a plaintiff can retain an expert to analyze whether protected class members receive systematically different ratings than similarly situated non-members. This analysis is powerful but also produces mixed results — statistical differences in ratings require a large enough sample to be statistically significant, which is not always available at the individual manager level.

Narrative sections are analyzed differently and often more effectively in litigation. An employment attorney reviewing performance review files for a termination case is reading the narrative text for:

Behavioral Specificity as the Primary Bias Reduction Tool

The most effective writing intervention for performance review bias is behavioral specificity: writing review narratives that describe specific, observable behaviors rather than evaluative conclusions. This intervention is supported by substantial organizational psychology research as effective for both reducing bias and improving the accuracy of the review.

The behavioral anchoring approach — developed from the Behaviorally Anchored Rating Scale (BARS) methodology introduced by Smith and Kendall in the Journal of Applied Psychology in 1963 — ties evaluative language to specific behavioral descriptions at each rating level. A BARS-anchored review of "project management" does not simply rate an employee as "excellent" — it describes what an "excellent" project manager does that distinguishes them from a "satisfactory" or "needs improvement" project manager, in behavioral terms, and asks the rater to identify which behavioral description most accurately describes the employee's behavior.

Even in reviews that do not use a formal BARS approach, the writing principle is the same: replace evaluative conclusions with behavioral observations. The revision requires more specificity from the manager — they must recall actual behaviors rather than generating an overall impression — which is precisely what reduces bias. The cognitive work of identifying specific behavioral examples forces the manager to examine whether their overall impression is supported by observable events.

Rating Inflation and Its Organizational Consequences

Performance review research consistently documents rating inflation — the tendency for ratings to cluster at the high end of rating scales regardless of actual performance distribution. A study published in the Journal of Applied Psychology by Egan and colleagues found that in a large sample of corporate performance ratings, more than 70% of employees received ratings of "meets expectations" or higher in organizations with five-point rating scales, despite the logical expectation of a normal distribution.

Rating inflation is partly a product of relationship dynamics — managers who work closely with their direct reports tend to rate them favorably regardless of performance, because the cost of a low rating (a damaged relationship, a demotivated employee) is felt immediately while the benefit of accurate rating (better organizational decisions) is diffuse. It is also a product of the writing environment: managers who must justify low ratings in written narratives, but who are not required to justify high ratings with equivalent specificity, face asymmetric writing demands that favor inflation.

The organizational consequence of systematic rating inflation is that the performance review data becomes unusable for its primary purpose — distinguishing high performers from low performers for promotion, compensation, and succession decisions. When 70% of employees "meet expectations," the review data provides almost no differentiation. Organizations that have discovered this pattern often respond with forced distribution systems, which solve the differentiation problem while introducing different biases. The underlying issue — that review writing norms do not require specific behavioral evidence for positive ratings — remains unaddressed.

Writing Reviews That Survive Legal Challenge

EEOC guidance and published employment law practitioner guidance identify specific writing practices that reduce legal vulnerability in performance reviews:

Get Your Performance Review Writing Evaluated

Our performance review evaluation examines your review narratives and rating documentation for the bias patterns, vague evaluative language, and legal exposure risks that HR professionals and employment attorneys identify — before the reviews are finalized and filed.

Get your Performance Review Legal Compliance