Performance reviews are one of the few workplace processes that leave a large written record, which makes them unusually amenable to analysis. Several studies have taken large samples of reviews and examined them systematically.

The findings are consistent enough across studies to be worth taking seriously, though it should be noted that most of this research comes from particular sectors — technology and professional services predominate — and generalisation beyond them is uncertain.

Finding one: personality versus work

The most frequently replicated finding is that critical feedback given to women contains more references to personality and interpersonal style, and less about specific technical or task performance, than critical feedback given to men.

Terms describing manner — abrasive, bossy, emotional, aggressive, difficult — appear substantially more often in reviews of women in the samples analysed.

The practical problem with personality feedback is that it is not actionable. "Be less abrasive" specifies no behaviour to change, provides no measure of success, and cannot be demonstrated to have been addressed. Feedback about a specific deliverable can be.

Finding two: specificity

Analyses have found that positive feedback given to women is more often vague — good year, great to work with, valuable contributor — while positive feedback given to men more often references specific achievements and technical accomplishments.

This matters at promotion time. Promotion cases are built from documented evidence, and a file of general praise is a weaker case than a file of specific accomplishments, regardless of the underlying performance.

The reviewer who wrote warm generalities did not intend to weaken the case. The case is weaker anyway.

Finding three: attribution of outcomes

Studies of how success is described find differences in attribution — successes described in terms of ability and insight versus successes described in terms of effort, helpfulness and team contribution.

The effort attribution is the same pattern found in classroom feedback research, appearing twenty years later in a different institution with the same consequence: it does not support a prediction about harder future work.

Finding four: the same behaviour, different words

Perhaps the most striking result comes from studies presenting evaluators with identical described behaviour attributed to different individuals.

The same assertive act is described using different vocabulary depending on the actor, with the vocabulary applied to women carrying more negative connotation.

This has an important implication: the difference is not necessarily in the assessment of effectiveness but in the language used to record it, and the record is what persists into the promotion file.

The reference letter parallel

The same pattern appears in academic and professional reference letters, where it has been studied extensively.

Letters for women contain more communal terms — helpful, caring, collaborative — and fewer standout terms — exceptional, outstanding, brilliant. Letters for women are on average shorter. Letters for women more often contain hedges and references to effort.

Studies presenting such letters to evaluators find the differences affect ratings. Since letters are among the most consequential documents in academic hiring, this is not a small matter.

What reduces it

Requiring evidence for every rating. A reviewer who must cite a specific instance for each claim produces more specific text, which addresses the specificity finding directly.

Separating behavioural feedback from outcome feedback, and requiring behavioural feedback to name an observable behaviour and a suggested alternative. "Be less abrasive" fails that test. "In the design review you interrupted three times; wait for a pause" passes it.

Calibration sessions where reviewers compare ratings across a group have shown some effect, though the evidence is mixed and poorly run calibration can amplify rather than reduce inconsistency.

Language checking tools that flag personality terms and vague praise exist and have been adopted by some organisations. Their evaluation is limited but the mechanism is plausible.

And simply showing reviewers an analysis of their own review text is reported to have substantial effect, in the same way the assessment-versus-blind-score comparison does in schools. Most reviewers are genuinely unaware of the patterns in what they write.

What to do if you receive this feedback

Ask for the behaviour. "Can you give me a specific instance and tell me what you would have preferred?" is a reasonable, professional request, and it is the question personality feedback cannot survive.

Frequently the answer reveals either a specific incident that can be addressed, or that no specific instance is available — which is itself information.

Keep your own record. A contemporaneous log of specific accomplishments, with dates and outcomes, provides the material a vague review does not, and it is the single most useful habit available to anyone whose promotion will depend on documented evidence.

And read the review as a document that will be read by people who were not there. That is what it is, and its usefulness to you depends far more on its specificity than on its warmth.