The famous version of this finding is from 1968: teachers told that certain randomly selected students were expected to bloom intellectually, and those students subsequently showing larger gains.
That study has been criticised heavily on methodological grounds, its effect sizes have not replicated at the original magnitude, and it should not be cited as settled. The broad phenomenon, however, has survived much better than that particular study.
What the modern evidence supports
Meta-analyses of expectancy effects find them real but modest — typically small effect sizes, meaningfully different from zero and considerably smaller than the popular account implies.
Two refinements matter more than the headline.
First, effects are larger where expectations are inaccurate to begin with. Teachers' expectations of most students are reasonably well calibrated to actual performance, and accurate expectations cannot produce a distortion. The effect concentrates in the cases where the initial judgement was wrong.
Second, effects are larger for students from stigmatised or lower-status groups. This is the finding with the most social significance: the students most likely to be misjudged are also the students most affected by being misjudged.
The measurement problem, and how it was solved
The strongest recent evidence comes from a specific design: comparing a teacher's assessment of a student against that student's score on an externally, blindly marked examination.
Both measures are available in many national systems. The difference between them is a direct estimate of assessment bias, and because the external mark is blind, differences by student group cannot be attributed to differing performance.
Studies using this design have found systematic patterns. In several countries, teacher assessments in mathematics have been found to favour boys relative to their blind scores, while assessments in language subjects have been found to favour girls. Findings differ by country and by cohort, and the effects are generally small, but the pattern of subject-specific divergence has been replicated enough to be taken seriously.
Similar designs have found assessment differences by ethnicity and by socioeconomic status, generally larger than the gender effects.
Why small effects still matter here
A small bias applied once is negligible. The same bias applied repeatedly at decision points is not.
Teacher assessment feeds into setting and streaming decisions, into predicted grades used for university applications, into recommendations for advanced courses, and into the informal encouragement that determines whether a student is told she should consider a subject.
Each of those is a gate. A small nudge at a gate produces a different path, and paths compound. This is the mechanism by which effects too small to matter in a single classroom become visible in national enrolment statistics.
There is a documented feedback loop as well: streaming decisions determine curriculum exposure, curriculum exposure determines subsequent attainment, and subsequent attainment appears to validate the original decision.
What the mechanism appears to be
The evidence does not support a picture of conscious differential treatment. Teachers who show measurable assessment differences generally do not report or endorse differential expectations, and the effects appear in populations with strong stated commitments to equity.
The better-supported account involves how ambiguity is resolved. Where performance is unambiguous — a correct answer, a completed proof — there is little room for expectation to operate. Where it is ambiguous — partially correct work, a borderline essay, a judgement about potential rather than achievement — prior expectation supplies the missing information.
This predicts that bias should concentrate in subjective assessment and near decision boundaries, and it does.
What reduces it
Blind marking is the most direct intervention and has the clearest evidence. It is also cheap, at least for written work, and its adoption in higher education produced measurable shifts in outcome distributions in several documented cases.
Rubrics with specified criteria reduce the ambiguity that expectation fills. The effect is modest but consistent.
Feeding assessment-versus-blind-score comparisons back to departments has been used successfully in several systems. Most teachers are genuinely surprised by their own data, which is consistent with the mechanism being non-conscious.
Separating potential judgements from performance judgements helps. "Is she capable of the higher tier?" invites expectation; "what did this piece of work score against these criteria?" invites less.
The practical version for students and parents
Where a decision is being made on teacher judgement rather than on a blind score, and the decision is consequential — set placement, predicted grades, entry to an advanced course — it is reasonable to ask what the judgement was based on and how it compares to externally marked results.
That is not an accusation. It is a request for the evidence, and it is exactly the kind of question that reduces the space in which unexamined expectation operates.