Ask a room of fifteen-year-olds to rate how good they are at maths, then compare those ratings against their actual test scores. The pattern that emerges has been replicated across dozens of countries and several decades, and it is consistent enough to be considered established.
Girls with identical scores to boys rate themselves lower. Not marginally lower in a way that could be noise — measurably, repeatedly lower.
This finding is usually reported as "the confidence gap," and it is usually followed by advice aimed at girls. That sequence contains a mistake worth unpicking, because the finding is real and the advice that follows from it mostly is not.
What is actually being measured
Self-assessment is a comparison. When a student rates herself as "quite good" at something, she is placing herself against a reference point — her class, her year group, her idea of what competence looks like.
The research finding is not that girls are less able. Attainment data in most Western school systems shows girls matching or exceeding boys in most subjects, including mathematics in many jurisdictions. The finding is that girls place themselves lower on the scale than their performance warrants.
The obvious reading is that girls are underconfident. The less obvious reading, which the calibration research supports at least as well, is that boys are overconfident. When you measure how accurately students predict their own scores, boys' predictions tend to overshoot. Girls' predictions tend to be closer to the mark, or to undershoot slightly.
Accuracy is not the same as confidence. A student who correctly estimates she will score 72 per cent is well calibrated. A student who predicts 85 and scores 72 is not more capable — he is worse at self-assessment.
Why the distinction matters
If the problem is framed as girls lacking confidence, the remedy is aimed at girls: workshops, affirmations, encouragement to speak up. If the problem is framed as a calibration difference in which one group systematically overestimates, the remedy points somewhere else entirely — at the systems that reward stated confidence over demonstrated accuracy.
And those systems are everywhere. Job applications ask candidates to rate their own proficiency. Interviews reward fluent self-description. Promotion processes in many organisations depend on someone putting themselves forward. Each of these converts a self-assessment difference directly into an outcome difference, regardless of underlying ability.
This is the mechanism that matters. The gap in self-rating would be a curiosity if nothing depended on it. It is consequential because so much does.
Where the gap appears and where it does not
The pattern is not uniform across every domain, which is itself informative.
Self-assessment gaps tend to be largest in domains stereotyped as male — mathematics, physics, computing, competitive leadership. They shrink or reverse in domains stereotyped as female, such as language and literature, where girls' self-ratings are often closer to or above their measured performance.
That domain-dependence is difficult to reconcile with a general personality explanation. If girls simply had a lower disposition to self-promote, the gap would be roughly constant across subjects. It is not. It tracks the stereotype associated with the field.
Studies that manipulate framing find similar effects. When a task is described as a test of an ability associated with men, self-ratings and sometimes performance shift; when the same task is described neutrally, the gap narrows. The effect sizes in this literature vary and some early findings have proved difficult to replicate at their original magnitude, but the direction of the effect has held up better than its size.
The age at which it appears
Self-assessment gaps are not present in the earliest school years in most measurements. They emerge during primary school and widen through adolescence.
That timing is a useful piece of evidence. Something that develops over a decade of schooling is being learned, not brought in at the start. Whatever produces it is available for examination — classroom practice, feedback style, peer comparison, media, family expectation — rather than being a fixed property of the students.
What follows from this
The practical implications run in two directions, and only one of them involves girls.
For institutions: any process that relies on self-report to allocate opportunity will systematically favour the group that over-reports. Structured assessment, work samples and blind evaluation reduce that effect. This is not a controversial claim; it is the basis on which orchestras adopted screened auditions.
For individuals: the useful advice is not "be more confident," which asks a person to become less accurate about herself. It is closer to "understand that the scale is not calibrated the same way for everyone, and that under-claiming is a real cost in systems that read claims as evidence."
That is a less inspiring sentence than most confidence advice. It has the advantage of describing what is actually happening.
What the research does not settle
It does not settle causation. The correlation between self-assessment and later choices is robust; the direction of the arrow is harder to establish, and both directions are plausible.
It does not settle the size of the downstream effect. Self-rating differences predict subject choice and application behaviour, but so do a dozen other things, and disentangling them is genuinely difficult.
And it does not support the strongest versions of the claim that appear in popular writing — that closing the confidence gap would close the pay gap, for instance. The pay gap has structural components that no amount of self-belief touches.
What the research does support is narrower and more useful: girls' self-assessments are better calibrated to reality, they are read as weaker claims by systems that reward strong claims, and the fix belongs at least as much in the reading as in the claiming.