The conversation about bias in peer review tends to gravitate toward the conspicuous cases — the reviewer who torpedoes a competitor's paper, the panelist with an undisclosed conflict of interest, the editorial board that systematically underrates work from certain institutions or demographics. These cases are real and important, but they are also relatively rare and, when identified, relatively easy to act on. You remove the reviewer. You flag the conflict. You investigate. The machinery for handling overt misconduct, however creaky, exists.
The far more consequential problem is the bias that does not look like bias. It lives in the cognitive shortcuts that every reviewer applies when translating a complex manuscript into a set of evaluative judgments under time pressure. These shortcuts — anchoring on first impressions, letting a strong methods section inflate the assessment of a weak discussion, applying stricter standards to unfamiliar subfields — operate below the threshold of the reviewer's awareness. They produce reviews that are internally coherent, confidently written, and systematically skewed in ways that neither the reviewer nor the editor can easily detect by reading the text once through.
This post catalogues seven bias types that appear with regularity in peer review contexts, explains the mechanism by which each one distorts evaluation, and describes what textual signatures editors and program officers can look for to identify them before they corrupt a decision.
1. Anchoring Bias
Anchoring is the tendency to over-weight the first piece of information encountered when forming a judgment. In peer review, the anchor is typically the abstract, the author's institutional affiliation (in non-blinded review), or the first section the reviewer reads in depth. Once an anchor is set, subsequent evaluation is pulled toward it — positive anchors produce more generous readings of ambiguous evidence, and negative anchors produce harsher ones.
The mechanism is well-documented in the decision-making literature and has been demonstrated empirically in grant panel settings. When panelists discuss proposals in sequence, the first few scores voiced in the room exert a measurable pull on subsequent discussion, not because the later panelists are deferential but because the initial framing shifts what counts as the reference point for "good" or "bad" within that discussion. The same dynamic operates within a single reviewer's reading process: if the abstract sets an expectation of a strong paper, the reviewer is more likely to interpret borderline methodological choices as "reasonable given the constraints" rather than "inadequately justified."
The textual signature of anchoring is evaluative momentum — a review whose tone and conclusions are set in the first paragraph and never substantively revised by anything that follows.
What to look for: reviews where the overall recommendation is predictable from the first two sentences of the report, where later sections receive progressively less critical attention, or where the reviewer references the abstract or introduction as justification for assessments of sections they appear not to have read closely. A review that quotes the abstract's claims about significance and then evaluates the results section by restating those claims rather than engaging with the data is anchored on the framing rather than the evidence.
2. Confirmation Bias
Where anchoring sets the initial reference point, confirmation bias sustains it. A reviewer who has formed a positive (or negative) first impression will — unconsciously but systematically — seek evidence that confirms that impression and discount evidence that contradicts it. The result is a review that feels comprehensive but is actually selective: the reviewer has read the whole paper, but they have metabolized only the parts that support the conclusion they had already reached.
Confirmation bias is particularly difficult to detect because it does not produce obvious errors. The reviewer's observations are typically accurate in isolation — they are just not representative. A reviewer who has decided the paper is strong will cite the three strongest tables and ignore the one that shows a null result. A reviewer who has decided the paper is weak will fixate on the limitations section and treat the positive results as incidental. Both reviewers have read the paper. Neither has reviewed it impartially.
What to look for: asymmetric treatment of evidence. If a reviewer discusses strengths at length with specific citations to the text but dismisses weaknesses (or vice versa) with generic language and no textual support, the asymmetry may reflect confirmation bias rather than a balanced reading. The ratio of specific-to-generic observations across the positive and negative dimensions of the review is a useful diagnostic. A balanced review should be roughly equally specific in both directions.
3. Halo and Horn Effects
The halo effect occurs when a strong impression on one evaluative dimension elevates the assessment of other, unrelated dimensions. The horn effect is the inverse: a weak impression on one dimension depresses the rest. In peer review, the most common halo triggers are a well-written introduction, a prestigious institutional affiliation, and a sophisticated statistical methodology. The most common horn triggers are poor English, a methodology the reviewer considers outdated, and a topic the reviewer regards as unfashionable.
The structural problem with halo and horn effects in review contexts is that evaluation rubrics typically treat criteria as independent: originality is scored separately from rigor, which is scored separately from significance. But the reviewer's cognitive process does not respect these boundaries. A reviewer who is impressed by the novelty of the research question will — often without realizing it — rate the rigor more favorably than they would have in a paper with a less interesting question. The criteria bleed into each other through the reviewer's overall impression, producing correlated scores that should, under an independent evaluation, show more variance.
The halo effect does not make a reviewer dishonest. It makes them consistent in a way that the rubric did not ask for — and that consistency is itself the distortion.
What to look for: scores that are suspiciously uniform across criteria (all high or all low, with minimal spread), combined with review text that is disproportionately focused on one dimension. A review that devotes 60% of its text to praising the novelty of the approach and then assigns high marks across rigor, clarity, and significance — with minimal substantive comment on those latter criteria — may be exhibiting a halo effect from the novelty dimension.
4. Leniency and Severity Bias
Some reviewers are consistently generous. Others are consistently harsh. Neither pattern is inherently problematic if the relative ordering of their evaluations is preserved — a lenient reviewer who ranks papers in the same order as a severe reviewer is contributing the same informational content, just on a different scale. The problem arises when leniency or severity bias is not uniform: a reviewer who is lenient toward work in their subfield and severe toward work outside it, or lenient toward established authors and severe toward junior ones, is not applying a consistent standard. They are applying two standards selectively.
Detecting leniency and severity bias at the individual-review level is difficult without a baseline. Within a panel context, it becomes more tractable: if Reviewer A's scores are systematically higher than Reviewers B and C across all proposals, the pattern is likely a calibration difference. If Reviewer A's scores are higher on some proposals and lower on others, with no consistent criterion-level explanation, the pattern may reflect conditional leniency or severity rather than a stable baseline offset.
What to look for: score distributions that deviate significantly from the panel mean without corresponding deviations in the specificity or substance of the review text. A reviewer who assigns the highest scores on the panel but writes the least detailed reviews is likely exhibiting leniency bias — high scores as a path of least resistance rather than a reflection of evaluative judgment.
5. Scope Mismatch
Scope mismatch occurs when a reviewer evaluates a paper or proposal for what they believe it should have done rather than what it set out to do. This is one of the most frustrating bias types from the author's perspective and one of the most consequential from the editor's perspective, because it can produce a technically detailed, superficially competent review that is fundamentally misdirected.
A common manifestation is the reviewer who criticizes a qualitative study for not using quantitative methods, or who faults a paper that explicitly claims to provide a theoretical framework for not including empirical validation. These are not minor quibbles — they represent a fundamental misalignment between the reviewer's evaluative expectations and the paper's stated scope. The reviewer is not wrong that empirical validation would be valuable; they are wrong that the absence of empirical validation is a flaw in a paper that never claimed to provide it.
What to look for: criticisms that address missing content rather than present content, especially when the "missing" content falls outside the paper's stated objectives. Language such as "the authors should have also examined" or "a major limitation is the absence of X" — where X is not a methodological flaw but an entirely different study — is the characteristic signature of scope mismatch. The editorial response should be to evaluate whether the reviewer's expectations are reasonable given the paper's framing, not to automatically weight the criticism against the authors.
6. Expertise Overreach
Expertise overreach is the complement of scope mismatch: rather than evaluating the wrong paper, the reviewer evaluates the right paper but renders judgments on aspects they are not qualified to assess. A domain expert in computational fluid dynamics who evaluates the statistical analysis of a paper using methods they have never employed is overreaching. A biologist reviewing a computational biology paper who critiques the biological framing (their strength) and the algorithmic implementation (not their strength) with equal confidence is overreaching on the latter.
The danger of expertise overreach is that it is invisible in the review text. The reviewer does not flag their own limitations — indeed, they may not be aware of them. The overreach manifests as technically flavored criticism that is either vague enough to sound plausible without being specific enough to be actionable, or specific but incorrect in ways that only a true expert in the evaluated dimension would recognize. Either form can mislead an editor who trusts the reviewer's overall domain credentials as evidence of competence across all dimensions of the review.
What to look for: a disproportionate drop in specificity or accuracy on certain criteria compared to others. A reviewer who writes precise, well-supported assessments of the experimental design but lapses into generalities ("the theoretical framework could be stronger") when addressing other dimensions may be operating outside their zone of competence on those dimensions. Cross-referencing the specificity of comments per criterion with the reviewer's known expertise profile (where available) can flag likely overreach.
7. Language Bias
Language bias is the systematic conflation of English proficiency with intellectual quality. A reviewer who notes that "the writing needs significant improvement" and then assigns low scores for clarity and rigor and significance is exhibiting language bias if the rigor and significance scores are depressed by the writing quality rather than by actual deficiencies in the research. The writing may indeed need improvement. The methodology may be entirely sound. Language bias collapses these independent assessments into a single negative judgment.
This bias disproportionately affects researchers from non-English-speaking countries and institutions, and it operates in tension with the stated goals of virtually every major journal and funding agency to evaluate work on its scientific merits rather than its presentational polish. The corrective is not to ignore writing quality — clear communication is a legitimate evaluative criterion — but to ensure that writing quality scores do not contaminate scores on substantive criteria that are independent of how well the paper is written.
What to look for: reviews that reference language quality in the general assessment (rather than in a dedicated "presentation" or "clarity" section), reviews where the severity of writing-related comments is inconsistent with the severity of substantive comments, and reviews where the overall recommendation is more negative than the sum of the substantive assessments would predict — with language criticism making up the difference.
From Detection to Action
Identifying bias in a review is necessary but not sufficient. The editorial or programmatic response depends on the type and severity of the detected bias, the decision context, and whether the bias is likely to have affected the outcome.
For minor or ambiguous cases — a possible anchoring pattern, a slight asymmetry in evidence treatment — the appropriate response is to note the flag and weight the review accordingly in the decision. Flagging does not mean discarding. A review that shows mild leniency bias may still contain valuable substantive observations. The bias flag alerts the editor to discount the overall recommendation while retaining the specific comments.
For more severe cases — clear expertise overreach on a criterion that is central to the decision, or a scope mismatch that renders the entire review non-responsive to the paper's actual contribution — the appropriate response may be to solicit an additional review, or to assign reduced weight to the flagged reviewer's assessment on the affected criteria while retaining their assessment on criteria where no bias was detected.
ReviewPanel.ai screens for all seven bias types using multi-model agreement: a bias flag is raised only when two or more AI models independently identify the same pattern in the same review text, with direct textual evidence cited for each flag. This design prioritizes precision over recall — the system flags fewer biases than a single model would, but the flags it does raise are corroborated and grounded in specific passages, giving the editor a concrete basis for evaluating the concern rather than a vague algorithmic suspicion.
The goal of bias detection is not to produce a verdict. It is to produce a question that the decision-maker can investigate with the benefit of specific textual evidence rather than a vague sense that something felt off.
The seven biases outlined here are not exhaustive, but they cover the patterns that recur most frequently in review contexts and that have the greatest potential to distort editorial and funding decisions. Measuring them systematically does not eliminate bias from peer review — it makes the bias visible, which is the prerequisite for managing it.