Funding Agencies

NSF Merit Review Criteria: What Panels Actually Get Wrong

The two NSF merit review criteria have been in place for decades. Panels still systematically underweight one of them — and program officers know it.

July 1, 2026 · 12 min read

The National Science Foundation's merit review system rests on two criteria — Intellectual Merit and Broader Impacts — that are among the most clearly articulated evaluation standards in research funding. Every panelist receives explicit guidance on both. Every proposal is required to address both. Every review is supposed to evaluate both. And yet, cycle after cycle, the same asymmetry appears: Intellectual Merit receives substantive, detailed evaluation from most panelists, while Broader Impacts receives perfunctory treatment — a few sentences of generic commentary that check the procedural box without providing the evaluative substance that a program officer needs to differentiate between proposals.

This asymmetry is not a minor calibration issue. It is a structural failure in the evaluation system that produces funding decisions based on effectively one criterion instead of two. When three panelists each write 500 words on Intellectual Merit and 50 words on Broader Impacts, the panel discussion — which draws on the written reviews as its raw material — inevitably gravitates toward the criterion that was evaluated in depth. The program officer, faced with detailed assessments of merit and superficial assessments of impact, is left to either make the Broader Impacts judgment independently (which bypasses the panel process) or to treat the superficial assessments as if they contained more information than they do.

Neither outcome is what the system intends, and both are avoidable if the completeness of panel reviews is measured before the evaluative conversation begins.

The Broader Impacts Problem Is a Completeness Problem

The undertreatment of Broader Impacts is often attributed to a definitional problem — panelists do not know what "broader impacts" means, or the definition is too vague to evaluate rigorously. This diagnosis is partially correct but mostly misleading. The NSF has refined the Broader Impacts criterion repeatedly, provided examples, and published guidance documents. The issue is not that panelists lack access to a definition but that the evaluation infrastructure does not enforce engagement with the criterion at a depth commensurate with its stated importance.

A panelist who writes "The broader impacts are adequate and include outreach to underrepresented groups" has technically addressed the criterion. But the assessment contains no evaluative substance: it does not describe what the proposed outreach involves, whether the approach is realistic, whether the PI has a track record of successful outreach, or how the proposed activities connect to the scientific work. An editor reading this comment would not know whether the reviewer spent thirty seconds or thirty minutes evaluating the broader impacts — and the answer, in most cases, is closer to the former.

The Broader Impacts criterion is not underspecified. It is underenforced. Panelists know what it asks. They simply know, from experience, that superficial treatment carries no penalty.

The remedy is measurement. If every panel review is mapped against both criteria and scored for depth — using the thorough/adequate/superficial/missing scale described in our post on the completeness problem — the asymmetry becomes visible and addressable before the panel discussion occurs. A program officer who can see that Reviewer 2 scored "thorough" on Intellectual Merit and "superficial" on Broader Impacts can either request a revision to the review, assign more weight to the other reviewers' Broader Impacts assessments, or note the gap for the panel discussion.

Intellectual Merit: Where the Problems Are Subtler

The Broader Impacts asymmetry is the most visible failure in NSF panel review, but Intellectual Merit evaluation is not immune to quality problems — they are simply harder to see because the volume of text devoted to merit creates an illusion of thoroughness.

The most common failure mode in Intellectual Merit evaluation is selective depth: the reviewer engages deeply with one aspect of the merit criterion (usually the aspect closest to their own expertise) and treats the remaining aspects with generic approval or silence. A reviewer who specializes in computational methods may write three paragraphs evaluating the proposed algorithms and one sentence noting that "the theoretical motivation is sound" — a judgment that may or may not be well-founded, and that provides the program officer with no basis for evaluating it.

This selective depth is a variant of the expertise overreach problem: the reviewer's detailed treatment of one aspect is excellent, but their cursory treatment of other aspects creates a false impression of comprehensive evaluation. The review appears thorough because it is long and technically detailed, even though the technical detail concentrates on a subset of the evaluative space.

A second failure mode is criteria conflation. The NSF defines Intellectual Merit in terms of several elements: the potential to advance knowledge, the qualifications of the investigators, the soundness of the approach, the adequacy of resources, and the integration of research and education. These are distinct evaluative dimensions, and a proposal can be strong on some and weak on others. But panelists frequently collapse them into a single global impression — "this is an excellent research program" — that obscures which dimensions were evaluated favorably and which were not evaluated at all.

What Program Officers Should Look For

Program officers reviewing panel assessments before the discussion should check three things in each review. First, whether both criteria received treatment at a depth that permits an independent evaluative judgment — not just a score, but enough evaluative reasoning to understand why the score was assigned. Second, whether the reviewer's merit assessment covers the full set of merit elements or concentrates on a subset corresponding to the reviewer's expertise. And third, whether the scores across criteria are internally consistent with the textual assessment — a reviewer who writes three paragraphs of criticism and assigns a score of "Excellent" may be exhibiting leniency bias or may have written their text and score at different times without reconciling them.

The Panel Discussion Gap

Even well-written individual reviews can fail to produce a well-functioning panel if the discussion does not resolve the disagreements and gaps that the reviews contain. The panel discussion is the mechanism that is supposed to catch incomplete reviews, surface contradictions, and produce a consensus assessment that is more reliable than any individual review. In practice, panel discussions are time-constrained (typically 15–30 minutes per proposal in large panels), dominated by the most vocal panelists, and focused on the proposals where disagreement is most visible — which are not necessarily the proposals where quality problems are most consequential.

A proposal with three uniformly positive reviews sails through the discussion in minutes, regardless of whether those reviews are complete. The discussion mechanism only engages when there is apparent disagreement, and apparent disagreement is driven by score divergence, not by completeness divergence. Three incomplete reviews that happen to agree on the subset of criteria they evaluated will generate less discussion than two complete reviews with a genuine threshold disagreement on one criterion — even though the former represents a worse evaluative outcome than the latter.

Panel discussions are designed to resolve disagreement. They are not designed to detect absence. A proposal that received three enthusiastic but incomplete reviews will be discussed less, not more, than one with a genuine split — and the decision will rest on less information.

The implication is that completeness analysis should precede the panel discussion, not emerge from it. If the panel chair can see, before the discussion begins, that Proposal X has a completeness gap on Broader Impacts across all three reviewers, the discussion can be steered to address that gap explicitly. Without this information, the discussion follows the path of least resistance: the panelists who wrote the most detailed reviews drive the conversation, the underevaluated criteria receive even less attention than they did in the written reviews, and the program officer makes a decision that is, unknowingly, based on a partial evaluation.

Calibration Across Panels

A less discussed but equally important quality dimension in NSF review is consistency across panels within the same program. Different panels, evaluating proposals in the same competition under the same criteria, may apply those criteria at different thresholds — with the result that a proposal's fate depends partly on which panel it lands in rather than on its intrinsic merit.

This variation is not random. Panels are composed of different individuals with different disciplinary norms, different interpretive frameworks, and different baseline expectations. A panel drawn primarily from experimentalists may apply different standards for "soundness of the approach" than a panel drawn primarily from theorists, even if both are evaluating proposals in the same program. The variation is not a sign of incompetence — it is an inevitable consequence of composing evaluation bodies from diverse experts whose norms differ in ways the rubric does not resolve.

Measuring this variation requires cross-panel consistency analysis: comparing the distribution of scores, the depth of criterion-level evaluation, and the frequency of identified quality problems (incomplete reviews, cross-reviewer contradictions, bias flags) across panels within the same program. If Panel A consistently produces more thorough Broader Impacts evaluations than Panel B, the difference is attributable to panel composition and norms rather than to differences in proposal quality — and the program officer should weight the two panels' assessments accordingly.

This kind of cross-panel analysis is beyond what any individual program officer can perform manually, but it is tractable as an automated quality assurance workflow: feed each panel's reviews through completeness scoring and consistency analysis, compare the distributions across panels, and flag panels whose quality profiles diverge significantly from the program mean. The findings inform both the current funding cycle (by highlighting decisions that rest on lower-quality evaluations) and future panel composition (by identifying the reviewer characteristics associated with more complete and consistent evaluations).

Toward Measurable Review Standards

The NSF merit review system is, by design, one of the most structured evaluation frameworks in research funding. The criteria are explicit, the guidance is detailed, and the process includes built-in mechanisms (panel discussion, program officer oversight) for catching evaluation failures. What the system lacks is measurement — systematic, criterion-level assessment of whether the reviews produced by panelists actually satisfy the evaluative requirements that the criteria define.

Adding measurement does not change the criteria, replace the panel process, or diminish the role of expert judgment. It makes the quality of that expert judgment visible and verifiable, so that program officers can identify evaluation gaps before they become funding errors, panel chairs can steer discussions toward the dimensions that need attention, and the agency can track evaluation quality across panels and cycles as a institutional metric rather than a matter of individual impression.

The tools for this measurement exist. Completeness scoring, cross-reviewer consistency analysis, and bias screening can be applied to NSF panel reviews with the same rigor they bring to any structured evaluation context — the rubric is well-defined (Intellectual Merit and Broader Impacts, with specified sub-criteria), the reviews are text-based and criterion-referenced, and the decision workflow has a clear integration point (between review submission and panel discussion) where quality findings can inform rather than disrupt the process.

What remains is the institutional decision to measure. The evidence that panel reviews are frequently incomplete on Broader Impacts, selectively deep on Intellectual Merit, and inconsistent across reviewers on both criteria is neither new nor contested. What is new is the ability to detect these problems at the speed and scale the system requires, and to do so in time to improve the decision rather than merely to document the deficiency after the fact.

ReviewPanel reads your manuscript and reviewer comments and drafts a structured response →