How Science Actually Gets Funded: Peer Review, NSF Criteria, and What "Consensus" Really Means
Most students picture "science" happening entirely inside a lab. But before almost any research project can start, it has to survive a much quieter process: a panel of scientists deciding whether it's worth funding at all. I've spent a good part of my career inside that process — as a program officer managing study committees at the National Research Council, as a fellow inside the National Science Foundation itself, and as a project manager at the American Association for the Advancement of Science (AAAS) from 2018 to 2019, where I helped manage the review panels that decide which scientific proposals get funded, in the U.S. and internationally. I'm a high school teacher and community college professor now, but that direct experience inside the review process is what this reading draws on. This reading is my attempt to pull that hidden machinery into the light: how a review panel actually works, how two major federal agencies score proposals differently, a real statistical trick used to make reviewer scoring fairer, and what scientists actually mean — and don't mean — when they say something has reached "consensus."
What a review panel actually is
A review panel is a group of outside scientists, assembled by a funding agency, to evaluate proposals competing for a limited pool of money. Building one starts with avoiding a conflict of interest — you can't review a proposal from your own institution, a former student, or a close collaborator, so agencies spend real effort just making sure each panel is free of those entanglements before a single proposal is read. Reviewers typically read and score proposals independently first, writing their own critique without seeing anyone else's opinion. Then the panel meets — in person or virtually — and each proposal is discussed out loud, reviewer by reviewer, before the group lands on a final recommendation. That discussion step matters more than people expect: a proposal that one reviewer misunderstood, or that a specialist in the room can vouch for, regularly moves up or down once the panel actually talks it through.
How the National Science Foundation decides
The National Science Foundation (NSF) scores every proposal against exactly two criteria, weighted equally: Intellectual Merit — how much the project would advance knowledge and understanding, and how well-qualified and well-organized the proposed work is — and Broader Impacts — how much the project would benefit society beyond the immediate research finding, such as training students, reaching underrepresented groups, or improving public understanding of science. In practice, agency surveys have found that many reviewers and program staff still tend to weigh Intellectual Merit more heavily and treat Broader Impacts as an afterthought, which is exactly why NSF's governing board has recently pushed to clarify and reinforce that the two criteria are supposed to carry equal weight. I spent 2007–2008 inside NSF's Division of Research on Learning as an AAAS Science & Technology Policy Fellow, sitting in on exactly this kind of proposal review and negotiation, and later served as an NSF review panelist myself.
How the National Institutes of Health does it differently
The National Institutes of Health (NIH) historically scored proposals on five separate criteria — Significance, Investigator(s), Innovation, Approach, and Environment — each given its own 1-to-9 score, with 1 being the best. Reviewers then also assigned a separate Overall Impact score, which was never a simple average of the five criterion scores but the reviewer's holistic judgment informed by them. Since January 2025, NIH has been transitioning most research grant applications to a Simplified Peer Review Framework that regroups those same underlying questions into three broader factors: Importance of the Research, Rigor and Feasibility, and Expertise and Resources. Either way, final scores across a whole study section get converted into a percentile ranking, and NIH institutes fund proposals down to a cutoff called the payline — meaning a genuinely well-reviewed proposal can still go unfunded in a tight budget year simply because too many other proposals scored slightly better. Comparing NSF's two-criteria model to NIH's multi-factor model is a useful exercise on its own: they're both trying to answer "is this good science," but they've built two genuinely different scoring architectures to get there.
When reviewers don't grade the same way
Here's a problem every review panel eventually runs into: some reviewers are naturally tougher graders than others, and some are more generous, even when judging the same quality of work. If you just compare raw scores side by side, you're not really comparing the proposals — you're partly just comparing how harshly each reviewer grades. When I chaired an AAAS Fellowship selection committee (deciding which scientists would become AAAS Science & Technology Policy Fellows placed at federal agencies), we used a real statistical fix for this: converting every reviewer's raw scores into z-scores — how many standard deviations each score fell above or below that specific reviewer's own average. A reviewer whose average score is a strict 6 out of 10 and a reviewer whose average is a generous 8 out of 10 can both flag their true favorite candidate the same way: as the candidate furthest above their own personal average, not the candidate with the highest raw number. Once every score is converted to the same standardized scale, you can fairly compare evaluations from reviewers who never agreed on what a "7" even means.
What "consensus" actually means (and doesn't)
From 2009 to 2010, I worked as a Program Officer at the National Research Council's Board on Science Education, where my job included being the study director for a committee tasked with a specific "charge" — a formal question the committee was convened to answer. That experience taught me that "scientific consensus" is one of the most misunderstood phrases in public science communication. It doesn't mean every single scientist agrees, and it isn't a single moment where everyone suddenly changes their mind at once. In my view — shaped directly by sitting in that room — consensus in practice functions more like an overwhelming supermajority than a strict unanimous vote, and National Academies committees actually build this in explicitly: any committee member keeps the right to file a formal dissenting opinion if they disagree with the group's conclusion, and that dissent is published alongside the report rather than hidden. Reaching consensus is also not instant. It's a genuine process — gathering evidence, hearing from outside experts, debating interpretations, revising drafts, responding to critical peer review of the report itself — that can take months, and it's built specifically to survive strong, respectful disagreement inside the room rather than to erase it.
Does the best idea actually win?
Mostly, yes — that's the honest answer, and it's also the whole point of building these systems this way rather than letting funding decisions get made by a single person's judgment. But "mostly" is an important word. Review panels have well-documented tendencies to favor safer, more incremental research over riskier ideas, simply because incremental proposals are easier for a room of strangers to evaluate confidently in a single sitting. Reviewer-to-reviewer variability is real, which is exactly why mechanisms like z-score normalization and structured, multi-criterion scoring exist — they're not there because the system is broken, they're there because good systems are built assuming individual human judgment alone isn't reliable enough, and they correct for that on purpose. A rejected proposal is also frequently revised and resubmitted successfully later, once a reviewer's specific critique has been addressed — rejection at one panel is rarely the end of an idea's story.
- Peer review
- The evaluation of scientific work by independent experts in the same field, used both to decide funding and to decide what gets published.
- Conflict of interest
- Any personal or professional relationship that could bias a reviewer's judgment of a specific proposal, requiring that reviewer to be excluded from evaluating it.
- Intellectual Merit / Broader Impacts
- NSF's two equally weighted review criteria: the potential to advance knowledge, and the potential to benefit society.
- Percentile ranking / payline
- NIH's system of ranking scored proposals against each other and setting a funding cutoff line, so being "well-reviewed" doesn't always guarantee funding.
- Z-score
- A statistical measure of how many standard deviations a value falls above or below its group's mean; used here to make reviewers with different scoring habits comparable.
- Scientific consensus
- A strong, evidence-based supermajority agreement among experts, reached through a structured, often lengthy process that explicitly preserves the right to formal dissent.
Think about it
- Explain the difference between NSF's Intellectual Merit and Broader Impacts criteria, and why NSF has recently pushed to make sure both are actually weighted equally in practice.
- Describe one meaningful difference between how NSF and NIH structure their proposal scoring.
- Using your own words, explain why converting raw reviewer scores into z-scores can produce a fairer ranking than comparing raw scores directly.
- Explain why "scientific consensus" is better described as a strong supermajority reached through a structured process than as total unanimous agreement.
- Review panels are sometimes criticized for favoring safer, incremental research over risky new ideas. Why might that bias exist, and what tools from this reading are used to try to counteract reviewer bias generally?
Sources: National Science Foundation, "How We Make Funding Decisions: Merit Review" (nsf.gov); National Science Board, "NSF Merit Review for a Changing Landscape" report, February 2026; NIH Office of Extramural Research, "Simplified Peer Review Framework" and NIAID "Scoring & Summary Statements" guidance (grants.nih.gov, niaid.nih.gov); National Academies of Sciences, Engineering, and Medicine, consensus study process documentation (nationalacademies.org). Personal professional experience: J. Reid Schwebach, Program Officer, National Research Council Board on Science Education (2009–2010); AAAS Science & Technology Policy Fellow, National Science Foundation, Division of Research on Learning (2007–2008); NSF review panelist (2015); Chair, AAAS Health, Education and Human Services Fellowship Selection Committee (2016); Project Manager, AAAS Research Competitiveness Program (2018–2019); now a high school teacher and community college professor. DRAFT — verify current Virginia Science SOL alignment (if any) with the current Curriculum Framework before publishing.