Skip to content
Book a Call
Menu
hello@abscilem.com

We answer within one business day

Book a Call

Peer Review vs AI: Will a Machine Judge Your Papers by 2027?

What is already happening, what AI can and cannot assess, and what researchers need to know about the system that evaluates their work.

The question in this article's title is not a hypothetical. In a meaningful sense, it has already been answered — the answer is partly yes, the transition is underway, and the researchers who understand what is happening are better positioned to navigate it than those who assume the peer review system they trained under is still the one evaluating their submissions.

Pangram Labs analysed all 70,000 reviews submitted to ICLR 2025 and found that roughly 21% were fully AI-generated — not polished with AI, not outlined with AI, but entirely written by a large language model, start to finish, for one of the most important machine learning conferences in the world. NeurIPS, the premier AI conference, saw a 75% increase in research submissions in 2025 compared with 2023, swamping its human review system and leading to AI-generated research being accepted into the programme.

These figures come from AI and machine learning conferences, where the convergence of AI expertise in both the submissions and the review pool makes the numbers higher than in most natural science fields. But the directional trend is consistent across fields: reviewer fatigue is real, submission volumes are rising, the demand on human reviewers is outpacing the supply of qualified reviewers willing to provide timely, substantive feedback, and AI tools are filling the gap — sometimes with journal knowledge and sometimes without it.

By 2027, the question is not whether AI will be involved in peer review. It already is. The question is which parts of peer review AI will perform reliably, which parts it will perform poorly, and what that means for the researchers whose work is being assessed.

The Crisis That Made AI Peer Review Inevitable

Manuscript submissions to peer-reviewed journals have seen an unprecedented 6.1% annual growth since 2013, with a considerable increase in retraction rates. It is estimated that over 15 million hours are spent every year on reviewing manuscripts previously rejected and then resubmitted to other journals.

21% ICLR 2025 reviews fully AI-generated
75% NeurIPS submission growth 2023 → 2025
15M Hours wasted yearly reviewing resubmitted rejects
81.8% AI reviewer accuracy vs 83.9% human average

This is the structural context in which AI peer review tools have emerged: not as an experiment in algorithmic assessment but as a response to a system that is genuinely struggling under the volume it is being asked to process. The reviewer shortage is documented, persistent, and worsening. Fraud and misconduct are scaling, reviewers are fatigued, and AI poses new challenges, yet the community is also building tools, policies, and collaborations that could make peer review more robust, transparent, and inclusive than ever before.

The response from the publishing community has been to deploy AI at the stages of the review process where automation is most defensible — and to attempt to hold the line on the stages where human judgment remains irreplaceable. Understanding where that line falls, and how it is shifting, is the most practically important thing a researcher can know about the review process in 2026.

What AI Is Already Doing in Peer Review

The involvement of AI in peer review exists on a spectrum from unambiguously legitimate to deeply problematic, and different parts of that spectrum are being deployed at different journals and at different stages of the review workflow.

  • Automated pre-screening Widely deployed

    The least controversial application and the most widely deployed. AI-driven reviewer-matching algorithms analyse manuscript keywords and compare them against massive databases of reviewer expertise profiles, factoring in publication histories, institutional affiliations, and recent conference presentations, while ensuring diversity in reviewer selection across geographic spread, career stage, and gender balance. This application of AI does not replace human judgment — it supports editorial logistics, reducing the time editors spend searching for appropriate reviewers and improving the quality of reviewer-manuscript matching.

  • Statistical & image validation Advancing rapidly

    Tools like StatReviewer, which performs automated statistical checks on submitted manuscripts, and Proofig, which screens images for manipulation and duplication, are being integrated into editorial workflows at a growing number of journals. These tools make reviews faster, more consistent, and more rigorous, with AI performing statistical analysis of models, p-values, and power calculations while human reviewers assess logic, applicability, and potential overfitting. AI performs checks that are algorithmic and well-defined; human experts evaluate judgements that require contextual reasoning.

  • Full AI-generated review reports Contentious

    A talk at the 2025 Peer Review Congress discussed a "Fast Track" peer review offering from NEJM AI, in which decisions are issued within a week of submission, based solely on the editors' assessment of the manuscript and two AI-generated reviews. This represents a fundamental shift in the peer review model — from a process in which qualified domain experts provide substantive critique to a process in which algorithmic assessment provides structured feedback at high speed.

The accuracy of these AI-generated reviews has been studied with increasing rigour. Research validating an AI reviewer system on 1,963 ICLR 2025 submissions found that the AI achieved 81.8% accuracy for the accept/reject classification task, compared to 83.9% for the average human reviewer. AI-generated reviews were rated as higher quality than the human average by an LLM judge, though still trailing the strongest expert contributions.

That finding deserves careful interpretation. An AI system performing at 81.8% accuracy against human average accuracy of 83.9% is not a negligible difference — it represents a meaningful gap in the quality of evaluation at the margin. And the aggregate accuracy figure conceals an asymmetry in what AI assesses well and what it assesses poorly.

What AI Cannot Evaluate: The Irreducible Human Contribution

The analysis highlights domains where AI reviewers excel — fact-checking, literature coverage — and where they struggle — assessing methodological novelty and theoretical contributions — underscoring the continued need for human expertise.

This distinction is foundational for understanding what AI peer review means for researchers submitting work to journals.

AI evaluates reliably
  • Citation accuracy and existence
  • Statistical methodology and reporting
  • Compliance with reporting standards
  • Internal consistency of claims and methods
  • Literature coverage and completeness
  • Image manipulation and duplication
AI evaluates poorly
  • Whether a claim is genuinely novel given the field
  • Whether a theoretical contribution is meaningful
  • Methodological ingenuity beyond formal criteria
  • Contextual fit with the specific laboratory practice
  • Unpublished work and emerging community results
  • Groundbreaking implications of incremental data

In groundbreaking studies involving new concepts, AI may fail to perceive the minute details, the originality of a new perspective, or a new theory based on existing data, resulting in insufficient critical analysis. A reviewer with expertise in the field is more likely to be open-minded about game-changing ideas.

This limitation has a specific consequence for the kinds of research this series has most consistently advocated: genuinely novel work in niche fields, by researchers at institutions without established reputations, addressing questions that are not well-represented in the mainstream literature. This is precisely the category of research that AI review systems are worst positioned to evaluate — because the novelty assessment requires knowing what has not been published, the context assessment requires knowing the specific field from the inside, and the judgement about significance requires the kind of expert intuition that comes from years of engagement with a specific scientific community.

AI handles the verifiable. Humans handle the evaluative. The frontier between them is where novelty actually lives.

The Confidentiality Problem

There is a structural ethical problem with AI peer review that has received less attention than the accuracy question but is, in some respects, more immediately consequential. For an article to be evaluated by AI, it must first be uploaded to an AI application. This constitutes a significant ethical violation because it compromises the confidentiality of a manuscript submitted to a journal. The confidentiality of the text uploaded to an AI application cannot be guaranteed.

ICMJE suggests that reviewers should not upload manuscripts to software or AI technology platforms that cannot guarantee confidentiality. Science prohibits the use of large language models during peer review and prohibits reviewers from uploading manuscripts to generative AI tools. The Lancet maintains that reviewers should refrain from using generative AI to assist in the scientific review of papers. Nature Nanotechnology, representing the Nature Portfolio policy, asks reviewers to "not upload manuscripts into generative AI tools," noting that using AI to improve the grammar or readability of human-generated review texts does not need to be declared.

For commercial-sensitive research

The confidentiality concern is particularly acute for research with commercial implications. A manuscript describing a novel synthetic route, a new device architecture, or a clinical result that could form the basis of a patent application is at risk if a reviewer uploads it to a commercial AI system — because the terms of service of most commercial AI platforms do not guarantee that uploaded content will not be used for model training or retained in accessible form.

The researcher submitting this work to peer review has no visibility into what a reviewer does with the manuscript, no mechanism to detect a confidentiality breach, and no recourse once the breach has occurred. The 21% figure from ICLR is, in this respect, not just a statistic about review quality — it is a statistic about confidentiality exposure on a scale the system was never designed to handle.

What 2027 Looks Like: The Hybrid System Taking Shape

The trajectory of AI in peer review is not toward full automation. The evidence from the fields where AI review has been most extensively tested suggests that the destination is a hybrid system in which AI performs the verifiable, algorithmic components of review and human experts retain responsibility for the evaluative judgements that require domain expertise and contextual reasoning.

Looking toward 2030, a hybrid review system where AI and human expertise work in synergy — each complementing the other's strengths in a fully transparent and auditable process — is emerging as the most defensible model. AI would perform triage based on topic, novelty, and ethics at the screening stage, with humans overriding AI errors and ensuring contextual fit. At the review assignment stage, AI suggests reviewers while humans ensure diversity and avoid conflicts of interest. AI will increasingly support, but not replace, reviewers and editors — speeding processes while humans retain the role of decision-makers and guarantors of integrity.

For researchers submitting work in 2026 and 2027, this hybrid trajectory has specific practical implications. The components of a manuscript that AI review systems assess most reliably — citation accuracy, statistical methodology, compliance with reporting standards, internal consistency — are the components where preparation and verification before submission have always mattered most. A manuscript that arrives at peer review with correct references, sound statistics, and properly documented methods is not merely more likely to survive AI pre-screening; it is also more likely to receive the focused expert attention on the evaluative questions where human judgement genuinely matters.

The components that AI assesses least reliably — novelty, theoretical significance, methodological ingenuity, and the kind of expert-context judgement that experienced reviewers bring — are the components where the quality of the science itself is the primary determinant. No amount of AI-proofing a manuscript changes the underlying intellectual contribution. But understanding that AI systems will evaluate the verifiable components with increasing rigour, while human experts will focus their remaining attention on the evaluative ones, is useful information for any researcher preparing a submission.

What Researchers Can Do

The transformation of peer review by AI is not something individual researchers control. But it has specific implications for manuscript preparation, submission strategy, and expectations about the review process that are worth understanding clearly.

Practical Framework

Four practices for submitting in the AI-augmented review era

  1. Prepare for algorithmic screening as rigorously as for human review

    The pre-screening layer that AI now performs at many journals is not a lower bar than human review — in some respects it is a higher one, because it is applied consistently to every submission rather than depending on whether a particular reviewer reads carefully. A manuscript with a hallucinated citation, an incorrectly reported statistical test, or a figure that a computational image integrity tool flags as potentially manipulated will face an automated challenge at the screening stage that would previously have been caught only if a particularly careful reviewer noticed it.

  2. Do not rely on peer review to catch what you missed

    When official review quality becomes less reliable, the cost of arriving at submission with unresolved scope-fit mismatch, citation-gap exposure, or figure-trust erosion goes up. Authors increasingly assume the formal peer-review system will catch the strategic scientific issues they missed. That assumption is no longer strong enough. The responsibility for catching errors, assessing novelty honestly, and verifying every claim against primary sources rests with the authors.

  3. Understand your rights when confidentiality is at risk

    If you are submitting work with commercial implications — patent-pending research, results forming the basis of a spinout, data that has not yet been publicly disclosed — know the specific AI policies of the journal you are submitting to, and know that those policies govern what the journal does, not necessarily what an individual reviewer does. The gap between journal policy and reviewer practice, documented in the 21% figure from ICLR, is real and relevant.

  4. Treat AI-generated reviews with calibrated scepticism

    If a review you receive is unusually generic — strong on formal criteria, weak on domain-specific insight, free of the specific contextual references that an expert in your field would naturally make — you may be reading an AI-generated report. Responding requires addressing the formal points while communicating directly with the editor about substantive scientific questions a domain expert would have engaged with. The mechanisms for escalating this concern exist at most journals; using them, where appropriate, is not a challenge to the review system but a participation in maintaining its integrity.

The Bottom Line

The peer review system is under genuine strain, and the AI tools being deployed to relieve that strain are a mixed development — genuinely useful in some applications, genuinely problematic in others, and at a stage of development where the outcomes for submitted manuscripts are less predictable than they have been in any previous era of scientific publishing.

The researcher who understands this clearly is not disadvantaged. They are informed — which is, in a system under transformation, the most useful thing to be.

Frequently Asked

Questions readers ask after this piece

Are AI tools already being used in peer review?

Yes. Pangram Labs analysed all 70,000 reviews submitted to ICLR 2025 and found that roughly 21% were fully AI-generated. NeurIPS saw a 75% increase in submissions in 2025 versus 2023, with AI-generated research being accepted into the programme. These figures come from AI conferences where adoption is highest, but the directional trend is consistent across fields. By 2027, AI involvement in peer review is not a hypothesis — it is the current state of the system.

How accurate are AI-generated peer reviews?

Research validating AI reviewers on 1,963 ICLR 2025 submissions found 81.8% accuracy for the accept/reject classification task, compared to 83.9% for the average human reviewer — a measurable gap at the margin. AI reviews are rated as higher quality than the human average by LLM judges but still trail the strongest expert contributions. The aggregate accuracy figure conceals an asymmetry: AI performs well on verification tasks and poorly on novelty assessment, theoretical significance, and methodological ingenuity.

What can AI evaluate reliably in peer review, and what cannot it?

AI evaluates reliably the verifiable, algorithmic components: citation accuracy, statistical methodology, compliance with reporting standards, internal consistency, literature coverage, and image integrity. AI evaluates poorly the components requiring expert judgement: whether a claim is genuinely novel given the current state of the field, whether a theoretical contribution is a meaningful advance, whether an experimental design is adequate by the practical standards of a specific laboratory context, and whether a result has groundbreaking implications. Novelty assessment is the area of greatest AI weakness.

Is it ethical for a reviewer to upload my manuscript to an AI tool?

It is generally not. ICMJE, Science, The Lancet, and Nature Portfolio policies all prohibit reviewers from uploading manuscripts to generative AI tools that cannot guarantee confidentiality. The terms of service of most commercial AI platforms do not guarantee that uploaded content will not be used for model training or retained in accessible form. This is particularly acute for research with commercial implications — patent-pending work, spinout-relevant results, or undisclosed clinical data. The gap between journal policy and reviewer practice is real.

What will peer review look like by 2027?

A hybrid system in which AI performs the verifiable, algorithmic components of review and human experts retain responsibility for evaluative judgements. AI will handle triage based on topic, novelty, and ethics at the screening stage; suggest reviewers at the assignment stage; perform statistical validation and image integrity checks; and produce structured feedback on formal criteria. Human reviewers will focus on the evaluative judgements that require domain expertise: novelty, theoretical contribution, methodological ingenuity, and contextual fit. Full automation of peer review is not the trajectory.

What can a researcher do to prepare for AI-augmented peer review?

Four practices materially help: prepare for algorithmic pre-screening as rigorously as for human review — verify every citation, statistic, and figure before submission; do not rely on peer review to catch what you missed, since the system increasingly cannot; understand the journal's specific AI policy if your work has commercial implications; and treat unusually generic reviews with calibrated scepticism — communicate with the editor directly about substantive scientific questions a domain expert would have engaged with.

Tell us what you are working on

Thirty minutes, no pitch. We will tell you plainly whether we are the right people for it.

We answer within one business day