Wednesday, October 7, 2026

If Authors Can Use AI, Why Can’t Referees?

 

The Reviewer Bottleneck  

Peer review has a very basic problem that predates artificial intelligence.

Reviewing a scientific paper properly is difficult work. It is also time-consuming work. A serious referee is expected to understand the central argument of a manuscript, check its important derivations, identify hidden assumptions, compare its claims with the existing literature, evaluate its novelty and significance, distinguish genuine errors from matters of presentation or interpretation, and determine whether the conclusions actually follow from the results. That is not a small task.

For a substantial theoretical paper, a careful review can consume many hours and sometimes several days of concentrated work. And who is supposed to do this work?

The same researchers who are already trying to write their own papers, conduct research, supervise students, teach courses, attend meetings, write reports and proposals, perform administrative duties, and maintain the rest of their academic lives.

Peer review depends overwhelmingly on scientists volunteering additional intellectual labor on top of everything else they already do.

It is therefore hardly surprising that journals increasingly struggle to find referees. Editors may have to invite many researchers before two agree to review a manuscript. Reviews can take weeks or months to arrive. Authors wait while editors search for another referee after several invitations have been declined. Sometimes the scientific evaluation itself takes much less time than simply finding someone willing and able to perform it.

This is not primarily a problem of bad referees. It is a problem of a system demanding substantial amounts of difficult intellectual work from people who often simply do not have enough time to perform it properly.

And this is precisely where artificial intelligence should enter the discussion.

AI Is More Urgently Needed by Reviewers Than by Authors

There is a curious asymmetry emerging in scientific publishing.

Authors are increasingly being given considerable latitude to use AI in their work, provided that they remain responsible for the final scientific content. They may use AI to improve language, organize arguments, summarize material, assist with coding, explore calculations, search literature, or interrogate complicated technical questions.

That development is understandable. But if AI assistance is useful to an author, it is arguably even more urgently useful to a reviewer.

The author is working on his or her own research. The author normally knows the calculation, the literature and the conceptual structure of the paper intimately. Months or years may have gone into producing it.

The referee enters from outside. The referee must understand somebody else's work, often under severe time constraints, reconstruct its logic, identify possible weaknesses, place it in the literature, and form a defensible scientific judgment.

This is exactly the kind of task for which intellectual assistance is valuable. Yet we are developing a publishing culture in which substantial AI assistance to the author is increasingly accepted while substantial AI assistance to the referee is often prohibited.

The priorities seem almost reversed.

If there is one part of the publication process where reducing the intellectual burden could immediately help the entire scientific system, it is peer review.

The Reviewer Does Not Need to Be Replaced

The issue is often framed incorrectly.

The choice is not

\[
\text{human reviewer}
\]

versus

\[
\text{AI reviewer}.
\]

The obvious model is

\[
\boxed{\text{human reviewer}+\text{AI assistance}.}
\]

The human referee remains responsible for the report.

That is exactly the principle already being adopted for authors. An author can use sophisticated computational tools without transferring authorship to those tools. The author remains responsible for the calculations and claims contained in the paper.

The same principle should apply to the referee.

Suppose I receive a theoretical-physics manuscript. Why should I not be allowed to give it to an AI system and ask:

“Reconstruct the central argument of this paper. Identify its principal assumptions. Check the important derivations where possible. Find possible inconsistencies. Compare the novelty claims with the cited literature. Separate potentially serious problems from minor ones.”

That does not replace me as referee. It gives me an extremely useful first analysis of the manuscript. I can then interrogate that analysis.

“Your objection to equation 27 seems incorrect because the large-(N) limit is taken before the infrared limit. Reanalyse it.”

“Derive equation 42 independently.”

“Where exactly is this assumption introduced?”

“Does reference 18 really establish what the authors claim?”

“Search for earlier work containing an equivalent result.”

“Construct the strongest argument against the main conclusion.”

“Now construct the strongest defense of the authors against those objections.”

At the end of this process, I decide what is correct. I decide what matters. I write the report. And I take responsibility for it.

This Could Make Reviews Better, Not Merely Faster

The obvious advantage is efficiency. A reviewer could perform in a few hours an analysis that might otherwise consume an entire weekend. Given the shortage of referees, that alone is important.

But speed is not the only issue. AI could also make reviews more thorough. A human referee may overlook an equation. A human referee may forget a relevant paper. A human referee may fail to notice that an assumption made on page 6 is incompatible with a conclusion on page 24. A human referee may simply become tired.

A sufficiently capable AI system can repeatedly traverse the manuscript, compare distant sections, track notation, examine limiting cases, summarize related literature and flag possible logical discontinuities.

Of course it can also make mistakes. Sometimes very convincing mistakes. But human referees make mistakes as well.

The correct response to this fact is not to prohibit assistance. It is to require the scientist to verify the criticisms that finally enter the referee report.

This is the same epistemic responsibility that scientists already exercise when using every other sophisticated tool.

We Do Not Demand That Referees Work Without Tools

Consider how strange such a principle would appear elsewhere. If an author evaluates a difficult integral using Mathematica, we do not insist that the referee reproduce the calculation by hand. If an author searches INSPIRE or Google Scholar, we do not require the referee to rely exclusively on papers remembered from years of reading. If an author performs a numerical computation on a powerful computer, nobody argues that the referee must reproduce it without computational assistance.

Scientific work has always advanced partly by developing better intellectual tools. AI is another such tool, although an unusually powerful one.

The relevant distinction should therefore not be

\[
\text{human work}\quad\text{versus}\quad\text{AI-assisted work}.
\]

It should be

\[
\text{responsible work}\quad\text{versus}\quad\text{irresponsible work}.
\]

Confidentiality Is a Solvable Problem

The strongest practical objection is confidentiality. An unpublished manuscript has been entrusted to the journal and to the referee. Clearly, a reviewer should not be free to distribute it indiscriminately to external parties.

But this is a technical and contractual problem. It is not a fundamental argument against AI-assisted reviewing.

Publishers could provide their own secure AI systems. A journal submission platform could make an AI assistant available directly to the referee. The manuscript would remain inside the publisher's protected infrastructure. The system could operate under explicit conditions that the manuscript is not used for training, is not retained beyond the required period, and is not disclosed externally. Universities could provide equivalent secure systems.

If confidentiality is genuinely the problem, then the rational response is to build confidential AI tools for reviewers. It is not to prevent reviewers from benefiting from AI.

Journals Should Give Every Referee an AI Assistant

This seems to me the natural direction in which peer review should develop. When a researcher accepts a review invitation, the journal system should provide an integrated AI assistant already authorized to examine the manuscript. The referee could ask:

What are the logically independent assumptions? 

Which conclusions depend on each assumption?

Check the derivation from equations 16 to 22.

Does the claimed limiting behaviour actually follow?

Which previous papers are closest to this result?

Are there important missing references?

Where does the manuscript move from demonstrated result to conjecture?

Construct the strongest possible objection to the central result.

Now defend the authors against that objection.

Such interaction would allow the referee to concentrate on the genuinely difficult part of scientific evaluation: understanding what matters. That is where human scientific experience remains indispensable.

The Referee Must Remain Responsible

There should nevertheless be a very clear boundary. A referee should not ask an AI system for a report and simply send that report to the editor. If the AI claims that an equation is wrong, the referee must satisfy himself or herself that the objection is valid before making it. If the AI claims that a result already exists, the referee should examine the relevant literature. If the AI misunderstands an unconventional argument, the referee must recognize the misunderstanding. The human reviewer remains scientifically and ethically responsible for the submitted report.

The governing principle is therefore simple:

\[
\boxed{\text{Do not outsource responsibility to AI. Outsource work to AI.}}
\]

This principle works equally well for authors, referees and editors.

AI Could Also Review the Reviewers

There is an additional possibility. AI could help improve not only manuscripts but referee reports themselves. Before sending a report to the author, the editorial system could ask:

Does this criticism accurately represent what the manuscript says?

Is the referee asking for work far outside the stated scope of the paper?

Does the claim that the result is already known have supporting references?

Are requested citations actually relevant?

Does the referee appear to have misunderstood a definition or convention used by the authors?

This would introduce some quality control into a part of scientific publishing that currently receives remarkably little.

Papers are scrutinized intensely. Authors are scrutinized intensely. Referee reports themselves are much less systematically scrutinized. AI could help correct that imbalance. 

A further problem is the imbalance of power built into the present system. A referee can sometimes become effectively dictatorial: demanding additional calculations, imposing personal preferences about scope or interpretation, insisting on unnecessary revisions, or repeatedly moving the conditions for acceptance, while facing very little accountability for those demands. 

Editors, especially when overloaded or working outside the referee's exact specialization, may simply defer to the report rather than independently judging whether the requests are scientifically justified. The result is an asymmetric process in which authors must answer every objection in detail, while referees are rarely required to justify the proportionality, relevance, or even correctness of their criticisms. 

AI-assisted editorial checking could provide a useful counterweight by asking whether each major referee demand is actually supported by the manuscript, relevant to its claims, and necessary for establishing the validity of the work.

The Strange Reversal of Priorities

This brings us back to the central paradox. Scientific publishing is suffering from a shortage of reviewers. Careful reviewing requires substantial time and intellectual effort. Researchers are increasingly reluctant or simply unable to accept additional referee assignments. Editors struggle to obtain reports. Authors wait. And yet, when one of the most powerful tools ever developed for reducing intellectual workload becomes available, we give considerable latitude to authors while being much more restrictive with reviewers. This seems backwards.

Authors certainly benefit from AI. But the publication system may need AI-assisted refereeing even more urgently.  The appropriate future of peer review is therefore not the replacement of scientists by machines. It is

\[
\boxed{
\text{AI assistance}
+
\text{human scientific judgment}
+
\text{human responsibility}.
}
\]

Used in this way, AI would not weaken peer review. It could help rescue a system that is already struggling under the amount of work we ask human referees to perform.

No comments:

Post a Comment

If Authors Can Use AI, Why Can’t Referees?

  The Reviewer Bottleneck   Peer review has a very basic problem that predates artificial intelligence. Reviewing a scientific paper properl...