Study Tests LLMs Against Experts on Microbial-Oncogenesis Reviews
A Frontiers study published August 6 tested four language models against domain experts on structured reviews of 24 microbial-oncogenesis papers. GPT-5 and GPT-5 Nano produced overall score distributions statistically indistinguishable from the experts, but methodology appraisal and contradiction detection remained the clearest weaknesses, limiting how far the result can be generalized.
Researchers from institutions including the University of the Witwatersrand and Emory University have tested whether large language models can perform structured evidence extraction and critical appraisal at a level comparable with biomedical domain experts. The peer-reviewed study appeared in Frontiers in Cellular and Infection Microbiology on August 6, 2026, with an arXiv version submitted the following day.
What the study measured
The team assembled 24 original research papers about a possible link between Mouse Mammary Tumor Virus-Like Virus and breast cancer. Three experts evaluated the papers with a 77-item template covering multiple-choice, Likert-scale, multi-select, and free-text questions. The same structured task was given to Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano.
The comparison focused on agreement distributions, not on a claim that a model can replace a biomedical reviewer in every setting. Across the complete set of question types, the authors reported that GPT-5 and GPT-5 Nano had score distributions statistically indistinguishable from those of the experts. Gemini models were broadly similar but tended to apply microbial-oncogenesis criteria more leniently.
Where the models still failed
The strongest caveat is that performance varied by task. Methodological appraisal and detecting contradictions in full papers were the most persistent weaknesses. In a manual review of 75 long-answer question instances, the authors found no hallucinations or distortions for GPT-5 or Gemini 2.5 Pro, compared with two for Gemini 2.5 Flash and seven for GPT-5 Nano. Omissions were common across every model, even when most added material was supported by the source paper.
The evaluation also covered one microbial-cancer pairing, so it does not establish performance across the wider biomedical literature. The authors say additional validation is needed across diseases and evidence profiles with different levels of support.
Why it matters for evidence workflows
The result is useful for teams building literature-review systems because it tests a constrained, auditable workflow rather than open-ended summarization. Structured questions, explicit criteria, and comparison with multiple experts made model errors visible. The practical lesson is narrower than the headline result: language models may help scale evidence extraction, but methodology, contradiction checks, and final scientific judgment still need deliberate expert oversight.
Key Points
- 1The evaluation used 24 research papers and a 77-item structured appraisal template.
- 2GPT-5 and GPT-5 Nano matched expert score distributions overall, while Gemini models were more lenient on the focal criteria.
- 3Methodology appraisal, contradiction detection, omissions, and single-case generalizability remain material limitations.
Scoring Rationale
The peer-reviewed study provides a concrete expert-comparison design and practical evidence-workflow implications, but its single microbial-oncogenesis case limits generalizability.
Sources
Primary source and supporting public references used for this report.
Practice with real Health & Insurance data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Health & Insurance problems

