Evaluating Large Language Models for Automated Evidence Synthesis in Neuroimaging AI: A Multi-Model Benchmark
Journal of Clinical Medicine, vol.15, no.11, 2026 (SCI-Expanded, Scopus)
- Publication Type: Article / Article
- Volume: 15 Issue: 11
- Publication Date: 2026
- Doi Number: 10.3390/jcm15114230
- Journal Name: Journal of Clinical Medicine
- Journal Indexes: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Chemical Abstracts Core, EMBASE, Academic Search Ultimate (EBSCO), Health Research Premium Collection (ProQuest)
- Keywords: artificial intelligence, benchmarking, evidence synthesis, information extraction, large language models, neuroimaging
- Acibadem Mehmet Ali Aydinlar University Affiliated: Yes
Abstract
Background: Data extraction for systematic reviews is highly resource-intensive. This study evaluated four frontier large language models (LLMs) on complex structured metadata extraction from specialized neuroimaging artificial intelligence (AI) literature to determine their performance in automated evidence synthesis. Methods: We compared Google Gemini 3 Pro Preview, Anthropic Claude Opus 4.5, Perplexity Sonar Pro, and OpenAI GPT 5.2. Using a standardized prompt, each model extracted 22 variables from 91 peer-reviewed neuroimaging AI articles. The variables were stratified into low-, medium-, and high-complexity tiers. The performance was measured via the exact-match accuracy against a consensus-based expert ground truth. Results: The overall exact-match accuracy was moderate. Gemini 3 Pro Preview achieved the highest overall rate (56.4%), followed by Sonar Pro (52.1%), Claude Opus 4.5 (51.3%), and GPT 5.2 (46.5%). Gemini significantly outperformed all other models (p < 0.001). The performance declined dramatically as the variable complexity increased. Across models, the accuracy was 88.9–92.9% for low-complexity categorical fields, 47.0–63.3% for medium-complexity text extraction, and 2.7–15.5% for high-complexity variables requiring clinical judgment or multi-section synthesis. The most common type of error was misclassification. All four models scored 0% on the main performance metric, but this reflected a representational mismatch with the ground truth rather than extraction failure, indicating that the exact-match accuracy underestimates the true semantic performance. Conclusions: Frontier LLMs can effectively automate the retrieval of simple categorical data, but have serious difficulties with methodological variables that are complex. Although extraction can be fully automated for low-complexity fields, human review remains essential for context-dependent variables that require clinical judgment.