Evaluating Large Language Models for Automated Evidence Synthesis in Neuroimaging AI: A Multi-Model Benchmark


Sulaimanov U., Sanlier N., Moniri A., Demir B., Serikkanov Y., Bayramoglu A. R., ...Daha Fazla

Journal of Clinical Medicine, cilt.15, sa.11, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 15 Sayı: 11
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/jcm15114230
  • Dergi Adı: Journal of Clinical Medicine
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Chemical Abstracts Core, EMBASE, Academic Search Ultimate (EBSCO), Health Research Premium Collection (ProQuest)
  • Anahtar Kelimeler: artificial intelligence, benchmarking, evidence synthesis, information extraction, large language models, neuroimaging
  • Acıbadem Mehmet Ali Aydınlar Üniversitesi Adresli: Evet

Özet

Background: Data extraction for systematic reviews is highly resource-intensive. This study evaluated four frontier large language models (LLMs) on complex structured metadata extraction from specialized neuroimaging artificial intelligence (AI) literature to determine their performance in automated evidence synthesis. Methods: We compared Google Gemini 3 Pro Preview, Anthropic Claude Opus 4.5, Perplexity Sonar Pro, and OpenAI GPT 5.2. Using a standardized prompt, each model extracted 22 variables from 91 peer-reviewed neuroimaging AI articles. The variables were stratified into low-, medium-, and high-complexity tiers. The performance was measured via the exact-match accuracy against a consensus-based expert ground truth. Results: The overall exact-match accuracy was moderate. Gemini 3 Pro Preview achieved the highest overall rate (56.4%), followed by Sonar Pro (52.1%), Claude Opus 4.5 (51.3%), and GPT 5.2 (46.5%). Gemini significantly outperformed all other models (p < 0.001). The performance declined dramatically as the variable complexity increased. Across models, the accuracy was 88.9–92.9% for low-complexity categorical fields, 47.0–63.3% for medium-complexity text extraction, and 2.7–15.5% for high-complexity variables requiring clinical judgment or multi-section synthesis. The most common type of error was misclassification. All four models scored 0% on the main performance metric, but this reflected a representational mismatch with the ground truth rather than extraction failure, indicating that the exact-match accuracy underestimates the true semantic performance. Conclusions: Frontier LLMs can effectively automate the retrieval of simple categorical data, but have serious difficulties with methodological variables that are complex. Although extraction can be fully automated for low-complexity fields, human review remains essential for context-dependent variables that require clinical judgment.