ORIGINAL ARTICLE
Effect of evaluation prompt strategies on LLM-as-a-judge reliability in critical care
,
 
,
 
,
 
,
 
,
 
 
 
More details
Hide details
1
School of Medicine, College of Medicine, National Taiwan University, Taipei, Taiwan
 
2
Department of Anesthesiology, Far Eastern Memorial Hospital, New Taipei City, Taiwan
 
3
Department of Anaesthesiology, National Taiwan University Hospital, Taipei, Taiwan
 
4
Department of Medicine, National Taiwan University Hospital, Taipei, Taiwan
 
These authors had equal contribution to this work
 
 
Submission date: 2025-12-11
 
 
Final revision date: 2026-04-21
 
 
Acceptance date: 2026-05-17
 
 
Publication date: 2026-07-30
 
 
Corresponding author
Yu-Chang Yeh   

3Department of Anaesthesiology, National Taiwan University Hospital, Taipei, Taiwan
 
 
Anaesthesiol Intensive Ther 2026;58(1):156-163
 
KEYWORDS
TOPICS
ABSTRACT
Introduction:
Large language model (LLM)-as-a-judge systems offer scalable evaluation of artificial intelligence (AI)-generated clinical outputs, yet their susceptibility to prompt variability raises concerns regarding reproducibility and alignment with expert judgement. This study examined whether evaluation prompt strategies influence scoring patterns and concordance with clinical raters in critical care.

Material and methods:
This post-hoc analysis used 90 structured clinical reports generated in a prior study using an XGBoost ICU mortality prediction model trained on the MIMIC-IV database. GPT-4o (Azure AI, version 2024-11-20) produced structured interpretations from risk estimates and SHAP attributions. These outputs were evaluated using the IMPACT framework under three evaluation prompt strategies: baseline (E1), top-down decremental (E2), and bottom-up incremental (E3). Agreement between clinician ratings and the automated o3-mini evaluator (Azure AI, version 2025-01-31) was assessed using intraclass correlation coefficients (ICC), with strategy comparisons by Fisher’s z-transformation. Score deviations were examined with repeated-measures ANOVA.

Results:
Mean IMPACT scores were 79.9 (SD 9.9) for E1, 83.3 (SD 9.6) for E2, 78.7 (SD 9.1) for E3, and 78.6 (SD 8.9) for clinicians. All strategies demonstrated substantial agreement (ICC > 0.80). E2 showed significantly lower agreement with clinicians (ICC = 0.82) than E1 and E3 (both ICC = 0.94, p < 0.001). Score deviations differed significantly across strategies (p < 0.001), with E3 showing the smallest mean deviation (0.1) and E2 the largest (4.7).

Conclusions:
Prompt design meaningfully affects both IMPACT scoring patterns and the reliability of LLM-based evaluators. Bottom-up incremental scoring showed the closest alignment with human assessment, underscoring the need for standardised prompt architectures in clinical AI evaluation.
REFERENCES (29)
1.
Sutton RT, Pincock D, Baumgart DC, Sadowski DC, Fedorak RN, Kroeker KI. An overview of clinical decision support systems: benefits, risks, and strategies for success. npj Digit Med 2020; 3: 17. DOI: https://doi.org/10.1038/s41746....
 
2.
Goh S, Goh RSJ, Chong B, Ng QX, Koh GCH, Ngiam KY, et al. Challenges in implementing artificial intelligence in breast cancer screening programmes: systematic review and framework for safe adoption. J Med Internet Res 2025; 27: e62941. DOI: 10.2196/62941.
 
3.
Hassan M, Kushniruk A, Borycki E. Barriers to and facilitators of artificial intelligence adoption in health care: scoping review. JMIR Hum Factors 2024; 11: e48633. DOI: 10.2196/48633.
 
4.
Bang Y, Cahyawijaya S, Lee N, Dai W, Su D, Wilie B, et al. A multitask, multilingual, multimodal evaluation of ChatGPT on reasoning, hallucination, and interactivity. arXiv:2302.04023; 2023. DOI: https://doi.org/10.48550/arXiv....
 
5.
Ji Z, Lee N, Frieske R, Yu T, Su D, Xu Y, et al. Survey of hallucination in natural language generation. ACM Comput Surv 2023; 55: 248. DOI: https://doi.org/10.1145/357173....
 
6.
Schramowski P, Turan C, Andersen N, Rothkopf CA, Kersting K. Large pre-trained language models contain human-like biases of what is right and wrong to do. Nat Mach Intell 2022; 4: 258-268.
 
7.
Tam TYC, Sivarajkumar S, Kapoor S, Stolyar AV, Polanska K, McCarthy KR, et al. A framework for human evaluation of large language models in healthcare derived from literature review. NPJ Digit Med 2024; 7: 258. DOI: https://doi.org/10.1038/s41746....
 
8.
Gu J, Jiang X, Shi Z, Tan H, Zhai X, Xu C, et al. A survey on LLM-as-a-judge. Innovation (Camb) 2026; 7: 101253. DOI: 10.1016/j.xinn.2025.101253.
 
9.
Liu Y, Iter D, Xu Y, Wang S, Xu R, Zhu C. G-Eval: NLG evaluation using GPT-4 with better human alignment. arXiv:2303.16634; 2023. DOI: https://doi.org/10.48550/arXiv....
 
10.
Ouyang L, Wu J, Jiang X, Almeida D, Wainwright C, Mishkin P, et al. Training language models to follow instructions with human feedback. Adv Neural Inform Proc Syst 2022; 35: 27730-27744.
 
11.
Zheng L, Chiang WL, Sheng Y, Zhuang S, Wu Z, Zhuang Y, et al. Judging LLM-as-a-judge with mt-bench and chatbot arena. Adv Neural Inform Proc Syst 2023; 36: 46595-46623.
 
12.
Sivarajkumar S, Kelley M, Samolyk-Mazzanti A, Visweswaran S, Wang Y. An empirical evaluation of prompting strategies for large language models in zero-shot clinical natural language processing: algorithm development and validation study. JMIR Med Inform 2024; 12: e55318. DOI: https://doi.org/10.2196/55318.
 
13.
Wei H, He S, Xia T, Liu F, Wong A, Lin J, Han M. Systematic evaluation of LLM-as-a-judge in LLM alignment tasks: Explainable metrics and diverse prompt templates. arXiv:2408.13006; 2025. DOI: https://doi.org/10.48550/arXiv....
 
14.
He J, Rungta M, Koleczek D, Sekhon A, Wang FX, Hasan S. Does prompt formatting have any impact on LLM performance? arXiv:241110541; 2024. DOI: https://doi.org/10.48550/arXiv....
 
15.
Pezeshkpour P, Hruschka E. Large language models sensitivity to the order of options in multiple-choice questions. arXiv:2308.11483; 2023. DOI: https://doi.org/10.48550/arXiv....
 
16.
Zheng C, Zhou H, Meng F, Zhou J, Huang M. Large language models are not robust multiple choice selectors. arXiv:2309.03882; 2024. DOI: https://doi.org/10.48550/arXiv....
 
17.
Goldberger AL, Amaral LAN, Glass L, Hausdorff JM, Ivanov PC, Mark RG, et al. PhysioBank, PhysioToolkit, and PhysioNet. Components of a new research resource for complex physiologic signals. Circulation 2000; 101: E215-E220.
 
18.
Johnson A, Bulgarelli L, Pollard T, Gow B, Moody B, Horng S, et al. ‘MIMIC-IV’ (version 3.1), PhysioNet RRID:SCR_007345. DOI: https://doi.org/10.13026/kpb9-....
 
19.
Johnson AEW, Bulgarelli L, Shen L, Gayles A, Shammout A, Horng S, et al. MIMIC-IV, a freely accessible electronic health record dataset. Sci Data 2023; 10: 1. DOI: 10.1038/s41597-022-01899-x.
 
20.
Gallifant J, Afshar M, Ameen S, Aphinyanaphongs Y, Chen S, Cacciamani G, et al. The TRIPOD-LLM reporting guideline for studies using large language models. Nat Med 2025; 31: 60-69. DOI: 10.1038/s41591-024-03425-5.
 
21.
Charnock D, Shepperd S, Needham G, Gann R. DISCERN: an instrument for judging the quality of written consumer health information on treatment choices. J Epidemiol Community Health 1999; 53: 105-111.
 
22.
Yeh YC, Yang HY, Chiu CT, Chao A, Chuang YC, Chan WS. Enhancing large language model clinical support information with machine learning risk and explainability: a feasibility study. Intensive Care Med Exp 2026; 14: 51.
 
23.
Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med 2016; 15: 155-163.
 
24.
Wang P, Li L, Chen L, Cai Z, Zhu D, Lin B, et al. Large language models are not fair evaluators. In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics 2024; 1: 9440-9450. DOI: 10.18653/v1/2024.acl-long.511.
 
25.
Wei J, Wang X, Schuurmans D, Bosma M, Ichter B, Xia F, et al. Chain-of-thought prompting elicits reasoning in large language models. arXiv:2201.11903; 2022. DOI: https://doi.org/10.48550/arXiv....
 
26.
An C, Zhang J, Zhong M, Li L, Gong S, Luo Y, et al. Why does the effective context length of LLMs fall short? arXiv:2410.18745; 2024. DOI: https://doi.org/10.48550/arXiv....
 
27.
Hsieh C-P, Sun S, Kriman S, Acharya S, Rekesh D, Jia F, et al. RULER: What’s the real context size of your long-context language models? arXiv:2404.06654; 2024. DOI: https://doi.org/10.48550/arXiv....
 
28.
Panickssery A, Bowman SR, Feng S. LLM evaluators recognise and favour their own generations. arXiv:2404.13076; 2024. DOI: https://doi.org/10.48550/arXiv....
 
29.
Wataoka K, Takahashi T, Ri R. Self-preference bias in LLM-as-a-judge. arXiv:2410.21819; 2025. DOI: https://doi.org/10.48550/arXiv....
 
eISSN:1731-2531
ISSN:1642-5758
Journals System - logo
Scroll to top