LLM-based virtual patient versus traditional teaching in ophthalmology clerkship: a controlled trial

Journal: Region - Educational Research and Reviews DOI: 10.32629/rerr.v8i4.5344

Dingqiao WANG, Jun MAO, Wei DU

Department of Ophthalmology, The Eighth Affiliated Hospital, Sun Yat-Sen University

Abstract

Background: Clinical history-taking is central to medical education, yet practice opportunities remain scarce in short specialty rotations such as ophthalmology. No study has directly compared two commercial large language models (LLMs) operating within the same virtual patient framework. Methods: We conducted a quasi-experimental sequential cohort study (TREND-reported). Forty-five medical students completing a five-day ophthalmology clerkship were allocated to one of three groups (n = 15 each): a control group receiving traditional teaching with real patients, and two intervention groups using a retrieval-augmented generation (RAG)-based virtual patient system powered by GPT-4o or Claude 3.5 Sonnet. Primary outcome was history-taking performance (100-point rubric, ICC = 0.89) on Day 2 and Day 5. Secondary outcomes were self-efficacy (SE-12 scale) and student satisfaction, assessed by paired t-test and Fisher's exact test. Results: History-taking improved by 17.11 points (GPT-4o, p < 0.001) and 15.17 points (Claude, p = 0.001), versus a non-significant 5.87-point change in controls (p = 0.078). Self-efficacy gains were 25.77 points (GPT-4o, p < 0.001), 21.19 points (Claude, p = 0.001), and 2.39 points (control, p = 0.536). All intervention students rated satisfaction as basically satisfied or above, versus 80% in the control group (Fisher's exact p < 0.001). The two LLM groups did not differ significantly on any measure. Conclusions: Both RAG-based virtual patient systems (GPT-4o and Claude 3.5 Sonnet) improved history-taking performance and clinical self-efficacy within five days; traditional teaching produced no significant change on either measure. The comparable outcomes across LLMs indicate that system architecture and case design matter more than model selection. These findings support the use of LLM-based virtual patients as a practical supplement to bedside teaching in compressed clerkships, subject to confirmation in randomized trials.

Keywords

large language model; virtual patient; history-taking; clinical self-efficacy; ophthalmology clerkship; medical education; retrieval-augmented generation

References

[1] Götz S, Fournier J, Tisjlar K, et al. Effects of the quality of medical history taking on diagnostic accuracy[J/OL]. Signa Vitae, 2023, 19(5): 68-74. https://doi.org/10.22514/sv.2023.081
[2] Flanagan O L, Cummings K M. Standardized patients in medical education: A review of the literature[J/OL]. Cureus, 2023, 15(7): e42027. https://doi.org/10.7759/cureus.42027
[3] Hamilton A, Molzahn A, McLemore K. The evolution from standardized to virtual patients in medical education[J/OL]. Cureus, 2024, 16(10): e71224. https://doi.org/10.7759/cureus.71224
[4] Li D, Lutfi S L. Large language model–based virtual patient systems for history-taking in medical education: Comprehensive systematic review[J/OL]. JMIR Medical Informatics, 2026, 1: e79039. https://doi.org/10.2196/79039
[5] Luo M J, Bi S, Pang J, et al. A large language model digital patient system enhances ophthalmology history taking skills[J/OL]. npj Digital Medicine, 2025. https://doi.org/10.1038/s41746-025-01841-6
[6] Axboe M K, Christensen K S, Kofoed P E, et al. Development and validation of a self-efficacy questionnaire (SE-12) measuring the clinical communication skills of health care professionals[J/OL]. BMC Medical Education, 2016, 16(1): 272. https://doi.org/10.1186/s12909-016-0798-7
[7] Wang D, Liang J, Ye J, et al. Enhancement of the performance of large language models in diabetes education through retrieval-augmented generation: comparative study[J/OL]. J Med Internet Res, 2024, 26: e58041.
https://doi.org/10.2196/58041
[8] Baker E A, Ledford C H, Fogg L, et al. The IDEA assessment tool: assessing the reporting, diagnostic reasoning, and decision-making skills demonstrated in medical students' hospital admission notes[J/OL]. Teach Learn Med, 2015, 27(2): 163-173. https://doi.org/10.1080/10401334.2015.1011654
[9] Malik T G, Mahboob U, Khan R A, et al. Virtual patients versus standardized patients for improving clinical reasoning skills in ophthalmology residents: a randomized controlled trial[J/OL]. BMC Med Educ, 2024, 24(1): 429. https://doi.org/10.1186/s12909-024-05241-4
[10] Cui Y, Wang L, Dong C, et al. A Bloom's Taxonomy-integrated rotation model enhances clinical reasoning and practical skills in optometry interns during ophthalmology rotation: a randomized controlled trial[J/OL]. Front Med (Lausanne), 2026, 13:1746533. https://doi.org/10.3389/fmed.2026.1746533
[11] Xu L, Xu Q, Liu C, et al. Virtual standardized patients for improving clinical thinking ability training in residents: Randomized controlled trial[J/OL]. JMIR Medical Education, 2025.
https://pmc.ncbi.nlm.nih.gov/articles/PMC12685284/

Copyright © 2026 Dingqiao WANG, Jun MAO, Wei DU

Creative Commons License
This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License