Purpose To develop and evaluate a large language model (LLM)-based learning tool, featuring virtual patients (VPs) and virtual assessors (VAs), and to assess its impact on medical students’ perceptions of history-taking education compared to conventional learning methods.
Methods A tool using the GPT-4 API was developed to provide seven clinical VP scenarios and a VA that delivered both immediate, reflective dialogue and comprehensive written feedback. First- and second-year medical students participated in a 6-day study. Pre- and post-participation surveys using a 5-point Likert scale assessed perceptions of the LLM tool versus conventional methods across usability, self-efficacy, and feedback quality domains.
Results Twenty-one students completed the study. The LLM-based tool demonstrated statistically significant improvements over conventional methods in all assessed domains. Students reported greater comfort during practice (mean 4.57 vs. 2.95, p=0.0002). Significant gains were seen in six of eight self-efficacy measures, including confidence in handling unfamiliar cases (4.00 vs. 2.90, p=0.0002). All nine feedback quality dimensions improved significantly, with feedback perceived as more specific (4.43 vs. 3.24, p=0.0005) and personalized (4.19 vs. 3.19, p=0.0001).
Conclusion An LLM-based learning tool featuring VPs and VAs can significantly enhance medical students’ perceived learning experience in history-taking education. It offers a scalable, accessible, and cost-effective complementary training method. Future research should validate these subjective improvements with objective performance metrics.
Citations
Citations to this article as recorded by
AI-Assisted Training for Teleconsultation Competencies in Undergraduate Medical Education: A Narrative Review Wojciech Michał Glinkowski, Barbara Jacennik, Aldona Katarzyna Jankowska, Tomasz Cedro, Szymon Wilk, Rafał Doniec Applied Sciences.2026; 16(10): 4858. CrossRef
Automation and Autonomy in the IVF Laboratory: Concepts and Implications for Embryologists Jacques Cohen, Gerardo Mendizabal-Ruiz, Giles Antony Palmer, Giuseppe Silvestri, Mina Alikani Reproductive BioMedicine Online.2026; : 105920. CrossRef
Opportunities, challenges, and future directions of large language models, including ChatGPT in medical education: a systematic scoping review Xiaojun Xu, Yixiao Chen, Jing Miao Journal of Educational Evaluation for Health Professions.2024; 21: 6. CrossRef
Effectiveness of Using ChatGPT as a Tool to Strengthen Benefits of the Flipped Learning Strategy Gilberto Huesca, Yolanda Martínez-Treviño, José Martín Molina-Espinosa, Ana Raquel Sanromán-Calleros, Roberto Martínez-Román, Eduardo Antonio Cendejas-Castro, Raime Bustos Education Sciences.2024; 14(6): 660. CrossRef
Exploring ChatGPT's Ability to Classify the Structure of Literature Reviews in Engineering Research Articles Maha Issa, Marwa Faraj, Niveen AbiGhannam IEEE Transactions on Learning Technologies.2024; 17: 1819. CrossRef
Exploring an LLM's Use in Supporting Journal Club Preparation and Discussion Among Residents Fahad Umer, Ayesha Mansoor, Azra Naseem, Syed Murtaza Raza Kazmi Journal of Dental Education.2026; 90(7): 962. CrossRef
Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions Jaehyun Byun, Hongik Kim, Jihan Lim, Junyeong Choi, Sangzin Ahn Korean Journal of Medical Education.2026; 38(1): 64. CrossRef
Make your teaching generative artificial intelligence-ready: a 1-hour workshop for health professions educators Anshul Kumar, Keri Barksdale Mans Korean Journal of Medical Education.2026; 38(1): 18. CrossRef
Artificial Intelligence in Undergraduate Medical Education: A Cross-Sectional Study of Utilization Patterns and Perceptions Among Medical Students Raju R Bokan, Rashmi Malhotra , Mukund Vatsa, Kanchan Bisht, Mukesh Singla, Rajeev Choudhary Cureus.2026;[Epub] CrossRef
A Multi-Chatbot Analysis: Strengths and Weaknesses in Neuroanatomy Learning Alessandro Naim, Sara Naim, Daniele Saverino Information.2026; 17(5): 475. CrossRef
Exploring communication self-efficacy and artificial intelligence generated assessment tools in primary care education Constanze Dietzsch, Johanna Klutmann, Aline Köhler, Sara Volz-Willems, Johannes Jäger, Fabian Dupont Discover Education.2026;[Epub] CrossRef
Exploring Radiology Postgraduate Students' Engagement with Large Language Models for Educational Purposes: A Study of Knowledge, Attitudes, and Practices Pradosh Kumar Sarangi, Braja Behari Panda, Sanjay P., Debabrata Pattanayak, Swaha Panda, Himel Mondal Indian Journal of Radiology and Imaging.2025; 35(01): 035. CrossRef
Using large language models (ChatGPT, Copilot, PaLM, Bard, and Gemini) in Gross Anatomy course: Comparative analysis Volodymyr Mavrych, Paul Ganguly, Olena Bolgova Clinical Anatomy.2025; 38(2): 200. CrossRef
A comprehensive survey of large language models and multimodal large language models in medicine Hanguang Xiao, Feizhong Zhou, Xingyue Liu, Tianqi Liu, Zhipeng Li, Xin Liu, Xiaoxuan Huang Information Fusion.2025; 117: 102888. CrossRef
Intelligenza generativa artificiale in medical education: ragionamento clinico artificiale vs ragionamento clinico umano Rosa Cera EDUCATION SCIENCES AND SOCIETY.2025; (2): 239. CrossRef
Application of large language models in healthcare: A bibliometric analysis Lanping Zhang, Qing Zhao, Dandan Zhang, Meijuan Song, Yu Zhang, Xiufen Wang DIGITAL HEALTH.2025;[Epub] CrossRef
Comparative Evaluation of Artificial Intelligence Models for Contraceptive Counseling Anisha V. Patel, Sona Jasani, Abdelrahman AlAshqar, Rushabh H. Doshi, Kanhai Amin, Aisvarya Panakam, Ankita Patil, Sangini S. Sheth Digital.2025; 5(2): 10. CrossRef
Application of large language models in medicine Fenglin Liu, Hongjian Zhou, Boyang Gu, Xinyu Zou, Jinfa Huang, Jinge Wu, Yiru Li, Sam S. Chen, Yining Hua, Peilin Zhou, Junling Liu, Chengfeng Mao, Chenyu You, Xian Wu, Yefeng Zheng, Lei Clifton, Zheng Li, Jiebo Luo, David A. Clifton Nature Reviews Bioengineering.2025; 3(6): 445. CrossRef
Artificial Intelligence in Medical Education: A Practical Guide for Educators Nivritti Gajanan Patil, Nga Lok Kou, Daniel T. Baptista‐Hon, Olivia Monteiro MedComm – Future Medicine.2025;[Epub] CrossRef
A Systematic Review of Gen-AI Applications in Education: Rewards, Challenges and Future Prospects Everleen Nekesa Wanyonyi, Millicent K. Murithi Pan-African Journal of Education and Social Sciences.2025; 6(1): 1. CrossRef
A Comprehensive Overview of Large Language Models Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, Ajmal Mian ACM Transactions on Intelligent Systems and Technology.2025; 16(5): 1. CrossRef
Comparative performance of ChatGPT, Gemini, and final-year emergency medicine clerkship students in answering multiple-choice questions: implications for the use of AI in medical education Shaikha Nasser Al-Thani, Shahzad Anjum, Zain Ali Bhutta, Sarah Bashir, Muhammad Azhar Majeed, Anfal Sher Khan, Khalid Bashir International Journal of Emergency Medicine.2025;[Epub] CrossRef
Semi-automated Systematic Review: Main applications and trends of foundation models German Cuaya Simbro, Emmanuel Ramírez Romero, Ismael Ortega García Revista Ingenierías Universidad de Medellín.2025; 24(47): 1. CrossRef
Comparative evaluation of AI platforms “Google Gemini 2.5 Flash, Google Gemini 2.0 Flash, DeepSeek V3 and ChatGPT 4o” in solving multiple-choice questions from different subtopics of anatomy Anjali Singal, Swati Goyal Surgical and Radiologic Anatomy.2025;[Epub] CrossRef
ChatGPT as a Virtual Peer: Enhancing Critical Thinking in Flipped Veterinary Anatomy Education Nieves Martín-Alguacil, Luis Avedillo, Rubén A. Mota-Blanco, Mercedes Marañón-Almendros, Miguel Gallego-Agúndez International Medical Education.2025; 4(3): 34. CrossRef
Large language models and their impact in medical imaging education Jiajia Zhu, Huanhuan Cai PeerJ Computer Science.2025; 11: e3433. CrossRef
Tıp Eğitiminde Yapay Zeka: Asistan Hekimlerin Kullanım Alanları ve Algıları Hilal Hatice Ülkü, Selcen Öncü, Fulya Torun Tıp Eğitimi Dünyası.2025; 24(74): 46. CrossRef
Large Language Models and Artificial Intelligence: A Primer for Plastic Surgeons on the Demonstrated and Potential Applications, Promises, and Limitations of ChatGPT Jad Abi-Rafeh, Hong Hao Xu, Roy Kazan, Ruth Tevlin, Heather Furnas Aesthetic Surgery Journal.2024; 44(3): 329. CrossRef
Utilizing GPT-4 and generative artificial intelligence platforms for surgical education: an experimental study on skin ulcers Ishith Seth, Bryan Lim, Jevan Cevik, Foti Sofiadellis, Richard J. Ross, Roberto Cuomo, Warren M. Rozen European Journal of Plastic Surgery.2024;[Epub] CrossRef
Emerging Voices in Drug Delivery – Breaking Barriers (Issue 1) Juliane Nguyen, Shawn C. Owen Advanced Drug Delivery Reviews.2024; 208: 115273. CrossRef
Quantitative Evaluation of Large Language Models to Streamline Radiology Report Impressions: A Multimodal Retrospective Analysis Rushabh Doshi, Kanhai S. Amin, Pavan Khosla, Simar Bajaj, Sophie Chheang, Howard P. Forman Radiology.2024;[Epub] CrossRef
Emerging Voices in Drug Delivery – Harnessing and Modulating Complex Biological Systems (Issue 2) Shawn C. Owen, Juliane Nguyen Advanced Drug Delivery Reviews.2024; 208: 115293. CrossRef
The role of artificial intelligence in Physical Therapy education Scott William Lowe Bulletin of Faculty of Physical Therapy.2024;[Epub] CrossRef
A systematic review of large language models and their implications in medical education Harrison C. Lucas, Jeffrey S. Upperman, Jamie R. Robinson Medical Education.2024; 58(11): 1276. CrossRef
ChatGPT and Other Large Language Models in Medical Education — Scoping Literature Review Alexandra Aster, Matthias Carl Laupichler, Tamina Rockwell-Kollmann, Gilda Masala, Ebru Bala, Tobias Raupach Medical Science Educator.2024; 35(1): 555. CrossRef
Embracing Large Language Models for Adult Life Support Learning Serena Patel, Rohit Patel Cureus.2024;[Epub] CrossRef
Large Language Models in Medical Education: Opportunities, Challenges, and Future Directions Alaa Abd-alrazaq, Rawan AlSaad, Dari Alhuwail, Arfan Ahmed, Padraig Mark Healy, Syed Latifi, Sarah Aziz, Rafat Damseh, Sadam Alabed Alrazak, Javaid Sheikh JMIR Medical Education.2023; 9: e48291. CrossRef
Data Science as a Core Competency in Undergraduate Medical Education in the Age of Artificial Intelligence in Health Care Puneet Seth, Nancy Hueppchen, Steven D Miller, Frank Rudzicz, Jerry Ding, Kapil Parakh, Janet D Record JMIR Medical Education.2023; 9: e46344. CrossRef
Performance of Large Language Models (ChatGPT, Bing Search, and Google Bard) in Solving Case Vignettes in Physiology Anup Kumar D Dhanvijay, Mohammed Jaffer Pinjar, Nitin Dhokane, Smita R Sorte, Amita Kumari, Himel Mondal Cureus.2023;[Epub] CrossRef
A use case of ChatGPT in a flipped medical terminology course Sangzin Ahn Korean Journal of Medical Education.2023; 35(3): 303. CrossRef
Assessing the Utilization of Large Language Models in Medical Education: Insights From Undergraduate Medical Students Sairavi Kiran Biri, Subir Kumar, Muralidhar Panigrahi, Shaikat Mondal, Joshil Kumar Behera, Himel Mondal Cureus.2023;[Epub] CrossRef
Transforming clinical trials: the emerging roles of large language models Jong-Lyul Ghim, Sangzin Ahn Translational and Clinical Pharmacology.2023; 31(3): 131. CrossRef
Large Language Model-Based Neurosurgical Evaluation Matrix: A Novel Scoring Criteria to Assess the Efficacy of ChatGPT as an Educational Tool for Neurosurgery Board Preparation Sneha Sai Mannam, Robert Subtirelu, Daksh Chauhan, Hasan S. Ahmad, Irina Mihaela Matache, Kevin Bryan, Siddharth V.K. Chitta, Shreya C. Bathula, Ryan Turlip, Connor Wathen, Yohannes Ghenbot, Sonia Ajmera, Rachel Blue, H. Isaac Chen, Zarina S. Ali, Neil Ma World Neurosurgery.2023; 180: e765. CrossRef
Assessment of the capacity of ChatGPT as a self-learning tool in medical pharmacology: a study using MCQs Woong Choi BMC Medical Education.2023;[Epub] CrossRef
Evaluation of the performance of GPT-3.5 and GPT-4 on the Polish Medical Final Examination Maciej Rosoł, Jakub S. Gąsior, Jonasz Łaba, Kacper Korzeniewski, Marcel Młyńczak Scientific Reports.2023;[Epub] CrossRef
Adapting to the Impact of Artificial Intelligence in Scientific Writing: Balancing Benefits and Drawbacks while Developing Policies and Regulations Ahmed Salem Bahammam, Khaled Trabelsi, Seithikurippu R. Pandi-Perumal, Haitham Jahrami Journal of Nature and Science of Medicine.2023; 6(3): 152. CrossRef
Performance of a Large Language Model in Medical Pharmacology Education: An Assessment Using Multiple-Choice Questions Benjamin S. Wright, Laura J. Kim, Nathan R. Coleman Annals of Pharmacy Education, Safety, and Public Health Advocacy.2023; 3(1): 232. CrossRef