Skip to main navigation Skip to main content
  • KSME
  • E-Submission

KJME : Korean Journal of Medical Education

OPEN ACCESS
ABOUT
BROWSE ARTICLES
FOR AUTHORS AND REVIEWERS

Articles

Original Research

Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions

Korean Journal of Medical Education 2026;38(1):64-73.
Published online: February 20, 2026

1Inje University College of Medicine, Busan, Korea

2Department of Pharmacology and Pharmacogenomics Research Center, Inje University College of Medicine, Busan, Korea

3Cardiovascular and Metabolic Diseases Medical Research Center, Inje University College of Medicine, Busan, Korea

Corresponding Author: Sangzin Ahn (https://orcid.org/0000-0003-2749-0014) Department of Pharmacology, Inje University College of Medicine, 75 Bokji-ro, Busanjin-gu, Busan 47392, Korea Tel: +82.51.890.5909 Fax: +82.51.893.1232 E-mail: sangzinahn@inje.ac.kr
• Received: September 9, 2025   • Revised: December 27, 2025   • Accepted: January 6, 2026

© The Korean Society of Medical Education.

This is an open-access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/3.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 1,246 Views
  • 81 Download
  • 1 Crossref
  • 1 Scopus
prev next
  • Purpose
    To develop and evaluate a large language model (LLM)-based learning tool, featuring virtual patients (VPs) and virtual assessors (VAs), and to assess its impact on medical students’ perceptions of history-taking education compared to conventional learning methods.
  • Methods
    A tool using the GPT-4 API was developed to provide seven clinical VP scenarios and a VA that delivered both immediate, reflective dialogue and comprehensive written feedback. First- and second-year medical students participated in a 6-day study. Pre- and post-participation surveys using a 5-point Likert scale assessed perceptions of the LLM tool versus conventional methods across usability, self-efficacy, and feedback quality domains.
  • Results
    Twenty-one students completed the study. The LLM-based tool demonstrated statistically significant improvements over conventional methods in all assessed domains. Students reported greater comfort during practice (mean 4.57 vs. 2.95, p=0.0002). Significant gains were seen in six of eight self-efficacy measures, including confidence in handling unfamiliar cases (4.00 vs. 2.90, p=0.0002). All nine feedback quality dimensions improved significantly, with feedback perceived as more specific (4.43 vs. 3.24, p=0.0005) and personalized (4.19 vs. 3.19, p=0.0001).
  • Conclusion
    An LLM-based learning tool featuring VPs and VAs can significantly enhance medical students’ perceived learning experience in history-taking education. It offers a scalable, accessible, and cost-effective complementary training method. Future research should validate these subjective improvements with objective performance metrics.
Accurate diagnosis by clinicians begins with effective history taking, a fundamental clinical skill that forms the cornerstone of patient care [1]. Medical students must master this skill through deliberate practice and structured feedback, yet significant challenges exist in providing adequate training opportunities. This study introduces and evaluates a novel large language model (LLM)-based tool designed to address these educational gaps through virtual patients (VPs) and virtual assessors (VAs).
Medical education research consistently demonstrates the superiority of active, experiential learning methods for developing history taking skills. A comprehensive review utilizing the Medical Education Research Study Quality Instrument found that “learning by doing” approaches significantly outperformed traditional methods such as lectures, demonstrations, and self-study [1]. Small group workshops incorporating role-playing, patient interviews, and structured feedback from peers and facilitators emerged as the most effective teaching strategies. The critical importance of practice with immediate feedback has been well established [2,3].
Standardized patients (SPs) have long served as the gold standard for clinical skills training, providing consistent, realistic patient encounters that enhance students’ confidence, clinical competence, and communication abilities [4]. However, the widespread implementation of SP programs faces substantial practical barriers. In South Korea, a comprehensive examination with eight SP cases costs approximately 100,000 Korean won per student, limiting the frequency and accessibility of these valuable learning experiences [5]. The COVID-19 (coronavirus disease 2019) pandemic accelerated adoption of digital learning solutions in medical education [6], coinciding with rapid advances in artificial intelligence (AI) and natural language processing capabilities.
Several alternative approaches have emerged to address SP program limitations, including Peer-led multi-role practice Objective Structured Clinical Examinations (PrOSCEs) [7] and computer-based VPs [8]. While these innovations offer partial solutions, they often lack the conversational flexibility and comprehensive feedback mechanisms necessary for optimal skill development. Recent developments in LLMs present unprecedented opportunities for medical education innovation. GPT-4 has shown strong performance on medical licensing examinations and clinical reasoning tasks [9]. A pioneering study proposed using chatbot technology for clinical interview training [10]. Since then, several groups have demonstrated the feasibility of GPT-4-powered VPs for history-taking education. Holderried et al. [11] showed that a GPT-4-based chatbot could provide medically plausible responses in over 99% of interactions, with automated feedback achieving almost perfect agreement (Cohen κ=0.832) with human raters. Rädel-Ablass et al. [12] reported that over 80% of health professions students rated GPT-4 VP accuracy as good to excellent, with students preferring AI-driven training over traditional role-plays. It has also been conceptualized that LLM-based VPs are a ‘disruptive innovation’ enabling scalable, low-cost clinical training that can reach previously underserved learners globally [13].
Building on these recent advances, our study develops and evaluates a student-centered, LLM-based learning tool that integrates VPs with VAs using the GPT-4 API. This tool was designed to offer key advantages over traditional methods, including practice opportunities without time or location constraints, consistent and structured feedback across encounters, and a private environment to potentially reduce performance anxiety, all at a comparatively low cost. The primary aim of this study was to examine how medical students perceive the learning experience offered by this tool in relation to conventional learning methods (CLMs). Accordingly, we compared students' perceptions across three key domains: usability, self-efficacy development, and feedback quality.
1. Study design and educational framework
The educational framework of our study is informed by the PrOSCE model, where students typically manage all aspects of the examination [7]. While effective, research on this model indicates that students strongly prefer the active role of ‘student doctor’ over other required peer roles. Our design leverages this insight by automating these less-preferred roles, creating VPs and VAs with an LLM. This adaptation allows participants to focus exclusively on the high-value task of clinical practice.
While the PrOSCE model informed the structural design of our tool, its pedagogical features were guided by established principles for enhancing history-taking skills [14]: (1) use of clinical cases representative of summative assessments, (2) simulation of examination conditions, (3) opportunities to practice structured probing, (4) provision of structured feedback, (5) prompts for student self-reflection, and (6) collection of student feedback on the tool.
2. The LLM-based learning tool
The learning tool consists of two core AI-driven components: a VP for interview practice and a VA for feedback. Both were developed using the GPT-4 API (gpt-4-0613 model; OpenAI, San Francisco, USA) with sophisticated prompt engineering, designing detailed instructions to guide the subsequent LLM’s output generation [15]. All clinical content for the VP’s scenario and VA’s assessment criteria was derived from “The guide to clinical performance” by the Korean Association of Medical Colleges [16] and “The patient history: evidence-based approach,” second edition [17]. Following contemporary validity frameworks for simulation-based assessment tools [18], content validity was established through expert review by medical specialists in relevant fields to ensure clinical accuracy and representativeness. Response process validity was examined through iterative pilot testing by the researchers to confirm that the VP responses aligned with intended clinical presentations and that the automated scoring correctly detected checklist elements from student-VP dialogues.

1) The virtual patient

The VP was designed to simulate a realistic patient encounter through a text-based chat interface (Fig. 1A). A foundational ‘base prompt’ provided general behavioral instructions, including: (1) role definition, (2) scenario grounding, (3) within-session consistency and memory, and (4) communication and disclosure style. This base prompt was combined with one of seven specific clinical scenarios selected from the 48 identified in the Korean Medical Licensure Examination: chest pain, hemoptysis, diarrhea, red urine, fatigue, joint pain, and headache. Each scenario prompt provided comprehensive patient information, including demographics and specific medical conditions (examples are available in the Supplements 14).

2) The virtual assessor and multifaceted feedback

The VA was engineered to provide automated, multifaceted feedback through two distinct mechanisms: an immediate, reflective dialogue and a comprehensive, delayed written report.
The first mechanism was an immediate, reflective dialogue. Following each patient encounter, the VA initiated a structured, interactive discussion with the student designed to foster critical self-reflection on their clinical approach and differential diagnosis. This process was guided by the five stages of critical thinking by Kamin et al. [19] (Fig. 1B).
The second mechanism was a comprehensive performance report delivered at the conclusion of the study period. To generate this, the VA analyzed the full transcript of the student-VP interaction, automatically scoring performance by comparing questions and actions against a case-specific checklist. The final report was structured according to high-quality objective structured clinical examination (OSCE) feedback principles [20] and the ‘enhanced written feedback’ template [21], presenting the analysis in clear columns for ‘checklist item,’ ‘performance evidence,’ ‘score,’ and detailed ‘feedback’ that included explanations of clinical reasoning (Fig. 1C).
3. Participant recruitment and study procedure
This study was approved by the Institutional Review Board of Busan Paik Hospital (BPIRB 2023-09-026-007). Participants were recruited from first- and second-year classes of a single medical school in Korea. Informed consent was obtained online at the beginning of the survey.
The study was conducted over a 6-day period in September 2023 and followed a structured procedure. First, participants completed a pre-participation survey and registered on a custom-developed web-based platform. Upon registration, each student was assigned a unique ID to ensure confidentiality while linking survey responses to platform usage.
The study was conducted remotely, with participants accessing the platform using personal computers or laptops. During the 6-day intervention period, participants could engage with the LLM-based learning tool as frequently as desired, with no restrictions on the number of sessions or repeated access to the same VP scenarios. Each history-taking session with a VP was limited to a maximum of 10 minutes, a typical time limit for a SP session, and was immediately followed by a reflective discussion with the VA. At the conclusion of the study, comprehensive written feedback reports and scores were provided on the platform, after which participants completed the final post-participation survey.
4. Data collection and analysis

1) Survey instrument

Pre- and post-participation surveys compared student perceptions of the LLM-based tool against CLMs. The core 16-item questionnaire assessed three domains: usability, self-efficacy, and feedback quality, using 5-point Likert scales (1=‘strongly disagree’ to 5=‘strongly agree’). The post-survey included nine additional questions assessing overall satisfaction and specific tool perceptions, including comparisons with printed materials commonly used for SP session and OSCE preparation.

2) Data analysis

Analysis of primary learning methods was based on 31 pre-survey responses. Core effectiveness analysis used 21 complete paired pre- and post-survey responses. Statistical analysis employed the Wilcoxon signed-rank test, with paired t-test as alternative if Shapiro-Wilk test confirmed normal distribution. Pearson and Spearman correlations examined relationships between dialogue metrics (turn count, duration) and automated assessment scores. Independent samples t-tests and Mann-Whitney U tests compared performance between first and repeat attempts.
1. Participant characteristics and study completion
Of 34 students initially expressing interest, 31 completed the pre-survey, 24 registered on the platform, and 21 completed both surveys. Among pre-survey respondents, passive learning methods predominated: 83.9% relied on “The guide to clinical performance” textbook, 51.6% used commercial clinical performance examination review books, while only 25.8% engaged in active peer practice and 6.5% used educational videos (specific results can be found in the Supplement 5).
2. Tool usage patterns
During the 6-day study period, the 21 participants who completed the study engaged with the LLM-based tool an average of 2.2 times (range, 1–7 sessions). Nine students (47.6%) used the tool at least twice, with six students (28.6%) engaging three or more times. A total of 47 conversation logs were collected.
3. Dialogue metrics and performance
Across all 41 VP sessions with complete dialogue data, students engaged in a mean of 19.24 turns (standard deviation [SD]=6.32) per session, with a mean duration of 6.58 minutes (SD=2.05). The mean assessment score was 58.27% (SD=16.72). A strong positive correlation was observed between the number of conversational turns and assessment score (r=0.66, p<0.001), while session duration showed a moderate correlation with score (r=0.42, p=0.006) (Fig. 2A, B). When comparing first attempts with repeat attempts, repeat attempts showed significantly lower scores (53.71%±17.18% vs. 63.56%±14.91%, p=0.035) (Fig. 2C). This performance decline was accompanied by significantly reduced engagement: students asked fewer questions (16.64±5.81 turns vs. 22.26±5.62 turns, p=0.003) and spent less time (5.92±1.91 minutes vs. 7.34±1.98 minutes, p=0.024) in repeat sessions. Paired analysis of students with multiple attempts (n=9) confirmed this pattern, showing a mean score decline of 10.30 percentage points (p=0.039) (Fig. 2D).
4. Comparative effectiveness of the LLM tool

1) Usability domain

Students reported significantly improved usability with the LLM-based tool compared to SPs (Table 1). Participants found it easier to focus on the patient and felt substantially more comfortable practicing history taking.

2) Self-efficacy domain

The LLM-based tool demonstrated significant improvements in six of eight measured self-efficacy domains (Table 2). Significant gains were observed in clinical reasoning ability, problem-solving skills, confidence in handling unfamiliar clinical presentations, overall confidence in history-taking abilities, history-taking skills, and confidence for exam performance. Communication skills and understanding of clinical presentations showed positive trends but did not reach statistical significance (p=0.0621 and p=0.0578, respectively).

3) Feedback quality domain

All nine dimensions of feedback quality showed statistically significant improvements with the LLM-based tool (Table 3). The most substantial improvements were observed in specificity and detail of feedback, personalization to individual needs, and provision of appropriate discussion topics.
5. Student satisfaction and acceptance
Students expressed strong overall satisfaction with the LLM-based learning experience (Table 4). The tool was rated highly useful for history taking practice and enabled more extensive practice opportunities. As printed materials represent the predominant method for OSCE preparation in our baseline data, we assessed the tool’s comparative value in this preparatory context. Students found the tool more helpful for improving clinical skills and more motivating. The tool’s helpfulness for acquiring medical knowledge and the realism of VP responses received moderate ratings.
6. Economic analysis
Total API usage cost, encompassing developmental testing, all 47 VP interactions, and generation of the detailed feedback reports, was US$245.48, translating to approximately US$11.69 per student for the entire study period, or US$5.22 per individual practice session. When calculated across eight cases [5], the LLM-based tool achieved an approximate 45% cost reduction compared to traditional SP sessions while providing personalized and flexible practice opportunities.
1. Principal findings
Our study demonstrates that an LLM-based learning tool featuring VPs and VAs significantly enhances medical students’ perceptions of history-taking education compared to CLMs across three critical domains: usability, self-efficacy, and feedback quality, while reducing costs by approximately 45%.
2. Educational context and barriers to active learning
Our baseline data revealed concerning predominance of passive learning methods among medical students, with only 25.8% engaging in active peer practice despite established evidence favoring experiential approaches (Supplement 5). This gap reflects multiple barriers: cultural factors in Asian educational contexts that foster a reluctance to expose knowledge gaps or provide critical feedback to peers [21]; the documented limitations of peer role-play, which can lack structure and result in significant variability [3]; and the high cost and limited availability of SP sessions [5].
Our LLM-based tool directly addresses these barriers by providing a private, non-judgmental learning environment available at all times. The significant improvements in comfort level (mean 4.57 vs. 2.95, p=0.0002) and focus on the patient can be interpreted through the lens of cognitive load theory [22]. Text-based interaction eliminates extrinsic cognitive load associated with processing multimodal cues and performance anxiety, allowing learners to concentrate cognitive resources on the intrinsic demands of clinical reasoning. This positions the tool as optimized for early skill acquisition before transitioning to higher-fidelity encounters.
3. Theoretical framework and learning mechanisms
The tool's effectiveness stems from integration of behaviorist principles (immediate environmental stimuli, reinforcement through scoring/feedback) and constructivist approaches (self-reflection, active knowledge construction through the critical thinking stages by Kamin et al. [19]). This dual approach explains observed improvements in both skill-based outcomes and higher-order cognitive abilities including clinical reasoning (p=0.0013) and problem-solving skills (p=0.0323).
4. Key strengths and educational innovations
The tool demonstrated substantial impact on self-efficacy development, with significant improvements in six of eight measured domains. Most notably, confidence in handling unfamiliar clinical presentations improved dramatically (p=0.0002). The positive but non-significant trends for communication skills and understanding of clinical presentations likely reflect the current limitations of text-based interaction.
The unanimous improvement across all nine feedback quality dimensions represents a fundamental advancement in medical education feedback delivery. This stems from several unique advantages: consistent detailed feedback for every practice session, structured approach using the ‘enhanced written’ format [21], and elimination of interpersonal dynamics that can compromise feedback quality [3]. The dual feedback mechanism combining immediate reflective discussion with delayed comprehensive reports caters to different learning styles and reinforces key concepts through multiple modalities, aligning with best practices emphasizing both immediacy and specificity [20]. These findings corroborate and extend recent international studies. While Holderried et al. [11] validated response accuracy and interrater reliability, and Rädel-Ablass et al. [12] demonstrated student acceptance across health disciplines, our study uniquely demonstrates comparative effectiveness against conventional methods through a multidimensional perception assessment. Furthermore, our dual feedback mechanism combining immediate reflective dialogue with comprehensive written reports, represents a pedagogical innovation not present in prior implementations.
The system’s scalability addresses critical challenges in medical education. Unlike SP programs requiring extensive recruitment, training, and coordination, the LLM-based system can be deployed instantly to unlimited students, particularly relevant given expanding medical school enrollments globally.
5. Limitations and challenges
Despite promising results, several important limitations must be acknowledged. From a pedagogical standpoint, the most significant technical limitation is the risk that the AI may model poor clinical habits or provide flawed feedback. This arises from the phenomenon of ‘hallucinations,’ where the system is incentivized by its training to generate confident-sounding falsehoods rather than express uncertainty [23]. While we grounded the system in expert-validated content to ensure clinical accuracy, this does not eliminate the model’s core behavioral tendency [24]. The high response latency noted by students in free-form feedback impacts interaction realism, though rapid technological advances are addressing this limitation. The text-based interface substantially deviates from real-world clinical settings where verbal communication, body language, and emotional cues play crucial roles in patient interaction. This limitation has also been acknowledged in previous studies by Holderried et al. [11] and Rädel-Ablass et al. [12]. Moreover, Vogel et al. [25] reported that non-verbal communication correlates significantly with empathy during history taking, whereas verbal communication does not, suggesting that empathy-related skills require explicit training that includes non-verbal cues. Accordingly, our tool should be interpreted as supporting history-taking structure and content, rather than comprehensive communication skills training.
A scope-related limitation concerns our comparative analysis. While SPs typically address multiple clinical skill domains including physical examination and professionalism, our tool focused exclusively on history-taking. Thus, the observed improvements (Tables 13) should be interpreted within this bounded context.
Methodological limitations also constrain our findings’ generalizability. The small sample size (n=21) from a single Korean medical institution limits broader applicability, particularly given specific cultural and educational contexts. The attrition rate from 34 initial participants to 21 completers may introduce selection bias, potentially overrepresenting motivated students more likely to benefit from self-directed learning tools. The brief 6-day study period provides only a snapshot of the tool’s potential impact, and the time gap between pre- and post-surveys may have influenced perception changes independent of the intervention itself.
Most critically, our focus on subjective perceptions without measuring objective performance metrics leaves unanswered whether improved satisfaction and self-efficacy translate to actual skill improvement. Furthermore, while dialogue metrics showed a strong correlation between conversational turns and assessment scores (r=0.66), this relationship warrants cautious interpretation. Because scoring was based on checklist completion, this correlation captures questioning quantity rather than clinical reasoning quality. While enhanced confidence and positive learning experiences are valuable outcomes, they do not guarantee competency in real clinical settings. The absence of objective assessments such as OSCE scores or standardized clinical performance evaluations represents a significant gap in validating educational effectiveness. The current system cannot replicate the complexity of human interaction, including managing difficult patients, recognizing subtle non-verbal cues, or developing the professional identity that emerges from authentic patient encounters.
6. Future directions and implementation strategies
Future research priorities include randomized controlled trials with objective performance metrics, long-term retention studies, and technical enhancements such as voice-based interaction and physical examination module integration [26]. Our observation that repeat attempts yielded both lower scores and reduced engagement (fewer turns and shorter duration) highlights the importance of incorporating strategies to sustain learner motivation, such as progressive difficulty, gamification elements, or explicit feedback on questioning depth. Furthermore, the per-session cost reported in our study could be further optimized in large-scale implementations by, for example, utilizing a database of pre-generated responses for recurring checklist items rather than generating each component on-demand. For curriculum integration, LLM-based tools should complement rather than replace traditional methods, serving primarily for skill building and confidence development before real patient encounters [27].
As API costs continue declining (the most recent gpt-5 model with higher performance and lower latency costs approximately one-seventh of that used in our study [28]), accessibility will increase substantially. Medical educators must develop guidelines ensuring AI-assisted learning enhances rather than compromises clinical competence and professional identity development [29]. Quality assurance mechanisms, regular content validation, and clear pedagogical frameworks will be essential for successful implementation.
7. Conclusion
Our study provides evidence that an LLM-based learning tool featuring VPs and VAs can significantly enhance medical students’ perceived learning experience in history-taking education while reducing costs and improving accessibility. The tool’s success validates the potential of AI-assisted medical education as a complementary training modality. While technical limitations and the need for objective performance validation remain important challenges, our findings support the value of LLM-based tools in addressing global challenges in clinical skills training. Future research must focus on validating these subjective improvements through objective performance metrics and determining optimal integration strategies that preserve the irreplaceable value of human interaction while leveraging the unique advantages of AI-assisted learning.
Supplementary files are available from https://doi.org/10.3946/kjme.2025.108
Supplement 1.
Example of Scoring Criteria.
kjme-2025-108-Supplement-1.pdf
Supplement 2.
Example of VP Case Design.
kjme-2025-108-Supplement-2.pdf
Supplement 3.
Example of Discussion (Immediate Feedback).
kjme-2025-108-Supplement-3.pdf
Supplement 4.
Examples of Delayed Feedback.
kjme-2025-108-Supplement-4.pdf
Supplement 5.
Reported Methods of Preparation for the CPX.
kjme-2025-108-Supplement-5.pdf

Data sharing statement

Please contact the corresponding author for data availability.

Acknowledgements

The authors acknowledge the valuable contributions of Drs. Bo Young Yoon, Sung Woo Cho, Hyung Koo Kang, Yoo Jin Lee, and Jong Wook Kim (Department of Internal Medicine), and Dr. Pamela Song (Department of Neurology), Inje University College of Medicine, whose expert review supported clinical validation of the tool.

Funding

This work was supported by the Inje University Research Grant and the National Research Foundation of Korea (NRF), funded by the Korean government (MSIT) (RS-2025-02214129).

Conflicts of interest

The authors declare the following potential non-financial conflict of interest: the study was conducted with student participants at the same institution where the authors are affiliated. To mitigate potential bias, participation was strictly voluntary and confidential, with no impact on academic assessment.

Author contributions

JB, HK, JL, JC, and SA conceptualized and designed the study. JB, HK, JL, and JC developed the survey and administered the study. HK developed the online tool. JB acquired the data and wrote the initial manuscript draft. JB and SA analyzed and interpreted the data. SA revised the draft. All authors read and approved the final submitted manuscript.

Fig. 1.
The large language model (LLM)-based learning tool interface and feedback mechanisms. (A) An example of a text-based interaction between a student doctor and the virtual patient, simulating an initial history-taking session for a 70-year-old woman presenting with knee pain. (B) An excerpt from the interactive dialogue where the virtual assessor guides a student through a structured self-reflection on their clinical reasoning immediately following a virtual patient encounter. (C) An excerpt of the comprehensive written feedback report, illustrating the structured format, automated scoring, and detailed, personalized feedback on clinical reasoning.
kjme-2025-108f1.jpg
Fig. 2.
Dialogue metrics and performance analysis. (A) Relationship between number of conversational turns and assessment score. (B) Relationship between session duration and assessment score. (C) Comparison of scores between first attempts and repeat attempts; diamonds indicate group means. (D) Paired comparison for students with multiple attempts; green lines indicate improvement, red lines indicate decline, diamonds represent group means.
kjme-2025-108f2.jpg
Table 1.
Comparison of Perceived Usability
Table 1.
Question SPs LLM-based tool p-value
I found it easy to focus on the patient. 3.29±0.85 3.86±0.91 0.0368
I felt comfortable practicing history taking. 2.95±0.80 4.57±0.60 0.0002

Data are presented as mean±standard deviation unless otherwise stated.

SP: Standard patient, LLM: Large language model.

Table 2.
Comparison of Perceived Self-efficacy
Table 2.
Question Conventional LLM-based tool p-value
My ability to perform clinical reasoning improved. 3.48±0.60 4.19±0.60 0.0013
My problem-solving skills were improved. 3.52±0.60 4.24±0.62 0.0323
My communication skills were improved. 3.38±0.74 3.76±1.00 0.0621a)
My understanding of the clinical presentations was improved. 3.62±0.67 4.38±0.59 0.0578
I became more confident in my history-taking abilities. 3.57±0.87 3.90±0.83 0.0006
I feel confident handling previously uncovered clinical presentations. 2.90±0.83 4.00±0.71 0.0002
My history-taking skills have improved. 3.81±0.40 3.95±0.80 0.0050
I feel confident about performing history taking in actual exams. 3.38±0.59 3.95±0.67 0.0016

Data are presented as mean±standard deviation unless otherwise stated.

LLM: Large language model.

a)Paired t-test used, Wilcoxon signed-rank test used otherwise.

Table 3.
Comparison of Perceived Feedback Quality
Table 3.
Question SPs LLM-based tool p-value
I received appropriate discussion topics regarding my results. 3.38±0.97 4.48±0.60 0.0007
I learned a significant amount of medical content. 3.43±0.60 4.24±0.70 0.0012a)
I received clear and specific feedback on areas needing improvement. 3.38±0.74 4.24±0.62 0.0012
My well-performed actions were sufficiently recognized. 3.52±0.81 4.43±0.68 0.0013
Feedback was well-balanced between areas for improvement and strengths. 3.52±0.81 4.38±0.59 0.0017
I was able to self-evaluate and improve. 3.76±0.54 4.43±0.60 0.0032
The feedback I received was specific and detailed. 3.24±1.04 4.43±0.75 0.0005
The feedback I received was tailored to my needs. 3.19±0.93 4.19±0.68 0.0001
Overall, I was satisfied with the feedback. 3.67±0.58 4.33±0.58 0.0039

Data are presented as mean±standard deviation unless otherwise stated.

SP: Standard patient, LLM: Large language model.

a)Paired-t test used, Wilcoxon signed-rank test used otherwise.

Table 4.
Overall Satisfaction of the Large Language Model Tool
Table 4.
Question Mean±SD
Using large language models for history taking was useful. 4.38±0.50
I was able to practice history taking more extensively than before. 4.19±0.60
The tool was more helpful for improving clinical skills than printed materials. 3.90±0.70
The tool was more motivating than learning from printed materials. 3.81±0.87
The tool was more helpful for acquiring medical knowledge than printed materials. 3.48±0.93
The tool’s responses were comparable to a human actor. 3.43±1.03

SD: Standard deviation.

Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:

Include:

Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions
Korean J Med Educ. 2026;38(1):64-73.   Published online February 20, 2026
Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:
Include:
Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions
Korean J Med Educ. 2026;38(1):64-73.   Published online February 20, 2026
Close

Figure

  • 0
  • 1
Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions
Image Image
Fig. 1. The large language model (LLM)-based learning tool interface and feedback mechanisms. (A) An example of a text-based interaction between a student doctor and the virtual patient, simulating an initial history-taking session for a 70-year-old woman presenting with knee pain. (B) An excerpt from the interactive dialogue where the virtual assessor guides a student through a structured self-reflection on their clinical reasoning immediately following a virtual patient encounter. (C) An excerpt of the comprehensive written feedback report, illustrating the structured format, automated scoring, and detailed, personalized feedback on clinical reasoning.
Fig. 2. Dialogue metrics and performance analysis. (A) Relationship between number of conversational turns and assessment score. (B) Relationship between session duration and assessment score. (C) Comparison of scores between first attempts and repeat attempts; diamonds indicate group means. (D) Paired comparison for students with multiple attempts; green lines indicate improvement, red lines indicate decline, diamonds represent group means.
Enhancing history-taking education through GPT-4-based virtual patients and automated assessment: a study of medical student perceptions
Question SPs LLM-based tool p-value
I found it easy to focus on the patient. 3.29±0.85 3.86±0.91 0.0368
I felt comfortable practicing history taking. 2.95±0.80 4.57±0.60 0.0002
Question Conventional LLM-based tool p-value
My ability to perform clinical reasoning improved. 3.48±0.60 4.19±0.60 0.0013
My problem-solving skills were improved. 3.52±0.60 4.24±0.62 0.0323
My communication skills were improved. 3.38±0.74 3.76±1.00 0.0621a)
My understanding of the clinical presentations was improved. 3.62±0.67 4.38±0.59 0.0578
I became more confident in my history-taking abilities. 3.57±0.87 3.90±0.83 0.0006
I feel confident handling previously uncovered clinical presentations. 2.90±0.83 4.00±0.71 0.0002
My history-taking skills have improved. 3.81±0.40 3.95±0.80 0.0050
I feel confident about performing history taking in actual exams. 3.38±0.59 3.95±0.67 0.0016
Question SPs LLM-based tool p-value
I received appropriate discussion topics regarding my results. 3.38±0.97 4.48±0.60 0.0007
I learned a significant amount of medical content. 3.43±0.60 4.24±0.70 0.0012a)
I received clear and specific feedback on areas needing improvement. 3.38±0.74 4.24±0.62 0.0012
My well-performed actions were sufficiently recognized. 3.52±0.81 4.43±0.68 0.0013
Feedback was well-balanced between areas for improvement and strengths. 3.52±0.81 4.38±0.59 0.0017
I was able to self-evaluate and improve. 3.76±0.54 4.43±0.60 0.0032
The feedback I received was specific and detailed. 3.24±1.04 4.43±0.75 0.0005
The feedback I received was tailored to my needs. 3.19±0.93 4.19±0.68 0.0001
Overall, I was satisfied with the feedback. 3.67±0.58 4.33±0.58 0.0039
Question Mean±SD
Using large language models for history taking was useful. 4.38±0.50
I was able to practice history taking more extensively than before. 4.19±0.60
The tool was more helpful for improving clinical skills than printed materials. 3.90±0.70
The tool was more motivating than learning from printed materials. 3.81±0.87
The tool was more helpful for acquiring medical knowledge than printed materials. 3.48±0.93
The tool’s responses were comparable to a human actor. 3.43±1.03
Table 1. Comparison of Perceived Usability

Data are presented as mean±standard deviation unless otherwise stated.

SP: Standard patient, LLM: Large language model.

Table 2. Comparison of Perceived Self-efficacy

Data are presented as mean±standard deviation unless otherwise stated.

LLM: Large language model.

a)Paired t-test used, Wilcoxon signed-rank test used otherwise.

Table 3. Comparison of Perceived Feedback Quality

Data are presented as mean±standard deviation unless otherwise stated.

SP: Standard patient, LLM: Large language model.

a)Paired-t test used, Wilcoxon signed-rank test used otherwise.

Table 4. Overall Satisfaction of the Large Language Model Tool

SD: Standard deviation.