Abstract
-
Purpose
Maintaining the quality of Single Best Answer (SBA) questions remains a challenge in medical education, especially as artificial intelligence (AI)-generated items become more common. While considerable attention has been paid to AI question generation, the vetting process is under-explored and difficult to scale.
-
Methods
This study investigates the feasibility and reliability of using a large language model to support the vetting of SBA questions. An AI-based reviewer, QA-bot, was developed using custom GPT and embedded with 25 criteria aligned with Bloom’s taxonomy (Levels 1–3). QA-bot and two experienced educators independently evaluated 32 AI-generated SBA questions using the shared evaluation rubric.
-
Results
The rubric showed high internal consistency (Cronbach’s alpha=0.878), and strong inter-rater reliability between human reviewers (intraclass correlation coefficient [ICC]=0.893). QA-bot demonstrated good alignment with human raters (ICC=0.861 and 0.840). While the AI performed well on objective, rule-based criteria, it was less consistent in detecting irrelevant complexity and accurately judging difficulty.
-
Conclusion
These findings suggest that AI can function as an efficient first-pass reviewer, improving consistency and reducing workload, with human oversight remaining essential for educational and clinical relevance.
-
Key Words: Assessment items, Artificial intelligence, Large language models, Question quality assurance, SBA questions
Introduction
Single Best Answer (SBA) questions are a widely adopted format in medical education, valued for their ability to assess clinical reasoning, diagnostic acumen and applied knowledge in a structured and scalable way [
1,
2]. Despite their widespread use, ensuring the quality and consistency of SBA questions remains a significant challenge, particularly as question banks grow in size and diversity. These challenges are magnified when items are co-authored across different institutions and subject areas, leading to variability in question quality and alignment with curricular objectives.
The process of reviewing and validating SBA questions, often referred to as question vetting, is critical to maintaining the integrity of assessments. This process typically involves human reviewers evaluating questions against established standards. However, this stage is frequently under-resourced, time consuming, and inconsistently applied. It requires expert input to identify common flaws such as item ambiguity, implausible distractors, and misalignment with learning outcomes [
3]. Given the increasing reliance on large-scale question banks to support both formative and summative assessments, conventional quality assurance (QA) approaches are becoming unsustainable.
Recent advances in artificial intelligence (AI), particularly in natural language processing and large language models (LLM), have streamlined the generation of assessment items [
4,
5]. We recognize that LLM are a subset of AI; however, for the purposes of this article, we use the terms AI and LLM interchangeably. This choice is intended to enhance readability and reflects the common usage of these terms in both academic literature and applied contexts, particularly within medical and educational domains.
While AI tools show promise in accelerating assessment workflow, their application has primarily focused on generating questions rather than vetting them. As AI-generated items enter question banks alongside human-authored content, the need for standardized, scalable QA tools becomes increasingly urgent. Without robust vetting mechanisms, both human and AI-generated questions risk undermining assessment validity due to issues such as factual inaccuracy, structural flaws, or cognitive mismatches.
Most existing AI-based educational tools either facilitate item generation [
4,
5] or offer descriptive analytics [
6], without engaging in deeper evaluation of item quality or alignment with pedagogical standards. This creates a critical gap in the assessment development process, where efficiency gains from AI in item creation are not matched by similar innovations in quality control.
To address the growing demand for scalable support in assessment QA, this study investigates whether LLM can feasibly and reliably assist in the vetting of SBA questions in medical education and explore the potential of AI to enhance and streamline the question review process. Specifically, we ask: To what extent can an LLM replicate human judgement when evaluating key quality elements of SBA items using a structured rubric?
Methods
To address the research question, we conducted a cross-sectional study in March 2025 comparing the evaluation of SBA questions by an AI-based QA-bot with those of experienced medical educators. The study examines the feasibility and reliability of using LLMs to support the vetting of SBA questions.
1. Question creation
A total of 32 SBA questions were generated using published LLM prompts, aligned with a predefined clinical domain such as Cardiovascular System, Renal System, Gastrointestinal System, Endocrine System and Neural System [
4]. These questions were designed to reflect varying levels of cognitive complexity and covered core content areas relevant to undergraduate medical curricula. The items adhered to the standard SBA format, consisting of a stem, a lead-in question, and five answer options [
1]. These questions can be found in
Supplement 1.
2. Evaluation rubric
The evaluation rubric was built based on local institutional standards and adapted from widely accepted item-writing guidelines [
1,
7], which included five domains: clarity of the question stem, quality of the lead-in, accuracy and relevance of content, plausibility of distractors, and cognitive level. The criteria were also aligned with Levels 1 to 3 of Bloom’s taxonomy [
7], encompassing knowledge recall, comprehension, and application. The evaluation rubric is available in
Supplement 2.
3. Expert review panel
Two experienced medical educators, with 5 and 15 years of experience in medical education, respectively, served as human reviewers. They independently evaluated the same set of questions using the shared rubric. Reviewers were blinded to the source of the questions and unaware of QA-bot’s evaluations. Reviewers used a simple scoring rubric: 1 for meets criteria, 0 for does not meet, and NA for not applicable.
4. AI-enabled QA-bot
We developed a customized QA-bot through custom GPT, a feature offered by OpenAI (
https://openai.com/) that allows the creation of tailored versions of ChatGPT without requiring coding or technical expertise. This interface allows for the integration of specific instructions, knowledge bases, and behavior settings that address domain-specific use cases [
8]. QA-bot is powered by GPT-4o and was developed to evaluate the quality of SBA questions. It assesses SBA question quality across 25 criteria, including the question stem, lead-in, answer options, and difficulty level [
9]. QA-bot generates a scoring rubric that is the same as that used by our human raters when a question is provided as input, helping to ensure consistency in evaluation. Details of the system prompt are available in
Supplement 3. QA-bot is accessible here:
https://chatgpt.com/g/g-67dd17793b588191b7a0f17817cb761a-sba-quality-checker.
5. Data analysis
To assess the level of agreement between QA-bot and human reviewers, the intraclass correlation coefficient (ICC) was calculated. A two-way random model with absolute agreement was used to evaluate overall scores across raters (human-human and human-AI) ensuring the degree of exact concordance while treating both raters as random samples to allow generalization [
10]. Internal consistency of the 25 criteria was assessed using Cronbach’s alpha to evaluate the reliability of the scoring framework. All statistical analyses were conducted using IBM SPSS Statistics ver. 29.0 (IBM Corp., Armonk, USA) with statistical significance set at p<0.05.
Results
The evaluation rubric demonstrated strong internal consistency, with a Cronbach’s alpha of 0.878, indicating reliable performance across the 25 quality criteria. Inter-rater agreement between the two medical educators was high (ICC=0.893), reflecting strong consistency in expert judgement.
QA-bot also showed good alignment with human reviewers (ICC between QA-bot and rater 1=0.861, ICC between QA-bot and rater 2=0.840), suggesting substantial agreement in rubric-based scoring and supporting the feasibility of AI-assisted vetting.
In terms of efficiency, QA-bot completed the full evaluation of 32 SBA questions in under 15 minutes. In contrast, each human reviewers required approximately 2 to 3 hours.
Discussion
QA-bot’s evaluations were generally consistent with those of the human raters, particularly in objective areas such as verb usage, spelling, and adherence to accepted clinical terminology. This strong agreement suggests that AI has potential to support routine QA tasks where structured, rule-based criteria are applied.
However, variability was observed in both human-to-human and human-to-AI comparisons, especially in domains requiring interpretive judgement, such as cognitive complexity and question relevance. Even among experienced educators, there were differences in scoring, reflecting the inherent subjectivity of educational assessment. These discrepancies highlight the difficulty of achieving complete consensus and point to the importance of using complementary human and AI approaches, having supported by clear rubrics and regular recalibration. For instance, an educator might first review a subset of items (such as a randomly selected percentage from each bulk upload) using a rubric and then compare or calibrate their judgments against the QA-bot’s output, or vice versa, using the tool to surface inconsistencies or overlooked issues before final validation.
Further differences emerged in more nuanced areas. Human reviewers were more likely to penalize questions that included unnecessary complexity or irrelevant detail, which the AI system did not consistently detect. Divergence was also seen in how difficulty levels were assessed. While human raters drew on contextual understanding and clinical reasoning, the AI occasionally misjudged the level of challenge, reflecting limitations in evaluating cognitive demands [
9].
These findings highlight the complementary roles of AI and human reviewers. At this point, AI tools appear to be more suited for tasks involving objective, clearly defined criteria. Therefore, they can function effectively as first-pass reviewers, identifying common or easily detectable issues. In contrast, human judgement remains essential for evaluating contextual relevance, pedagogical appropriateness, and clinical validity.
The significant time difference between AI and human reviewers highlights AI’s potential to improve scalability and reduce the resource demands of QA in medical assessment design. This complementary approach positions AI as a valuable aid rather than a replacement in the assessment development process. When integrated thoughtfully, AI can streamline item review, reduce workload, and improve consistency. At the same time, human oversight ensures that questions meet the deeper educational and clinical standards required for high-quality assessment [
11]. Ongoing refinement of AI tools, including improved rubric alignment and domain-specific training, may enhance their ability to handle more complex evaluative tasks over time.
1. Limitations
While this study highlights the potential of LLM to support the vetting of SBA questions, several limitations should be acknowledged. The evaluation focused exclusively on AI-generated questions across selected clinical domains, which may limit the generalizability of findings to other subject areas or question types. The two human reviewers, who had differing levels of experience, may have contributed to some variability in judgment. Expanding the reviewer pool in future work could help provide a broader basis for comparison. Although the rubric provided a structured framework for evaluation, it may not fully capture the subtleties of educational intent, or the interpretive reasoning applied by expert reviewers.
2. Conclusion
This study demonstrates that LLM can assist in the vetting of SBA questions, showing strong alignment with human reviewers on objective criteria. While AI offers clear benefits in efficiency and consistency, human oversight remains essential for evaluating clinical relevance and pedagogical appropriateness for SBA question QA. A combined approach, where AI serves as a first-pass reviewer and educators provide interpretive judgement, offers a practical, scalable solution for maintaining assessment quality in medical education.
Supplementary materials
Acknowledgements
None.
Funding
The authors declare that no funds, grants, or other financial support were received for the conduct of this study.
Conflicts of interest
No potential conflict of interest relevant to this article was reported.
Author contributions
ON conceptualized the study in collaboration with SPH and DHP. MHML and DHP contributed as medical educators supporting the evaluation process, with DHP also leading the development of the evaluation rubric. ON developed the AI-enabled QA-bot. Formal analysis was conducted by ON and DHP. ON prepared the first draft of the manuscript, with contributions from SPH. All authors reviewed and approved the final version of the manuscript.
Notes on contributors
Olivia Ng, PhD, is a lecturer in Medical Education, Assessment and Analytics at Lee Kong Chian School of Medicine (LKCMedicine), Nanyang Technological University, Singapore, with an interest in innovative assessment methods and technology-enhanced learning in medical education. Siew Ping Han, PhD, is a lecturer in Medical Education and Physiology at the Lee Kong Chian School of Medicine, Nanyang Technological University, Singapore. Her research interests include co-creation, student-staff partnerships, and technology-enhanced learning in medical education. Magdalene Hui Min Lee, MBBS, MMed (EM Med), is an emergency physician at Tan Tock Seng Hospital, Singapore. She also serves as an adjunct lecturer at LKC Medicine, Nanyang Technological University, Singapore. Her research interests include medical education assessment. Dong Haur Phua, MBBS, MMed (EM Med), MRCP Ed (A&E), FAMS, is an emergency physician and toxicologist at Tan Tock Seng Hospital, Singapore. He is also serving as assistant dean of Assessment at LKC Medicine, Nanyang Technological University, Singapore. His research interests include toxicology, cognitive errors, and education assessment.
References
- 1. Billings MS, DeRuchie K, Go S, et al. NBME Item Writing Guide. https://www.nbme.org/. 2021. Accessed May 10, 2025.2025.
- 2. Case SM, Swanson DB. Constructing written test questions for the basic and clinical sciences. Philadelphia, USA: National Board of Medical Examiners; 1998.
- 3. Downing SM. The effects of violating standard item writing principles on tests and students: the consequences of using flawed test items on achievement examinations in medical education. Adv Health Sci Educ Theory Pract. 2005;10(2):133-143. https://doi.org/10.1007/s10459-004-4019-5
- 4. Kıyak YS, Kononowicz AA. Case-based MCQ generator: a custom ChatGPT based on published prompts in the literature for automatic item generation. Med Teach. 2024;46(8):1018-1020. https://doi.org/10.1080/0142159X.2024.2314723
- 5. Lam G, Shammoon Y, Coulson A, et al. Utility of large language models for creating clinical assessment items. Med Teach. 2025;47(5):878-882. https://doi.org/10.1080/0142159X.2024.2382860
- 6. Ng O, Phua DH, Chu J, Wilding LVE, Mogali SR, Cleland J. Answering Patterns in SBA Items: Students, GPT3.5, and Gemini. Med Sci Educ. 2025;35(2):629-632. https://doi.org/10.1007/s40670-024-02232-4
- 7. Grainger R, Osborne E, Dai W, Kenwright D. The process of developing a rubric to assess the cognitive complexity of student-generated multiple choice questions in medical education. Asia Pac Sch. 2018;3(2):19-24. https://doi.org/10.29060/TAPS.2018-3-2/OA1049
- 8. Masters K, Benjamin J, Agrawal A, MacNeill H, Pillow MT, Mehta N. Twelve tips on creating and using custom GPTs to enhance health professions education. Med Teach. 2024;46(6):752-756. https://doi.org/10.1080/0142159X.2024.2305365
- 9. Herrmann-Werner A, Festl-Wietek T, Holderried F, et al. Assessing ChatGPT’s mastery of Bloom’s taxonomy using psychosomatic medicine exam questions: mixed-methods study. J Med Internet Res. 2024;26:e52113. https://doi.org/10.2196/52113
- 10. Koo TK, Li MY. A guideline of selecting and reporting intraclass correlation coefficients for reliability research. J Chiropr Med. 2016;15(2):155-163. https://doi.org/10.1016/j.jcm.2016.02.012
- 11. Ng O, Phua DH. Rethinking assessment in the context of AI. Med Teach. 2025;47(6):1056. https://doi.org/10.1080/0142159X.2024.2434105