Purpose This study investigated the relationship between the item response time (iRT) and classic item analysis indicators obtained from computer-based test (CBT) results and deduce students’ problem-solving behavior using the relationship.
Methods We retrospectively analyzed the results of the Comprehensive Basic Medical Sciences Examination conducted for 5 years by a CBT system in Dankook University College of Medicine. iRT is defined as the time spent to answer the question. The discrimination index and the difficulty level were used to analyze the items using classical test theory (CTT). The relationship of iRT and the CTT were investigated using a correlation analysis. An analysis of variance was performed to identify the difference between iRT and difficulty level. A regression analysis was conducted to examine the effect of the difficulty index and discrimination index on iRT.
Results iRT increases with increasing difficulty index, and iRT tends to decrease with increasing discrimination index. The students’ effort is increased when they solve difficult items but reduced when they are confronted with items with a high discrimination. The students’ test effort represented by iRT was properly maintained when the items have a ‘desirable’ difficulty and a ‘good’ discrimination.
Conclusion The results of our study show that an adequate degree of item difficulty and discrimination is required to increase students’ motivation. It might be inferred that with the combination of CTT and iRT, we can gain insights about the quality of the examination and test behaviors of the students, which can provide us with more powerful tools to improve them.
Citations
Citations to this article as recorded by
Conditional Dependencies Between Response Time and Item Discrimination: An Item-Level Meta-Analysis Joshua B. Gilbert, William S. Young, Zachary Himmelsbach, Esther Ulitzsch, Benjamin W. Domingue Educational and Psychological Measurement.2026;[Epub] CrossRef
Comparison of the response time-based effort-moderated IRT model and three-parameter logistic model according to computerized adaptive test performances: a simulation study Yusuf Kemal Arslan, Afra Alkan, Atilla Halil Elhan Communications in Statistics - Simulation and Computation.2025; 54(1): 44. CrossRef
Toward sustainable farming: Assessing and validating green skills for agricultural professionals in China Bowei Hu, Khahan Na-Nan, Yotsaphat Kittichotsatsawat Environmental Challenges.2025; 18: 101067. CrossRef
Knowledge, attitude and practice towards multiple myeloma among medical staff in Enshi Region Luping Zou, Jinhua Li, Hang Xiang, Jun Tan, Yan Zeng Scientific Reports.2025;[Epub] CrossRef
The Competent Computational Thinking Test (cCTt): A Valid, Reliable and Gender-Fair Test for Longitudinal CT Studies in Grades 3–6 Laila El-Hamamsy, María Zapata-Cáceres, Estefanía Martín-Barroso, Francesco Mondada, Jessica Dehler Zufferey, Barbara Bruno, Marcos Román-González Technology, Knowledge and Learning.2025; 30(3): 1607. CrossRef
Comparison of a generative large language model to pharmacy student performance on therapeutics examinations Christopher J. Edwards, Bernadette Cornelison, Brian L. Erstad Currents in Pharmacy Teaching and Learning.2025; 17(9): 102394. CrossRef
Designing Biochemical Visual Literacy Assessments: Insights from Classroom Testing and Student Interviews Kristen Procko, Josh T. Beckham, Roderico Acevedo, Swati Agrawal, Shane Austin, Charmita Burch, Shelly Engelman, Kristin M. Fox, Lauren A. Genova, Pamela S. Mertz, Rachel M. Mitton-Fry, Didem Vardar-Ulu Journal of Chemical Education.2025; 102(12): 5045. CrossRef
The impact of repeated item development training on the prediction of medical faculty members’ item difficulty index Hye Yoon Lee, So Jung Yune, Sang Yeoup Lee, Sunju Im, Bee Sung Kam BMC Medical Education.2024;[Epub] CrossRef
Identification of parameters for electronic distance examinations Robin Richter, Andrea Tipold, Elisabeth Schaper Frontiers in Veterinary Science.2024;[Epub] CrossRef
Validation of the Perceived Islamophobia Scale (PIS) among Muslims living in the United States Khulud Almutairi, Salman Shaheen Ahmad, Merranda Marie McLaughlin, Karina Gattamorta, Amy Weisman de Mamani Social Sciences & Humanities Open.2024; 10: 101054. CrossRef
Analysis of Nutrition Knowledge After One Year of Intervention in a National Extracurricular Athletics Program: A Cross-Sectional Study with Pair-Matched Controls of Polish Adolescents Dominika Skolmowska, Dominika Głąbska, Dominika Guzek, Jakub Grzegorz Adamczyk, Hanna Nałęcz, Blanka Mellová, Katarzyna Żywczyk, Krystyna Gutkowska Nutrients.2024; 17(1): 64. CrossRef
Differences in Multiple-Choice Questions of Opposite Stem Orientations Based on a Novel Item Quality Measure Samuel Olusegun Adeosun American Journal of Pharmaceutical Education.2023; 87(2): ajpe8934. CrossRef
Examination of response time effort in TIMSS 2019: Comparison of Singapore and Türkiye Esin YILMAZ KOĞAR, Sümeyra SOYSAL International Journal of Assessment Tools in Education.2023; 10(Special Is): 174. CrossRef
The development and validation of a questionnaire to assess relative energy deficiency in sport (RED-S) knowledge Namratha N. Pai, Rachel C. Brown, Katherine E. Black Journal of Science and Medicine in Sport.2022; 25(10): 794. CrossRef
Comparing the psychometric properties of two primary school Computational Thinking (CT) assessments for grades 3 and 4: The Beginners' CT test (BCTt) and the competent CT test (cCTt) Laila El-Hamamsy, María Zapata-Cáceres, Pedro Marcelino, Barbara Bruno, Jessica Dehler Zufferey, Estefanía Martín-Barroso, Marcos Román-González Frontiers in Psychology.2022;[Epub] CrossRef
Evaluation of usefulness of smart device-based testing: a survey study of Korean medical students Youngsup Christopher Lee, Oh Young Kwon, Ho Jin Hwang, Seok Hoon Ko Korean Journal of Medical Education.2020; 32(3): 213. CrossRef
Effect of Smart Device Ability on the Smart Device-Based Testing National Board Examination for Optometry Students Eun Joo Kim, Koon-Ja Lee, Jung Un Jang The Korean Journal of Vision Science.2019; 21(4): 631. CrossRef
PURPOSE In 2002, extended-matching type (R-type) items were introduced to the Korean Medical Licensing Examination. To evaluate the usability of R-type items, the results of the Korean Medical Licensing Examination in 2002 and 2003 were analyzed based on item types and knowledge levels. METHODS: Item parameters, such as difficulty and discrimination indexes, were calculated using the classical test theory.
The item parameters were compared across three item types and three knowledge levels. RESULTS: The values of R-type item parameters were higher than those of A- or K-type items. There was no significant difference in item parameters according to knowledge level, including recall, interpretation, and problem solving. The reliability of R-type items exceeded 0.99. With the R-type, an increasing number in correct answers was associated with a decreasing difficulty index. CONCLUSION: The introduction of R-type items is favorable from the perspective of item parameters.
However, an increase in the number of correct answers in pick 'n'-type questions results in the items being more difficult to solve.
Citations
Citations to this article as recorded by
Reforms of the Korean Medical Licensing Examination regarding item development and performance evaluation Mi Kyoung Yim Journal of Educational Evaluation for Health Professions.2015; 12: 6. CrossRef