Skip to main navigation Skip to main content
  • KSME
  • E-Submission

KJME : Korean Journal of Medical Education

OPEN ACCESS
ABOUT
BROWSE ARTICLES
FOR AUTHORS AND REVIEWERS

Articles

Original Research

The National Medical Admission Test and medical student selection: rethinking cutoff validity and fairness

Korean Journal of Medical Education 2026;38(1):74-81.
Published online: February 20, 2026

1Department of Clinical Epidemiology, College of Medicine, University of the Philippines Manila, Manila, Philippines

2Professional Regulatory Board of Medicine, Professional Regulation Commission, Manila, Philippines

3Department of Office of Research and Development, National Teacher Training Centre for the Health Profession, University of the Philippines Manila, Manila, Philippines

4Department of Health Policy and Administration, College of Public Health, University of the Philippines Manila, Manila, Philippines

5Research and Statistics Division, Professional Regulation Commission, Manila, Philippines

6SEAMEO Regional Center for Educational Innovations and Technology, Quezon City, Philippines

7Institute of Clinical Epidemiology, National Institutes of Health, University of the Philippines Manila, Manila, Philippines

Corresponding Author: Godofreda Ruiz Vergeire-Dalmacion (https://orcid.org/0000-0001-9212-9087) Department of Clinical Epidemiology, College of Medicine, University of the Philippines Manila, Room 103, Paz Mendoza Bldg. Pedro Gil Ermita Manila, 1000, Philippines Tel: +63.917.8408556 Fax: +63.85244098 Email: gvdalmacion1@up.edu.ph
• Received: June 6, 2025   • Revised: January 6, 2026   • Accepted: January 19, 2026

© The Korean Society of Medical Education.

This is an open-access article distributed under the terms of the Creative Commons Attribution Non-Commercial License (http://creativecommons.org/licenses/by-nc/3.0/) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited.

  • 1,235 Views
  • 33 Download
prev next
  • Purpose
    The Commission on Higher Education of the Philippines mandates a minimum National Medical Admission Test (NMAT) percentile rank—typically the 40th percentile—for medical school admission. However, percentile rank is cohort-dependent and varies in meaning across testing years. This study re-examined its validity as an admissions criterion and evaluated whether the NMAT General Performance Score (GPS), a standardized z-score, offers a more stable and valid basis for predicting academic performance.
  • Methods
    We conducted a retrospective analysis of 42,261 first-time Physician Licensure Examination (PLE) takers from 2012 to 2022. NMAT and PLE records were linked, and logistic regression models were used to assess associations between NMAT scores and PLE outcomes. Predictive performance was evaluated using both receiver operating characteristic (ROC–Youden) and Precision–Recall (PR–F1) analyses to identify optimal cutoffs.
  • Results
    Percentile ranks exhibited substantial year-to-year variability, with the same percentile corresponding to different GPS scores. A pooled GPS-to-percentile crosswalk is provided for interpretive reference but does not indicate fixed rank equivalence. In contrast, PR-F1 analysis of GPS appropriate for an imbalanced dataset showed consistent predictive validity (area under the ROC curve=0.918). The ROC–Youden index identified a cutoff at GPS=581, while the F1-optimized threshold was lower (GPS=377), favoring inclusivity. A midpoint cutoff (GPS=435) balanced stringency and access.
  • Conclusion
    The NMAT GPS is a more stable and equitable predictor of licensure performance than percentile rank. Its use may improve the fairness and consistency of medical school admissions and better align selection with long-term academic outcomes.
The National Medical Admission Test (NMAT) remains the principal gatekeeper for entry into Philippine medical schools. It consists of two domains: the ‘Aptitude Area,’ which assesses verbal, quantitative, inductive reasoning, and perceptual acuity; and the ‘Subject Area,’ covering biology, chemistry, physics, and the social sciences. Since 1985, it has been administered under the supervision of the Center for Educational Measurement, authorized by the Commission on Higher Education (CHED) [1]. The NMAT result is reported as a standardized NMAT General Performance Score (GPS) ranging from 200 to 800. At the same time, examinees are ranked using NMAT percentile rank, a percentage from 1 to 99+ that shows how well they performed relative to other test-takers and thus unstable and cohort-relative.
In the 1987–1988 school year, the Department of Education set the NMAT at the 40th percentile or higher for admission to the medical degree course [2]. However, compliance with this order was not maintained over the years due to its negative impact on medical student enrollment. The CHED currently uses the 40th percentile rank as the threshold for academic eligibility as provided in CHED Memorandum Order No. 18, Series of 2016 [3]. Yet, this benchmark has not been empirically validated against actual exam outcomes. Currently, individual medical schools are allowed to set higher, binding cutoff scores. Schools often use the 40th Percentile as a universal marker of readiness, but the absence of periodic validation raises concerns that potentially qualified applicants may be excluded. Furthermore, high cutoffs can increase inequity by favoring students from resource-rich institutions.
The Physician Licensure Examination (PLE) legally requires a passing average of 75%, with no grade below 50 in any subject, under the Philippine Medical Act of 1959. It is a high-stakes credentialing exam that grants its successful takers a license to practice medicine [4]. Concerns have been raised that excessively high cutoffs could disadvantage applicants from resource-limited backgrounds or schools outside metropolitan centers, thereby restricting diversity and inclusivity in medical education. This study aims (1) to evaluate the effectiveness of using the NMAT percentile rank of 40 in predicting PLE passers, (2) to compare cutoff determination methods using receiver operating characteristic (ROC)–Youden index versus Precision–Recall (PR)-F1 score, and (3) to identify cutoff thresholds balancing accuracy and fairness.
This study is grounded on predictive validity theory, a component of classical test theory that evaluates how well an assessment forecasts performance on a criterion measure, in this case, the PLE [4]. Policymakers need to consider the broader impact on the medical workforce, especially given declining interest in medical careers [5,6]. Several local and international studies have examined the correlation between medical school admission test scores and licensure exam performance, often using correlation and logistic regression [7-10]. However, the findings were inconsistent. A valid admission test should exhibit a moderate-to-strong correlation with later professional outcomes while ensuring fairness across subgroups. Guided by the equity–validity framework for high-stakes testing, the NMAT was examined as both a psychometric predictor and a policy instrument [11,12]. Overly stringent cutoffs may introduce construct-irrelevant variance, rejecting capable applicants for non-academic reasons, while overly lenient standards might weaken the link between selection and competence [13,14].
1. Study design and data source
A retrospective analytic study was conducted using a consolidated dataset of medical graduates who took the PLE between 2012 and 2022. The dataset linked each examinee’s NMAT performance to subsequent licensure outcomes, with data anonymized to ensure confidentiality. Ethical approval was obtained from the University of the Philippines Manila Research Ethics Board (No. UPMREB 2023-083501). The study complied with the principles of the Declaration of Helsinki regarding the use of secondary data.
2. Variables and definitions
The dependent variable was the PLE outcome, categorized as 1 for pass and 0 for fail. The principal independent variables were NMAT GPS and NMAT percentile rank. Because percentile rank varies by cohort, a pooled GPS-to-percentile rank crosswalk is provided in Supplement 1 for interpretive reference. To maintain cohort comparability, predictive analyses were conducted using standardized NMAT GPS scores. Meanwhile, the NMAT percentile rank was used in the model regression to reflect current admissions practice. This allows institutions to map GPS-based cutoffs onto their cohort-specific percentile distributions, without undermining the stability of the GPS-derived thresholds.
Demographic covariates included age group (<30 years, 30–39 years, and >39 years), sex at birth (male and female), relationship status (single and in a relationship), exam year (2012 to 2022), and type of medical school (public and private). Age referred to the examinee’s age at the time of the PLE, not at the time of NMAT administration. The cutoff at 30 years distinguished typical graduates (early to mid-20s) from older or delayed examinees, repeat takers, and career shifters, categories known to vary in academic continuity and test familiarity.
3. Statistical analysis
Means, standard deviations, medians, and interquartile ranges were calculated for continuous variables, while frequencies and percentages were calculated for categorical variables. Multicollinearity was formally assessed in Stata ver. 17.0 (Stata Corp., College Station, USA) using variance inflation factors (VIF). All calculated VIFs were less than 4 (mean VIF=2.28), confirming no significant collinearity (see Supplement 2 for details).
Logistic regression analysis is well-suited for setting cutoff scores of high-stakes exams, particularly for predicting pass/fail outcomes based on test performance [15]. Additionally, modified Poisson regression models were used to estimate relative risk ratios (RRRs) because the outcome variable, the number of passers in the 2012–2022 PLE, was common rather than rare [16]. Thus, multivariate logistic and modified Poisson regression analyses were performed to estimate the odds ratios (ORs) and RRRs, respectively, for the associations between PLE performance and the following study characteristics: age group, sex, relationship status, type of institution, exam year, and NMAT percentile rank. Effect size estimates were reported as crude and adjusted ORs and RRRs, with 95% confidence intervals (95% CIs), for associations between NMAT percentile rank and independent variables with PLE performance. Missing data were handled using complete-case analysis. The variable “type of pre-med courses” was excluded because it accounted for 11,666 missing values.
4. Determination of cutoff scores
Cutoff points for NMAT GPS were derived using two implementable statistical approaches: (1) PR analysis to identify the threshold maximizing the F1 score [2×precision×recall/(precision+recall)]; and (2) ROC analysis to identify the Youden index (maximum sensitivity+specificity–1).
While regression analysis is a specific algorithm for classification, PR is a general metric used to assess the performance of a classifier, especially when the dataset is imbalanced, as in the current study. On the other hand, the F1 score combines precision and recall to identify the threshold that maximizes this balanced measure, indicating the best cutoff point for the admission test [15,17].
In addition, the area under the ROC curve (AUC) is a single metric derived from the ROC curve that quantifies performance, with higher values closer to 1 indicating better model performance [18]. It is commonly used in medicine to assess the overall diagnostic accuracy of a test, particularly in radiology and laboratory medicine. If the F1 score identifies a threshold for the best cutoff of the PR curve, the Youden index is a specific, single-value summary measure of the ROC curve, calculated as maximum (sensitivity+specificity–1) [18]. In summary, PR, F1 score, Youden’s index, and ROC curves are metrics for evaluating classification models, but they differ in what they measure and in their optimal use cases. All analyses were performed using STATA ver. 17.0 (Stata Corp.).
1. Participant characteristics and predictors of PLE success
A total of 42,261 first-time PLE examinees were included, representing medical graduates who took the PLE from 2012 to 2022. The mean age at the time of PLE was 26.8±2.7 years (range, 21–42 years). Females constituted 62.1% of the sample, and 75.4% graduated from private medical schools. In addition, the median NMAT percentile rank was 71 (range, 48–89). Significant gender and age group differences were observed between PLE pass and fail groups (see Table 1 for the sample’s demographics and academic characteristics).
Adjusted effect estimates for the association between NMAT percentile rank and the likelihood of passing the PLE among first-time examinees were estimated using both logistic regression and modified Poisson regression models. In both models, NMAT percentile rank was positively associated with PLE success (Table 2). Specifically, the adjusted logistic regression showed that for every 1-point increase in NMAT percentile rank, the odds of passing the PLE increased by approximately 4.6% (adjusted OR, 1.046; 95% CI, 1.044–1.047; p<0.001). Similarly, the adjusted modified Poisson regression yielded a consistent finding, with a RRR of 1.006 (95% CI, 1.006–1.007; p<0.001). These results indicate that higher NMAT percentile ranks are independently associated with a greater likelihood of licensure success, reinforcing the test’s potential utility as a marker of academic readiness.
Although modest, these incremental effect estimates are meaningful at the population level, where even small predictive improvements can impact overall passing rates. Additionally, being 30 years or older or having studied at a private medical school were negatively associated with PLE success, while being male was positively associated. Notably, crude estimates indicated that years coinciding with the COVID-19 (coronavirus disease 2019) pandemic were associated with significantly lower outcomes than the reference year, 2012. The corresponding crude regression results are provided to enhance transparency and reproducibility (see Supplement 3 for details).
2. Optimal NMAT GPS/NMAT percentile rank cutoff
The estimated AUC for the PR analysis of NMAT GPS as predictor for passing the PLE is 0.918. Fig. 1 graphically illustrates the trade-off between precision, recall, and F1 score across various NMAT GPS cutoffs for predicting passing the PLE. The performance metrics are further detailed in Supplement 4. The highest F1 score (0.925) was observed at a GPS cutoff of 377, corresponding to the NMAT 5th percentile rank, indicating the point with the best overall balance between including those who will pass and excluding those who will not. At this threshold, recall was very high (97.5%), meaning almost all future PLE passers were included. However, precision was 87.9%, indicating that 12.1% of those admitted at this cutoff may not ultimately pass. The reader is reminded that even if GPS 377 aligns with the 5th percentile in pooled data, it corresponds to approximately the 15th percentile in 2022, reflecting cohort variation in score distribution.
A slightly higher cutoff of GPS 435 (10th percentile) offers a more conservative and pragmatic trade-off. It maintains a strong F1 score (0.915), with precision increasing to 89.3% (fewer unqualified students admitted) while recall slightly decreased to 93.7% (excluding 6.3% of actual passers compared with 2.5% with GPS=377). This balance may be preferable for institutions that prioritize minimizing false positives while still capturing the majority of capable students.
In contrast, the CHED recommends a cutoff at the 40th percentile (GPS=532). Although intended to reflect a national benchmark, our analysis suggests that this threshold is suboptimal in predictive performance. At GPS 532, recall drops to 65.8%, and although precision is high at 93.8%, it comes at the cost of excluding more than one-third of students who would likely pass the PLE (false negatives). This suggests that the 40th percentile rank cutoff, while strict, may inadvertently disenfranchise many capable students and may not be the most equitable or effective criterion for admission.
The ROC analysis in Table 3 indicates that NMAT exhibits good overall discriminative ability (AUC=0.777). The optimal cutoff was identified at approximately the 60th percentile (NMAT GPS=581), with a sensitivity of 0.677 and a false-positive rate of 0.276. However, with a sensitivity of 67.7%, this threshold would still exclude approximately one-third (32.3%) of students who are capable of passing the PLE. Meanwhile, 27.6% of those admitted at this threshold may not ultimately pass, reflecting the inherent trade-offs of relying solely on NMAT scores for admissions decisions. While ROC analysis is valuable in clinical settings, its application for determining admission cutoffs is limited not only by dataset imbalance but also by its limited applicability to decision thresholds in educational selection contexts, where missing competent students carries a heavier consequence.
We finally compared the performance of the two classification models: PR and ROC curves by graphically illustrating their respective cutoffs for determining the optimal NMAT admission threshold.
As shown in Fig. 2, both the ROC–Youden and PR–F1 analyses identify optimal NMAT thresholds that balance prediction accuracy and fairness in admission decisions. However, the two frameworks differ in their priorities. The ROC curve functions much like a diagnostic test, balancing the detection of true positives against the false positive rate, thereby emphasizing strict screening. In contrast, the PR curve focuses on recall, or the correct identification of true positives, able examinees, even if this increases the number of false positives. The PR–F1–based cutoff shown in Fig. 2A demonstrates superior inclusiveness, maintaining high predictive accuracy (F1=0.925) while reducing the exclusion of potentially competent students. In contrast, Fig. 2B shows the ROC–Youden analysis (AUC=0.777), with the optimal cutoff at GPS=581 (approximately the 60th percentile rank), with a sensitivity=0.691, a false-positive rate=0.291, and the optimum Youden index of 0.399, representing a stricter but less inclusive threshold. The ROC curve cutoff most closely corresponds to the traditional NMAT cutoff at the NMAT 40th percentile rank.
Finally, we generated a calibration plot comparing the predicted probabilities with the observed PLE pass rates in our study cohort. The predicted probabilities closely aligned with the observed outcomes, as indicated by a near-perfect calibration slope of 0.990 (see Supplement 5 for details).
Although regression models showed modest effect-size estimates, the findings are meaningful in large-scale admission systems, where even small predictive gains can influence aggregate success rates. The moderate but significant association between NMAT percentile rank and PLE outcomes reinforces the test’s utility as an indicator of academic readiness, though its limitations must be acknowledged. Reliance solely on NMAT GPS/NMAT percentile ranks, particularly when rigid cutoffs are imposed, risks excluding students who could otherwise achieve licensure-level competence.
The comparison between the ROC–Youden and PR–F1 frameworks demonstrates how analytic priorities shape admission policy decisions. ROC analysis, which aims to balance sensitivity and specificity, identified an optimal cutoff at the 60th percentile (NMAT GPS=581). This stricter threshold reduces the likelihood of admitting unqualified applicants (false-positive rate=29.1%) but increases the risk of excluding capable candidates (false-negative rate=30.9%).
In contrast, the PR–F1 approach prioritizes identifying as many qualified candidates as possible. This method suggested a more inclusive cutoff of GPS=377 (5th percentile rank), while maintaining high recall and minimizing the rejection of potentially competent students. On the other hand, GPS=435 (10th percentile rank) sets a more stringent threshold (excluding more marginal performers) while still capturing over 93% of likely PLE passers and maintaining a balanced, evidence-based standard. This makes it a strong candidate for institutions seeking a policy compromise that is more selective than GPS=377 but less exclusionary than GPS=462 or 481. Taken together, the ROC approach emphasizes caution and selectivity, whereas the PR–F1 framework values opportunity and fairness in access. These contrasting perspectives underscore that cutoff selection should align with the institution’s mission, whether by prioritizing risk minimization or by expanding opportunities for capable applicants.
These results demonstrate that statistical methods reflect underlying value judgments. The ROC–Youden approach aligns with a standardization-oriented model that prioritizes accuracy, whereas the PR–F1 method supports an equity-validity framework, which acknowledges that fairness and opportunity are integral to the interpretation of validity in educational assessment [18]. When the policy goal is to balance predictive precision with inclusiveness, the PR–F1 cutoff provides a more equitable, evidence-based benchmark for NMAT-based admissions particularly for imbalanced dataset [19].
The NMAT demonstrates consistent, though moderate, predictive validity for PLE performance. While a cutoff based on the ROC–Youden index (approximately GPS=581 or 60th percentile) optimizes the statistical balance between sensitivity and specificity, it carries the risk of excluding a substantial number of potentially qualified candidates. In contrast, the PR-F1 approach identifies a much lower threshold (GPS=377, 5th percentile) that maintains acceptable validity while significantly improving inclusiveness.
These contrasting thresholds are not merely technical; they reflect differing educational values. Institutions seeking to broaden opportunity and enhance social accountability may prefer the PR-F1-informed cutoff. Conversely, schools focused on academic efficiency and board exam performance may lean toward the ROC–Youden approach—provided they undertake periodic internal validation to ensure fairness.
Importantly, this study highlights that NMAT GPS, rather than NMAT percentile ranks, should be used to define admission cutoffs, as NMAT GPS is standardized and comparable across cohorts. In contrast, NMAT percentile ranks vary annually and are less stable.
We recommend that the CHED, while retaining NMAT as an admissions requirement, promote evidence-based flexibility in the application of score thresholds across institutions. Medical schools may be empowered to align NMAT GPS cutoffs with their specific missions, whether those emphasize academic performance, equitable access, or service to underserved populations.
Ultimately, periodic review and validation of NMAT thresholds are essential. Admission standards must serve not only as indicators of academic readiness but also as instruments of fairness, reflecting both the predictive validity framework and the broader goals of national medical education policy.
Supplementary files are available from https://doi.org/10.3946/kjme.2025.070
Supplement 1.
NMAT GPS to Percentile Rank Crosswalk (Official GPS).
kjme-2025-070-Supplement-1.pdf
Supplement 2.
Assessment of Multicollinearity Using VIF.
kjme-2025-070-Supplement-2.pdf
Supplement 3.
Crude (Unadjusted) Associations between Covariates, Including NMAT Percentile Rank, and the Likelihood of Passing the PLE among First-Time Takers.
kjme-2025-070-Supplement-3.pdf
Supplement 4.
Performance Metrics of Different Cutoffs of NMAT GPS versus Passing the PLE (2012–2022) Using Precision–Recall Analysis.
kjme-2025-070-Supplement-4.pdf
Supplement 5.
Calibration Plot Comparing Predicted Probabilities of Passing the PLE with Observed Outcomes Based on NMAT GPS Scores.
kjme-2025-070-Supplement-5.pdf

Data sharing statement

Please contact the corresponding author for data availability.

Acknowledgements

I am deeply grateful to Ms. Glenda H. Pedrosa for her generous assistance and invaluable guidance, especially in managing and setting up the dataset structure. I also thank Dr. Grace Aquiling-Dalisay, President of the Center for Educational Measurement Inc., and her hardworking team, particularly Ms. Armi S. Lantano, Research Section Head, for their cooperation. I would also like to express my gratitude for the support of Commissioner Jose Y. Cueto of the Professional Regulation Commission (PRC), as well as the assistance of other officials: Ms. Henrietta P. Narvaez, Chief of the Archives and Records Division; Ms. Gina A. Consignado, Director of the Information and Communication Technology Service; and Mr. Demosthenes N. Mistal, Chief of the Research and Statistics Division. The author also thanks the reviewers and editors of the Korean Journal of Medical Education for their thoughtful and constructive comments, which substantially enhanced the clarity and quality of this manuscript.

Funding

None.

Conflicts of interest

No potential conflict of interest relevant to this article was reported.

Author contributions

GRVD conceptualized and designed the protocol with guidance from ESB. GRVD led data acquisition. ALD assisted with data management and ROC analysis. GRVD analyzed and interpreted the data under ESB's supervision. GRVD drafted the initial manuscript. Each author contributed substantially to and revised the final version, and all have given explicit consent for publication. Each author takes responsibility for the content, having read and approved the final draft. ESB approved the final revision. GRVD is responsible for the overall content as the guarantor.

Fig. 1.
Trade-off between precision, recall, and F1 score at various National Medical Admission Test (NMAT) General Performance Score (GPS) cutoffs for Physician Licensure Examination passers of 2012–2022. The best performing NMAT GPS cutoff is 377 with precision: 87.9%, recall: 97.5%, and F1 score: 92.5%. A trade-off to balance precision and recall can be an NMAT GPS score cutoff of 435 (NMAT 10th percentile rank), with precision: 89.3%, recall: 93.7%, and F1: 91.5%. The Commission on Higher Education recommendation for the NMAT GPS score cutoff is 532 (NMAT 40th percentile rank), which falls within a low-performance zone across metrics, indicating poor balance between precision and recall.
kjme-2025-070f1.jpg
Fig. 2.
Comparison of precision–recall (PR-F1) and receiver operating characteristic (ROC)–Youden analyses for determining optimal National Medical Admission Test (NMAT) cutoff thresholds predicting passing the Physician Licensure Examination (PLE) (2012–2022). (A) Precision–recall curve (area under the ROC curve [AUC]=0.918) shows the highest F1 score (0.925) at a General Performance Score (GPS) cutoff of 377 (5th percentile), balancing precision (0.879), and recall (0.975). (B) ROC curve (AUC=0.777) identifies the 60th percentile (GPS=581) as optimal via the Youden index (0.399), with sensitivity=0.691 and false positive rate=0.291.
kjme-2025-070f2.jpg
Table 1.
Study Characteristics of First-Time PLE Takers, Stratified by PLE Performance (N=42,261)
Table 1.
Characteristic Total PLE passed PLE failed
Total 42,261 (100.0) 36,222 (85.7) 6,039 (14.3)
Age group (yr)
 <30 40,758 (96.4) 35,113 (97.0) 5,625 (93.2)
 30–39 236 (0.6) 126 (0.4) 110 (1.8)
 >39 1,264 (3.0) 961 (2.6) 303 (5.0)
 Missing 3 2 1
Sex
 Female 26,248 (62.1) 22,246 (61.4) 4,002 (66.3)
 Male 16,013 (37.9) 13,976 (38.6) 2.037 (33.7)
Relationship status
 Single 40,458 (95.7) 34,652 (95.7) 5,806 (96.1)
 In a relationship 1,803 (4.3) 1,570 (4.3) 233 (3.9)
Type of institution
 Public 9,886 (24.6) 9,005 (26.0) 881 (15.7)
 Private 30,386 (75.4) 25,653 (24.0) 4,733 (84.3)
 Missing 1989 1564 425
Exam year
 2012 1,664 (3.9) 1,515 (4.2) 149 (2.5)
 2013 2,070 (4.9) 1,944 (5.4) 126 (2.1)
 2014 2,643 (6.2) 2,439 (6.7) 204 (3.4)
 2015 2,940 (7.0) 2,773 (7.7) 167 (2.8)
 2016 3,796 (9.0) 3,230 (8.9) 566 (9.4)
 2017 4,008 (9.5) 3,562 (9.8) 446 (7.4)
 2018 4,749 (11.2) 4,193 (11.6) 556 (9.2)
 2019 4,947 (11.7) 4,514 (12.5) 433 (7.2)
 2020 4,907 (11.6) 4,062 (11.2) 845 (14.0)
 2021 4,165 (9.9) 3,345 (9.2) 820 (13.6)
 2022 6,372 (15.1) 4,645 (12.8) 1,727 (28.6)
General percentage average, PLE
 Mean±SD 78.8±5.2 80.3±3.1 69.2±4.7
 Median (IQR) 79.7 (76.3–82.2) 80.3 (78.0–82.6) 70.5 (67.2–72.3)
NMAT average score
 Mean±SD 556.9±98.2 571.1±92.1 472.0±90.4
 Median (IQR) 556 (494–624) 568 (510–633) 480 (406–533)
NMAT percentile rank
 Mean±SD 66.0±26.6 69.9±24.4 42.0±26.9
 Median (IQR) 71 (48–89) 75 (54–91) 42 (17–63)

Values are presented as number (%), mean±SD, or median (IQR) unless otherwise stated.

PLE: Physician Licensure Examination, SD: Standard deviation, IQR: Interquartile range.

Table 2.
Adjusted Logistic and Modified Poisson Regression Estimates of the Association between NMAT Percentile Rank and PLE Performance among First-Time Examinees
Table 2.
Characteristic Adjusted OR (95% CI)a) p-value Adjusted RRR (95% CI)a) p-value
NMAT percentile rank 1.046 (1.044–1.047) <0.001 1.006 (1.006–1.007) <0.001

NMAT: National Medical Admission Test, PLE: Physician Licensure Examination, OR: Odds ratio, RRR: Relative risk ratio, CI: Confidence interval.

a)Adjusted ORs and RRRs with 95% CIs are presented. Models controlled for age category, sex at birth, type of medical school, exam year, and relationship status. Both models demonstrated a positive association between NMAT percentile rank and the likelihood of passing the PLE (p<0.001).

Table 3.
ROC Curve Metrics by NMAT GPS/NMAT Percentile Rank
Table 3.
Cut off percentile rank Sensitivity Specificity Youden FPR
5 0.996 0.053 0.049 0.947
10 0.977 0.183 0.160 0.817
20 0.957 0.271 0.228 0.729
30 0.924 0.366 0.290 0.634
40 0.871 0.470 0.341 0.53
50 0.788 0.599 0.386 0.402
60 0.691 0.709 0.399 0.291
70 0.577 0.811 0.388 0.189
80 0.444 0.895 0.339 0.105
90 0.276 0.959 0.235 0.041
100 0.02 0.999 0.022 0.001

Optimal ROC cutoff: Approximately the NMAT 60th percentile rank (NMAT GPS=581) with the highest Youden index (0.399), sensitivity (69.1%), and FPR (29.1%). Data validated by 1,000-sample bootstrap; 95% CI for area under the ROC curve=0.77–0.78 and Youden index=0.399.

ROC: Receiver operating characteristic, NMAT: National Medical Admission Test, GPS: General Performance Score, FPR: False-positive rate, CI: Confidence interval.

Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:

Include:

The National Medical Admission Test and medical student selection: rethinking cutoff validity and fairness
Korean J Med Educ. 2026;38(1):74-81.   Published online February 20, 2026
Download Citation

Download a citation file in RIS format that can be imported by all major citation management software, including EndNote, ProCite, RefWorks, and Reference Manager.

Format:
Include:
The National Medical Admission Test and medical student selection: rethinking cutoff validity and fairness
Korean J Med Educ. 2026;38(1):74-81.   Published online February 20, 2026
Close

Figure

  • 0
  • 1
The National Medical Admission Test and medical student selection: rethinking cutoff validity and fairness
Image Image
Fig. 1. Trade-off between precision, recall, and F1 score at various National Medical Admission Test (NMAT) General Performance Score (GPS) cutoffs for Physician Licensure Examination passers of 2012–2022. The best performing NMAT GPS cutoff is 377 with precision: 87.9%, recall: 97.5%, and F1 score: 92.5%. A trade-off to balance precision and recall can be an NMAT GPS score cutoff of 435 (NMAT 10th percentile rank), with precision: 89.3%, recall: 93.7%, and F1: 91.5%. The Commission on Higher Education recommendation for the NMAT GPS score cutoff is 532 (NMAT 40th percentile rank), which falls within a low-performance zone across metrics, indicating poor balance between precision and recall.
Fig. 2. Comparison of precision–recall (PR-F1) and receiver operating characteristic (ROC)–Youden analyses for determining optimal National Medical Admission Test (NMAT) cutoff thresholds predicting passing the Physician Licensure Examination (PLE) (2012–2022). (A) Precision–recall curve (area under the ROC curve [AUC]=0.918) shows the highest F1 score (0.925) at a General Performance Score (GPS) cutoff of 377 (5th percentile), balancing precision (0.879), and recall (0.975). (B) ROC curve (AUC=0.777) identifies the 60th percentile (GPS=581) as optimal via the Youden index (0.399), with sensitivity=0.691 and false positive rate=0.291.
The National Medical Admission Test and medical student selection: rethinking cutoff validity and fairness
Characteristic Total PLE passed PLE failed
Total 42,261 (100.0) 36,222 (85.7) 6,039 (14.3)
Age group (yr)
 <30 40,758 (96.4) 35,113 (97.0) 5,625 (93.2)
 30–39 236 (0.6) 126 (0.4) 110 (1.8)
 >39 1,264 (3.0) 961 (2.6) 303 (5.0)
 Missing 3 2 1
Sex
 Female 26,248 (62.1) 22,246 (61.4) 4,002 (66.3)
 Male 16,013 (37.9) 13,976 (38.6) 2.037 (33.7)
Relationship status
 Single 40,458 (95.7) 34,652 (95.7) 5,806 (96.1)
 In a relationship 1,803 (4.3) 1,570 (4.3) 233 (3.9)
Type of institution
 Public 9,886 (24.6) 9,005 (26.0) 881 (15.7)
 Private 30,386 (75.4) 25,653 (24.0) 4,733 (84.3)
 Missing 1989 1564 425
Exam year
 2012 1,664 (3.9) 1,515 (4.2) 149 (2.5)
 2013 2,070 (4.9) 1,944 (5.4) 126 (2.1)
 2014 2,643 (6.2) 2,439 (6.7) 204 (3.4)
 2015 2,940 (7.0) 2,773 (7.7) 167 (2.8)
 2016 3,796 (9.0) 3,230 (8.9) 566 (9.4)
 2017 4,008 (9.5) 3,562 (9.8) 446 (7.4)
 2018 4,749 (11.2) 4,193 (11.6) 556 (9.2)
 2019 4,947 (11.7) 4,514 (12.5) 433 (7.2)
 2020 4,907 (11.6) 4,062 (11.2) 845 (14.0)
 2021 4,165 (9.9) 3,345 (9.2) 820 (13.6)
 2022 6,372 (15.1) 4,645 (12.8) 1,727 (28.6)
General percentage average, PLE
 Mean±SD 78.8±5.2 80.3±3.1 69.2±4.7
 Median (IQR) 79.7 (76.3–82.2) 80.3 (78.0–82.6) 70.5 (67.2–72.3)
NMAT average score
 Mean±SD 556.9±98.2 571.1±92.1 472.0±90.4
 Median (IQR) 556 (494–624) 568 (510–633) 480 (406–533)
NMAT percentile rank
 Mean±SD 66.0±26.6 69.9±24.4 42.0±26.9
 Median (IQR) 71 (48–89) 75 (54–91) 42 (17–63)
Characteristic Adjusted OR (95% CI)a) p-value Adjusted RRR (95% CI)a) p-value
NMAT percentile rank 1.046 (1.044–1.047) <0.001 1.006 (1.006–1.007) <0.001
Cut off percentile rank Sensitivity Specificity Youden FPR
5 0.996 0.053 0.049 0.947
10 0.977 0.183 0.160 0.817
20 0.957 0.271 0.228 0.729
30 0.924 0.366 0.290 0.634
40 0.871 0.470 0.341 0.53
50 0.788 0.599 0.386 0.402
60 0.691 0.709 0.399 0.291
70 0.577 0.811 0.388 0.189
80 0.444 0.895 0.339 0.105
90 0.276 0.959 0.235 0.041
100 0.02 0.999 0.022 0.001
Table 1. Study Characteristics of First-Time PLE Takers, Stratified by PLE Performance (N=42,261)

Values are presented as number (%), mean±SD, or median (IQR) unless otherwise stated.

PLE: Physician Licensure Examination, SD: Standard deviation, IQR: Interquartile range.

Table 2. Adjusted Logistic and Modified Poisson Regression Estimates of the Association between NMAT Percentile Rank and PLE Performance among First-Time Examinees

NMAT: National Medical Admission Test, PLE: Physician Licensure Examination, OR: Odds ratio, RRR: Relative risk ratio, CI: Confidence interval.

a)Adjusted ORs and RRRs with 95% CIs are presented. Models controlled for age category, sex at birth, type of medical school, exam year, and relationship status. Both models demonstrated a positive association between NMAT percentile rank and the likelihood of passing the PLE (p<0.001).

Table 3. ROC Curve Metrics by NMAT GPS/NMAT Percentile Rank

Optimal ROC cutoff: Approximately the NMAT 60th percentile rank (NMAT GPS=581) with the highest Youden index (0.399), sensitivity (69.1%), and FPR (29.1%). Data validated by 1,000-sample bootstrap; 95% CI for area under the ROC curve=0.77–0.78 and Youden index=0.399.

ROC: Receiver operating characteristic, NMAT: National Medical Admission Test, GPS: General Performance Score, FPR: False-positive rate, CI: Confidence interval.