Objective: The objective of this study was to evaluate the performance of ChatGPT-4o on the Korean Physical TherapistLicensing Examination(KPTLE) and to explore its potential as a supplementary tool in physical therapy education and assessment.
Design...
Objective: The objective of this study was to evaluate the performance of ChatGPT-4o on the Korean Physical TherapistLicensing Examination(KPTLE) and to explore its potential as a supplementary tool in physical therapy education and assessment.
Design: Quantitative experimental study using publicly available national examination data.
Methods: ChatGPT-4o was tested on 960 multiple-choice questions from the 47th to 51st KPTLE, withoutany provision of sourceor domain-specific information. Correct answer rates and pass/fail status were assessed. Additionally, its performance wascompared with that of examinees who passed the same examinations.Statistical analyses included one-sample t-tests, effect sizecalculationsusing Cohen’s d. and percentile rankings.
Results: ChatGPT-4o achieved an average correct answer rate of 88.9%, consistently exceeding the passing criteria across allyears. Compared to students who passed the same examinations, ChatGPT-4o performed significantly better in overall andsubject-specific scores (p<0.001), with large effect sizes (Cohen’s d>0.8) and top percentile rankings. Although its performancein the medical law section was relatively poor, the overall results indicated stable and strong performance.
Conclusions: ChatGPT-4o demonstrated sufficient domain knowledge to pass the written KPTLE often surpassing humanexaminee performance in standardized multiple-choice formats. These findings suggest the potential for its use in theoreticaleducation and physical therapy assessments. Additional, further research is required to assess its applicability in testing practicalskills and real-world clinical environments, particularly given the limitations inherent in current language models.