Classification accuracy, which can be considered as the validity of criterion-referenced test, is a useful and essential information in test development and utilization of results, as it represents the accuracy of classifying examinees based on their ...
Classification accuracy, which can be considered as the validity of criterion-referenced test, is a useful and essential information in test development and utilization of results, as it represents the accuracy of classifying examinees based on their achievement levels. As a result, several methods have been proposed to estimate classification accuracy and particularly, research on classification accuracy estimation using IRT has been actively conducted since the 2000s. According to previous studies, it has been observed that the estimation of IRT classification accuracy is influenced by several factors associated with the test. Additionally, examinees’ ability distribution has also been identified as another influencing factor in this regard. In practice, the assumption of normality for the examinees’ ability distribution is often inappropriate in educational evaluation or psychometric data.
Despite the frequent occurrence of violations of the normality assumption regarding the examinees’ ability distribution, which can impact the estimation of classification accuracy, there has been a lack of systematic attempts to investigate the influence of various examinees’ ability distributions on the IRT classification accuracy estimation methods. In particular, the Guo method, which is difficult to find previous studies on, is also susceptible to the non-normality of examinees’ ability distribution. The Guo method assumes a non-informative prior distribution when estimating the posterior distribution, which represents the latent distribution of examinee. This makes it difficult to reflect the non-normality of examinees’ ability distribution. Furthermore, IRT classification accuracy estimation methods, including the Guo method, require accurate parameter estimation. To achieve this the estimation of parameters can be performed using the MMLE-EM method, and in MMLE, the analysis is typically conducted under the assumption that the latent distribution follows a normal distribution. However, the assumption of normality for the examinees’ ability distribution is often inappropriate. In such cases, assuming that the latent distribution follows a normal distribution can increase bias in parameter estimation and hinder accurate estimation of classification accuracy.
In other words, it was necessary to verify whether accurate classification accuracy estimation is possible through the IRT classification accuracy estimation methods in cases where the assumption of normality is inappropriate. Moreover, in MMLE, it is possible to estimate the latent distribution along with parameter estimation in order to reduce the bias in parameter estimation. It was expected that by using the estimated latent distribution to estimate classification accuracy, it would better reflect the non-normality of the examinees’ ability distribution compared to the Guo method, which assumes a non-informative prior distribution.
In this study, a new method (referred to as the 2NM method) for estimating classification accuracy was proposed by using the the estimated latent distribution obtained through the latent distribution estimation method assuming the two-component normal mixture distribution.
The purpose of the simulation study was to examine the accuracy of the 2NM-latent distribution estimation, which is a prerequisite for utilizing the 2NM method. Additionally, the performance of the 2NM method was compared to that of the existing IRT classification accuracy estimation methods under the simulation conditions. To this end, the simulation study was designed with various conditions including examinees’ ability distribution, test length, sample size, ability estimation method, and cut-score location. The goal was to examine the accuracy of the 2NM-latent distribution estimation method based on different examinees' ability distributions and test lengths. Additionally, the study aimed to investigate the differences in classification accuracy estimates obtained from three methods based on variations in the examinees’ ability distribution, cut-score location, test length, and ability estimation method. Furthermore, to assess the stability and accuracy of the three classification accuracy estimation methods, the study calculated the standard error of estimate, bias, and root mean square error based on the true classification accuracy index under different simulation conditions.
The results and conclusions of this study are as follows : First, 2NM-latent distribution estimation method can handle not only normal distribution but also various non-normal distributions with flexibility. It has demonstrated stable and accurate estimation of the latent distribution. Secondly, in non-normal situations where skewness and bimodality are present, the 2NM method has been shown to outperform other methods. Thirdly, the difference between the 2NM method and other methods became even more pronounced when the test length was shorter. Lastly, the 2NM method did not show any significant difference in ability estimation method, whereas the choice of ability estimation method should be carefully considered for the other methods depending on the situation.