The purpose of this study is to examine the accuracy of variable selection methods used in machine-learning–based predictive analyses of educational panel data. Educational panel data have been widely used to predict various outcomes, such as academ...
The purpose of this study is to examine the accuracy of variable selection methods used in machine-learning–based predictive analyses of educational panel data. Educational panel data have been widely used to predict various outcomes, such as academic achievement, career decisions, and participation in private tutoring, by employing machine-learning–based predictive models. In this line of research, selected predictors are often interpreted to derive educational implications and policy recommendations. Penalized regression models, particularly LASSO, have been extensively adopted in educational studies because they simultaneously perform coefficient estimation and variable selection.
However, high predictive performance does not necessarily guarantee that the selected variables are truly associated with the outcome. Previous studies have shown that penalized regression models may over-select noise variables, which can undermine the interpretability of the selected variables. In this context, variables such as suppressor variables can be selected due to their associations with other predictors, potentially improving prediction accuracy without a direct association with the outcome.
Motivated by these concerns, this study shifts the focus from predictive performance to variable selection accuracy, examining how precisely machine-learning–based variable selection methods identify true predictors and how effectively they control the selection of noise variables. To this end, the variable selection accuracy of penalized regression models—LASSO, Elastic Net, and Adaptive Lasso—is compared with that of error-control–based methods, namely Derandomized knockoff (DRM) and Multiple Data Splitting (MDS). While penalized regression models prioritize predictive accuracy in the variable selection process, error-control–based methods explicitly regulate selection errors.
In the simulation study, experimental conditions included sample size, the number of predictors, and the degree of imbalance in a binary outcome variable. Simulation data were generated to reflect key structural characteristics of educational panel data, such as Likert-scale item bundles, complex correlation structures among predictors, and imbalanced binary outcomes. Simulation Study 1 varied sample size (500, 2,000, 5,000), number of predictors (50, 200, 500), and imbalance ratios (1:1, 1:3, 1:9), and evaluated variable selection accuracy using statistical power as an index of true variable identification and the false discovery rate (FDR) as an index of noise variable control.
Simulation Study 2 focused on the most severely imbalanced condition (1:9) and examined the effects of oversampling techniques—Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN)—on variable selection accuracy.
Finally, to assess whether the simulation findings generalize to real-world contexts, we conducted analyses using data from the Korean Education Longitudinal Study 2013, focusing on the 4th- and 5th-wave data of middle school students in Grades 8 and 9. Variable selection stability across penalized regression models and error-control–based methods was examined according to sample size, imbalance ratio, and imbalance data preprocessing.
The results are summarized as follows.
First, penalized regression models identified true variables relatively accurately and maintained noise variable selection at a moderate level under conditions with smaller sample sizes and fewer predictors. Although statistical power was generally low under small-sample and highly imbalanced conditions, increasing sample size mitigated the adverse effects of imbalance and improved power. However, in high-dimensional settings with many predictors and a fixed number of true variables, the proportion of noise variables among the selected variables increased substantially, leading to very high FDRs even with large samples. As imbalance became more severe, the total number of selected variables decreased, resulting in reduced selection of both true and noise variables. Among penalized regression models, Elastic Net selected the largest number of noise variables and exhibited the highest FDR, whereas Adaptive Lasso showed the most conservative selection behavior, yielding the lowest power and FDR, particularly in high-dimensional conditions.
Second, error-control–based methods maintained FDR at a controlled level even in high-dimensional settings. Although variable selection was limited under small-sample conditions, both DRM and MDS accurately identified true variables as sample size increased. DRM performed relatively better in medium- to high-dimensional settings, whereas MDS showed superior performance in low-dimensional settings. Under severe imbalance, both methods experienced substantial declines in power, with MDS being more sensitive to outcome imbalance than DRM.
Third, regarding imbalance data preprocessing, SMOTE and ADASYN produced limited gains in power for penalized regression models while substantially increasing the selection of noise variables, resulting in markedly higher FDRs. In contrast, DRM and MDS exhibited clear improvements in power under low-dimensional conditions when preprocessing was applied, while maintaining low FDRs, indicating the potential utility of preprocessing in such contexts. However, in high-dimensional settings, both DRM and MDS showed degraded error control, suggesting that imbalance data preprocessing can reduce variable selection accuracy in these conditions.
Fourth, the real data analysis largely replicated the simulation findings. When sample sizes were sufficiently large, all methods selected variables stably, whereas increasing imbalance led to overall reductions in selection frequency. Applying imbalance data preprocessing increased the number of selected variables across both penalized regression models and error-control–based methods.
In summary, variable selection accuracy varies depending on analytical conditions such as sample size, number of predictors, and outcome imbalance. In educational data analysis, variable selection methods should therefore be chosen carefully by considering both data structure and research objectives. The findings indicate that penalized regression models tend to over-select noise variables in high-dimensional settings, whereas error-control–based methods can more accurately identify true variables while stably controlling FDR when sample sizes are sufficient. These results suggest that error-control–based methods provide an effective alternative for addressing the limitations of penalized regression models in educational research.
Moreover, although imbalance preprocessing may enhance predictive performance, it does not necessarily improve the accuracy or stability of variable selection; in some cases, it substantially increases noise variable selection. Accordingly, in research contexts that prioritize accurate variable selection and interpretability, imbalance preprocessing should be applied cautiously, and variable selection results should be interpreted using higher selection-frequency thresholds established through repeated analyses.
By integrating simulation designs that reflect the structural characteristics of educational data with empirical analyses, this study provides practical guidelines for selecting appropriate variable selection methods in educational research. In particular, this study highlights the limitations of machine-learning–based predictive models that have primarily been utilized with an emphasis on predictive accuracy. Given that selected variables are often used as the basis for educational and policy decision-making, the findings underscore the need to prioritize not only predictive performance but also variable selection accuracy. From this perspective, improving the accuracy and interpretability of selected variables is essential for ensuring valid inferences in educational data analysis.