RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    교육 자료 분석에서 벌점회귀모형과 오류 통제 기반 방법의 변수 선택 정확도 비교 연구 = A Comparative Study of Variable Selection Accuracy between Penalized Regression Models and Error-Control?Based Methods in Educational Data Analysis

    한글로보기

    https://www.riss.kr/link?id=T17451262

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The purpose of this study is to examine the accuracy of variable selection methods used in machine-learning–based predictive analyses of educational panel data. Educational panel data have been widely used to predict various outcomes, such as academic achievement, career decisions, and participation in private tutoring, by employing machine-learning–based predictive models. In this line of research, selected predictors are often interpreted to derive educational implications and policy recommendations. Penalized regression models, particularly LASSO, have been extensively adopted in educational studies because they simultaneously perform coefficient estimation and variable selection.
    However, high predictive performance does not necessarily guarantee that the selected variables are truly associated with the outcome. Previous studies have shown that penalized regression models may over-select noise variables, which can undermine the interpretability of the selected variables. In this context, variables such as suppressor variables can be selected due to their associations with other predictors, potentially improving prediction accuracy without a direct association with the outcome.
    Motivated by these concerns, this study shifts the focus from predictive performance to variable selection accuracy, examining how precisely machine-learning–based variable selection methods identify true predictors and how effectively they control the selection of noise variables. To this end, the variable selection accuracy of penalized regression models—LASSO, Elastic Net, and Adaptive Lasso—is compared with that of error-control–based methods, namely Derandomized knockoff (DRM) and Multiple Data Splitting (MDS). While penalized regression models prioritize predictive accuracy in the variable selection process, error-control–based methods explicitly regulate selection errors.
    In the simulation study, experimental conditions included sample size, the number of predictors, and the degree of imbalance in a binary outcome variable. Simulation data were generated to reflect key structural characteristics of educational panel data, such as Likert-scale item bundles, complex correlation structures among predictors, and imbalanced binary outcomes. Simulation Study 1 varied sample size (500, 2,000, 5,000), number of predictors (50, 200, 500), and imbalance ratios (1:1, 1:3, 1:9), and evaluated variable selection accuracy using statistical power as an index of true variable identification and the false discovery rate (FDR) as an index of noise variable control.
    Simulation Study 2 focused on the most severely imbalanced condition (1:9) and examined the effects of oversampling techniques—Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN)—on variable selection accuracy.
    Finally, to assess whether the simulation findings generalize to real-world contexts, we conducted analyses using data from the Korean Education Longitudinal Study 2013, focusing on the 4th- and 5th-wave data of middle school students in Grades 8 and 9. Variable selection stability across penalized regression models and error-control–based methods was examined according to sample size, imbalance ratio, and imbalance data preprocessing.
    The results are summarized as follows.
    First, penalized regression models identified true variables relatively accurately and maintained noise variable selection at a moderate level under conditions with smaller sample sizes and fewer predictors. Although statistical power was generally low under small-sample and highly imbalanced conditions, increasing sample size mitigated the adverse effects of imbalance and improved power. However, in high-dimensional settings with many predictors and a fixed number of true variables, the proportion of noise variables among the selected variables increased substantially, leading to very high FDRs even with large samples. As imbalance became more severe, the total number of selected variables decreased, resulting in reduced selection of both true and noise variables. Among penalized regression models, Elastic Net selected the largest number of noise variables and exhibited the highest FDR, whereas Adaptive Lasso showed the most conservative selection behavior, yielding the lowest power and FDR, particularly in high-dimensional conditions.
    Second, error-control–based methods maintained FDR at a controlled level even in high-dimensional settings. Although variable selection was limited under small-sample conditions, both DRM and MDS accurately identified true variables as sample size increased. DRM performed relatively better in medium- to high-dimensional settings, whereas MDS showed superior performance in low-dimensional settings. Under severe imbalance, both methods experienced substantial declines in power, with MDS being more sensitive to outcome imbalance than DRM.
    Third, regarding imbalance data preprocessing, SMOTE and ADASYN produced limited gains in power for penalized regression models while substantially increasing the selection of noise variables, resulting in markedly higher FDRs. In contrast, DRM and MDS exhibited clear improvements in power under low-dimensional conditions when preprocessing was applied, while maintaining low FDRs, indicating the potential utility of preprocessing in such contexts. However, in high-dimensional settings, both DRM and MDS showed degraded error control, suggesting that imbalance data preprocessing can reduce variable selection accuracy in these conditions.
    Fourth, the real data analysis largely replicated the simulation findings. When sample sizes were sufficiently large, all methods selected variables stably, whereas increasing imbalance led to overall reductions in selection frequency. Applying imbalance data preprocessing increased the number of selected variables across both penalized regression models and error-control–based methods.
    In summary, variable selection accuracy varies depending on analytical conditions such as sample size, number of predictors, and outcome imbalance. In educational data analysis, variable selection methods should therefore be chosen carefully by considering both data structure and research objectives. The findings indicate that penalized regression models tend to over-select noise variables in high-dimensional settings, whereas error-control–based methods can more accurately identify true variables while stably controlling FDR when sample sizes are sufficient. These results suggest that error-control–based methods provide an effective alternative for addressing the limitations of penalized regression models in educational research.
    Moreover, although imbalance preprocessing may enhance predictive performance, it does not necessarily improve the accuracy or stability of variable selection; in some cases, it substantially increases noise variable selection. Accordingly, in research contexts that prioritize accurate variable selection and interpretability, imbalance preprocessing should be applied cautiously, and variable selection results should be interpreted using higher selection-frequency thresholds established through repeated analyses.
    By integrating simulation designs that reflect the structural characteristics of educational data with empirical analyses, this study provides practical guidelines for selecting appropriate variable selection methods in educational research. In particular, this study highlights the limitations of machine-learning–based predictive models that have primarily been utilized with an emphasis on predictive accuracy. Given that selected variables are often used as the basis for educational and policy decision-making, the findings underscore the need to prioritize not only predictive performance but also variable selection accuracy. From this perspective, improving the accuracy and interpretability of selected variables is essential for ensuring valid inferences in educational data analysis.
    번역하기

    The purpose of this study is to examine the accuracy of variable selection methods used in machine-learning–based predictive analyses of educational panel data. Educational panel data have been widely used to predict various outcomes, such as academ...

    The purpose of this study is to examine the accuracy of variable selection methods used in machine-learning–based predictive analyses of educational panel data. Educational panel data have been widely used to predict various outcomes, such as academic achievement, career decisions, and participation in private tutoring, by employing machine-learning–based predictive models. In this line of research, selected predictors are often interpreted to derive educational implications and policy recommendations. Penalized regression models, particularly LASSO, have been extensively adopted in educational studies because they simultaneously perform coefficient estimation and variable selection.
    However, high predictive performance does not necessarily guarantee that the selected variables are truly associated with the outcome. Previous studies have shown that penalized regression models may over-select noise variables, which can undermine the interpretability of the selected variables. In this context, variables such as suppressor variables can be selected due to their associations with other predictors, potentially improving prediction accuracy without a direct association with the outcome.
    Motivated by these concerns, this study shifts the focus from predictive performance to variable selection accuracy, examining how precisely machine-learning–based variable selection methods identify true predictors and how effectively they control the selection of noise variables. To this end, the variable selection accuracy of penalized regression models—LASSO, Elastic Net, and Adaptive Lasso—is compared with that of error-control–based methods, namely Derandomized knockoff (DRM) and Multiple Data Splitting (MDS). While penalized regression models prioritize predictive accuracy in the variable selection process, error-control–based methods explicitly regulate selection errors.
    In the simulation study, experimental conditions included sample size, the number of predictors, and the degree of imbalance in a binary outcome variable. Simulation data were generated to reflect key structural characteristics of educational panel data, such as Likert-scale item bundles, complex correlation structures among predictors, and imbalanced binary outcomes. Simulation Study 1 varied sample size (500, 2,000, 5,000), number of predictors (50, 200, 500), and imbalance ratios (1:1, 1:3, 1:9), and evaluated variable selection accuracy using statistical power as an index of true variable identification and the false discovery rate (FDR) as an index of noise variable control.
    Simulation Study 2 focused on the most severely imbalanced condition (1:9) and examined the effects of oversampling techniques—Synthetic Minority Over-sampling Technique (SMOTE) and Adaptive Synthetic Sampling (ADASYN)—on variable selection accuracy.
    Finally, to assess whether the simulation findings generalize to real-world contexts, we conducted analyses using data from the Korean Education Longitudinal Study 2013, focusing on the 4th- and 5th-wave data of middle school students in Grades 8 and 9. Variable selection stability across penalized regression models and error-control–based methods was examined according to sample size, imbalance ratio, and imbalance data preprocessing.
    The results are summarized as follows.
    First, penalized regression models identified true variables relatively accurately and maintained noise variable selection at a moderate level under conditions with smaller sample sizes and fewer predictors. Although statistical power was generally low under small-sample and highly imbalanced conditions, increasing sample size mitigated the adverse effects of imbalance and improved power. However, in high-dimensional settings with many predictors and a fixed number of true variables, the proportion of noise variables among the selected variables increased substantially, leading to very high FDRs even with large samples. As imbalance became more severe, the total number of selected variables decreased, resulting in reduced selection of both true and noise variables. Among penalized regression models, Elastic Net selected the largest number of noise variables and exhibited the highest FDR, whereas Adaptive Lasso showed the most conservative selection behavior, yielding the lowest power and FDR, particularly in high-dimensional conditions.
    Second, error-control–based methods maintained FDR at a controlled level even in high-dimensional settings. Although variable selection was limited under small-sample conditions, both DRM and MDS accurately identified true variables as sample size increased. DRM performed relatively better in medium- to high-dimensional settings, whereas MDS showed superior performance in low-dimensional settings. Under severe imbalance, both methods experienced substantial declines in power, with MDS being more sensitive to outcome imbalance than DRM.
    Third, regarding imbalance data preprocessing, SMOTE and ADASYN produced limited gains in power for penalized regression models while substantially increasing the selection of noise variables, resulting in markedly higher FDRs. In contrast, DRM and MDS exhibited clear improvements in power under low-dimensional conditions when preprocessing was applied, while maintaining low FDRs, indicating the potential utility of preprocessing in such contexts. However, in high-dimensional settings, both DRM and MDS showed degraded error control, suggesting that imbalance data preprocessing can reduce variable selection accuracy in these conditions.
    Fourth, the real data analysis largely replicated the simulation findings. When sample sizes were sufficiently large, all methods selected variables stably, whereas increasing imbalance led to overall reductions in selection frequency. Applying imbalance data preprocessing increased the number of selected variables across both penalized regression models and error-control–based methods.
    In summary, variable selection accuracy varies depending on analytical conditions such as sample size, number of predictors, and outcome imbalance. In educational data analysis, variable selection methods should therefore be chosen carefully by considering both data structure and research objectives. The findings indicate that penalized regression models tend to over-select noise variables in high-dimensional settings, whereas error-control–based methods can more accurately identify true variables while stably controlling FDR when sample sizes are sufficient. These results suggest that error-control–based methods provide an effective alternative for addressing the limitations of penalized regression models in educational research.
    Moreover, although imbalance preprocessing may enhance predictive performance, it does not necessarily improve the accuracy or stability of variable selection; in some cases, it substantially increases noise variable selection. Accordingly, in research contexts that prioritize accurate variable selection and interpretability, imbalance preprocessing should be applied cautiously, and variable selection results should be interpreted using higher selection-frequency thresholds established through repeated analyses.
    By integrating simulation designs that reflect the structural characteristics of educational data with empirical analyses, this study provides practical guidelines for selecting appropriate variable selection methods in educational research. In particular, this study highlights the limitations of machine-learning–based predictive models that have primarily been utilized with an emphasis on predictive accuracy. Given that selected variables are often used as the basis for educational and policy decision-making, the findings underscore the need to prioritize not only predictive performance but also variable selection accuracy. From this perspective, improving the accuracy and interpretability of selected variables is essential for ensuring valid inferences in educational data analysis.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    머신러닝 기법을 활용한 교육 자료 분석은 학업 성취, 진로 결정 여부 등 다양한 결과 변수 예측을 목적으로 이루어지고 있다. LASSO와 같은 벌점회귀모형은 변수 선택과 회귀계수 추정이 동시에 가능하다는 장점으로 교육 연구에서 널리 사용되고 있으며, 연구자들은 예측 과정에서 선택된 변수들을 바탕으로 교육적 시사점을 도출하거나 정책적 제안에 활용하였다. 그러나 예측 성능이 우수한 모형에서 선택된 변수가 반드시 종속변수와 관련된 참변수(true variable)라고 보장할 수 없다. 실제로 벌점회귀모형은 변수 선택 과정에서 다수의 잡음변수(noise variable)를 함께 선택한다는 한계가 지속적으로 제기되고 있다. 따라서 예측모형을 통해 선택된 변수를 교육적·정책적으로 해석하는 데에는 주의가 필요하다.
    이러한 문제의식을 바탕으로 이 연구는 기존 연구에서 주로 강조된 예측 성능이 아닌 변수 선택의 정확도에 주목하여, 머신러닝 기반 변수 선택 방법이 참변수를 얼마나 정확하게 식별하고 잡음변수 선택을 얼마나 효과적으로 통제할 수 있는지 분석하였다. 이를 위해 벌점회귀모형인 LASSO, Elastic Net, Adaptive Lasso와 오류 통제 기반 변수 선택 방법인 Derandomized knockoff(DRM), Multiple Data Splitting(MDS)의 변수 선택 정확도를 비교하였다. 벌점회귀모형은 예측 정확도를 중심으로 변수를 선택하는 반면, 오류 통제 기반 방법은 변수 선택 과정에서 발생하는 오류를 명시적으로 통제하며 변수를 선택한다는 점에서, 이 연구는 서로 다른 변수 선택 기준을 갖는 방법들의 특성을 실증적으로 검토하는 데 의의가 있다.
    모의실험 연구에서는 표본 수, 투입 변수의 수, 종속변수의 불균형 비율을 실험 조건으로 설정하고, 교육 패널 자료의 구조적 특성인 리커트 척도 기반 문항 묶음 구조, 변수 간 복잡한 상관 구조, 불균형한 이분형 종속변수의 특성을 반영한 모의실험 자료를 생성하였다. 모의실험 연구 1에서는 표본 수, 투입 변수의 수, 종속변수 불균형 비율을 변화시키며, 참변수 식별 성능인 검정력(power)과 잡음변수 통제 성능인 거짓 발견율(False Discovery Rate, FDR)을 중심으로 각 분석 방법의 변수 선택 정확도를 평가하였다. 모의실험 연구 2에서는 종속변수가 가장 불균형한 조건에 오버샘플링 기법인 SMOTE(Synthetic Minority Over-sampling Technique)와 ADASYN(Adaptive Synthetic Sampling)을 적용하여 예측 성능 향상을 목적으로 활용되어 온 불균형 데이터 전처리가 변수 선택 정확도에 미치는 영향을 검증하였다. 마지막으로 모의실험 결과가 실제 자료 분석에서도 재현되는지 검증하기 위해 한국교육종단연구2013의 4, 5차년도 중학교 2, 3학년 자료를 활용하여 벌점회귀모형과 오류 통제 기반 방법의 변수 선택 안정성을 비교함으로써 모의실험 결과의 재현 가능성을 검증하였다.
    연구 결과를 요약하면 다음과 같다.
    첫째, 벌점회귀모형은 표본 수와 투입 변수의 수가 적은 조건에서 참변수를 비교적 정확하게 식별하며, 잡음변수 선택을 일정 수준 이하로 유지하였다. 표본 수가 적고 불균형이 심한 조건에서는 검정력이 전반적으로 낮으나, 표본 수가 증가할수록 불균형으로 인한 성능 저하가 완화되며 검정력이 높아졌다. 그러나 투입 변수가 많아지고 잡음변수의 비율이 높아지는 고차원 조건에서는 표본 수를 늘려도 잡음변수가 함께 선택될 가능성이 커져 거짓 발견율이 매우 높게 나타났다. 반면, 종속변수의 불균형이 심해질수록 선택되는 변수의 수가 전반적으로 감소하여 참변수와 잡음변수 선택이 모두 감소하는 경향을 보였다. 분석 기법 간 비교 결과, Elastic Net은 가장 많은 변수를 선택하여 검정력과 거짓 발견율이 높았던 반면, Adaptive Lasso는 고차원 조건에서도 상대적으로 보수적인 변수 선택을 하여 검정력과 거짓 발견율 모두 가장 낮았다.
    둘째, 오류 통제 기반 방법은 투입 변수가 많고 잡음변수 비율이 높은 고차원 조건에서 거짓 발견율을 일정 수준 이하로 유지하였다. 소표본 조건에서는 변수 선택이 제대로 이루어지지 않아 변수 선택 성능이 낮았으나, 표본 수가 증가하면서 참변수를 정확하게 식별하였다. 분석 기법 간 비교 결과, DRM은 중·고차원 조건에서, MDS는 저차원 조건에서 상대적으로 높은 성능을 나타냈다. 또한 불균형이 심한 조건에서는 DRM과 MDS 모두 검정력이 크게 감소하였고, MDS가 DRM보다 종속변수 불균형에 더 민감하게 반응하였다.
    셋째, 불균형 데이터 전처리 기법 중 오버샘플링이 변수 선택 정확도에 미치는 영향을 검증한 결과, 벌점회귀모형은 SMOTE와 ADASYN을 적용하더라도 검정력 향상이 제한적이었으며, 오히려 전처리 후 잡음변수가 함께 선택될 가능성이 커져 거짓 발견율이 크게 증가하였다. 반면, DRM과 MDS는 저차원 조건에서 불균형 데이터 전처리를 적용하면 검정력이 뚜렷하게 높아졌고, 거짓 발견율도 낮은 수준으로 유지되어 저차원 자료에서 전처리 기법의 활용 가능성을 확인하였다. 그러나 고차원 조건에서는 DRM과 MDS 모두 오류 통제 성능이 저하되어 불균형 데이터 전처리가 변수 선택의 정확도를 오히려 낮출 수 있음을 확인하였다.
    넷째, 실제 자료 분석에서도 모의실험 결과가 상당 부분 재현되었다. 표본 수가 일정 수준 이상 충분히 확보되면 모든 분석 방법이 변수를 안정적으로 선택하였고, 불균형 비율이 심해지면 선택 빈도가 전반적으로 감소하였다. 또한 불균형 데이터 전처리 적용 시 벌점회귀모형뿐만 아니라 오류 통제 기반 방법 모두에서 선택되는 변수의 수가 증가하였다.
    종합하면, 변수 선택의 정확도는 표본 수, 투입 변수의 수, 종속변수의 불균형 비율과 같은 분석 조건에 따라 다르게 나타나며, 교육 자료 분석에서는 자료의 구조적 특성과 분석 목적을 고려하여 변수 선택 방법을 신중하게 선택할 필요가 있다. 이 연구 결과, 벌점회귀모형은 투입 변수가 많아지고 잡음변수 비율이 높아질수록 잡음변수를 과도하게 선택하였다. 반면, 오류 통제 기반 방법은 표본 수가 충분하면 거짓 발견율을 안정적으로 통제하면서 참변수를 더 정확히 식별할 수 있음을 확인하였다. 이는 교육 자료 분석에서 오류 통제 기반 방법이 벌점회귀모형의 과도한 변수 선택이라는 한계를 보완하는 효과적인 대안이 될 수 있음을 시사한다.
    또한 불균형 데이터 전처리는 예측 성능의 향상에는 기여할 수 있으나 변수 선택의 정확성과 안정성을 개선하는 것은 아니며, 전처리 기법의 특성에 따라 자료 구조가 변형되면서 잡음변수 선택이 오히려 크게 증가할 수 있음을 확인하였다. 특히 고차원 조건에서는 오버샘플링 기법의 적용이 검정력 향상과 동시에 거짓 발견율도 크게 증가시키는 경향이 나타났다. 이에 따라 높은 예측 성능뿐만 아니라 정확한 변수 선택과 해석 가능성이 중요한 연구 맥락에서는 불균형 데이터 전처리 기법을 신중하게 적용하고, 반복 분석을 통해 선택 빈도 임곗값을 높게 설정한 기준으로 변수 선택 결과를 해석할 필요가 있다.
    이 연구는 교육 자료의 특성을 고려한 모의실험 연구와 실제 자료 분석을 통해 벌점회귀모형과 오류 통제 기반 방법의 변수 선택 결과를 비교함으로써, 변수 선택 방법 적용 시 고려해야 할 분석 전략에 대한 실질적 가이드라인을 제시한다. 특히 예측 정확도 중심으로 활용되어 온 머신러닝 기반 예측모형의 한계를 지적하고, 선택된 변수가 정책적 의사결정의 근거로 활용된다는 점에서 예측 성능뿐만 아니라 변수 선택의 정확도를 높이려는 노력이 필요함을 강조한다.
    번역하기

    머신러닝 기법을 활용한 교육 자료 분석은 학업 성취, 진로 결정 여부 등 다양한 결과 변수 예측을 목적으로 이루어지고 있다. LASSO와 같은 벌점회귀모형은 변수 선택과 회귀계수 추정이 동...

    머신러닝 기법을 활용한 교육 자료 분석은 학업 성취, 진로 결정 여부 등 다양한 결과 변수 예측을 목적으로 이루어지고 있다. LASSO와 같은 벌점회귀모형은 변수 선택과 회귀계수 추정이 동시에 가능하다는 장점으로 교육 연구에서 널리 사용되고 있으며, 연구자들은 예측 과정에서 선택된 변수들을 바탕으로 교육적 시사점을 도출하거나 정책적 제안에 활용하였다. 그러나 예측 성능이 우수한 모형에서 선택된 변수가 반드시 종속변수와 관련된 참변수(true variable)라고 보장할 수 없다. 실제로 벌점회귀모형은 변수 선택 과정에서 다수의 잡음변수(noise variable)를 함께 선택한다는 한계가 지속적으로 제기되고 있다. 따라서 예측모형을 통해 선택된 변수를 교육적·정책적으로 해석하는 데에는 주의가 필요하다.
    이러한 문제의식을 바탕으로 이 연구는 기존 연구에서 주로 강조된 예측 성능이 아닌 변수 선택의 정확도에 주목하여, 머신러닝 기반 변수 선택 방법이 참변수를 얼마나 정확하게 식별하고 잡음변수 선택을 얼마나 효과적으로 통제할 수 있는지 분석하였다. 이를 위해 벌점회귀모형인 LASSO, Elastic Net, Adaptive Lasso와 오류 통제 기반 변수 선택 방법인 Derandomized knockoff(DRM), Multiple Data Splitting(MDS)의 변수 선택 정확도를 비교하였다. 벌점회귀모형은 예측 정확도를 중심으로 변수를 선택하는 반면, 오류 통제 기반 방법은 변수 선택 과정에서 발생하는 오류를 명시적으로 통제하며 변수를 선택한다는 점에서, 이 연구는 서로 다른 변수 선택 기준을 갖는 방법들의 특성을 실증적으로 검토하는 데 의의가 있다.
    모의실험 연구에서는 표본 수, 투입 변수의 수, 종속변수의 불균형 비율을 실험 조건으로 설정하고, 교육 패널 자료의 구조적 특성인 리커트 척도 기반 문항 묶음 구조, 변수 간 복잡한 상관 구조, 불균형한 이분형 종속변수의 특성을 반영한 모의실험 자료를 생성하였다. 모의실험 연구 1에서는 표본 수, 투입 변수의 수, 종속변수 불균형 비율을 변화시키며, 참변수 식별 성능인 검정력(power)과 잡음변수 통제 성능인 거짓 발견율(False Discovery Rate, FDR)을 중심으로 각 분석 방법의 변수 선택 정확도를 평가하였다. 모의실험 연구 2에서는 종속변수가 가장 불균형한 조건에 오버샘플링 기법인 SMOTE(Synthetic Minority Over-sampling Technique)와 ADASYN(Adaptive Synthetic Sampling)을 적용하여 예측 성능 향상을 목적으로 활용되어 온 불균형 데이터 전처리가 변수 선택 정확도에 미치는 영향을 검증하였다. 마지막으로 모의실험 결과가 실제 자료 분석에서도 재현되는지 검증하기 위해 한국교육종단연구2013의 4, 5차년도 중학교 2, 3학년 자료를 활용하여 벌점회귀모형과 오류 통제 기반 방법의 변수 선택 안정성을 비교함으로써 모의실험 결과의 재현 가능성을 검증하였다.
    연구 결과를 요약하면 다음과 같다.
    첫째, 벌점회귀모형은 표본 수와 투입 변수의 수가 적은 조건에서 참변수를 비교적 정확하게 식별하며, 잡음변수 선택을 일정 수준 이하로 유지하였다. 표본 수가 적고 불균형이 심한 조건에서는 검정력이 전반적으로 낮으나, 표본 수가 증가할수록 불균형으로 인한 성능 저하가 완화되며 검정력이 높아졌다. 그러나 투입 변수가 많아지고 잡음변수의 비율이 높아지는 고차원 조건에서는 표본 수를 늘려도 잡음변수가 함께 선택될 가능성이 커져 거짓 발견율이 매우 높게 나타났다. 반면, 종속변수의 불균형이 심해질수록 선택되는 변수의 수가 전반적으로 감소하여 참변수와 잡음변수 선택이 모두 감소하는 경향을 보였다. 분석 기법 간 비교 결과, Elastic Net은 가장 많은 변수를 선택하여 검정력과 거짓 발견율이 높았던 반면, Adaptive Lasso는 고차원 조건에서도 상대적으로 보수적인 변수 선택을 하여 검정력과 거짓 발견율 모두 가장 낮았다.
    둘째, 오류 통제 기반 방법은 투입 변수가 많고 잡음변수 비율이 높은 고차원 조건에서 거짓 발견율을 일정 수준 이하로 유지하였다. 소표본 조건에서는 변수 선택이 제대로 이루어지지 않아 변수 선택 성능이 낮았으나, 표본 수가 증가하면서 참변수를 정확하게 식별하였다. 분석 기법 간 비교 결과, DRM은 중·고차원 조건에서, MDS는 저차원 조건에서 상대적으로 높은 성능을 나타냈다. 또한 불균형이 심한 조건에서는 DRM과 MDS 모두 검정력이 크게 감소하였고, MDS가 DRM보다 종속변수 불균형에 더 민감하게 반응하였다.
    셋째, 불균형 데이터 전처리 기법 중 오버샘플링이 변수 선택 정확도에 미치는 영향을 검증한 결과, 벌점회귀모형은 SMOTE와 ADASYN을 적용하더라도 검정력 향상이 제한적이었으며, 오히려 전처리 후 잡음변수가 함께 선택될 가능성이 커져 거짓 발견율이 크게 증가하였다. 반면, DRM과 MDS는 저차원 조건에서 불균형 데이터 전처리를 적용하면 검정력이 뚜렷하게 높아졌고, 거짓 발견율도 낮은 수준으로 유지되어 저차원 자료에서 전처리 기법의 활용 가능성을 확인하였다. 그러나 고차원 조건에서는 DRM과 MDS 모두 오류 통제 성능이 저하되어 불균형 데이터 전처리가 변수 선택의 정확도를 오히려 낮출 수 있음을 확인하였다.
    넷째, 실제 자료 분석에서도 모의실험 결과가 상당 부분 재현되었다. 표본 수가 일정 수준 이상 충분히 확보되면 모든 분석 방법이 변수를 안정적으로 선택하였고, 불균형 비율이 심해지면 선택 빈도가 전반적으로 감소하였다. 또한 불균형 데이터 전처리 적용 시 벌점회귀모형뿐만 아니라 오류 통제 기반 방법 모두에서 선택되는 변수의 수가 증가하였다.
    종합하면, 변수 선택의 정확도는 표본 수, 투입 변수의 수, 종속변수의 불균형 비율과 같은 분석 조건에 따라 다르게 나타나며, 교육 자료 분석에서는 자료의 구조적 특성과 분석 목적을 고려하여 변수 선택 방법을 신중하게 선택할 필요가 있다. 이 연구 결과, 벌점회귀모형은 투입 변수가 많아지고 잡음변수 비율이 높아질수록 잡음변수를 과도하게 선택하였다. 반면, 오류 통제 기반 방법은 표본 수가 충분하면 거짓 발견율을 안정적으로 통제하면서 참변수를 더 정확히 식별할 수 있음을 확인하였다. 이는 교육 자료 분석에서 오류 통제 기반 방법이 벌점회귀모형의 과도한 변수 선택이라는 한계를 보완하는 효과적인 대안이 될 수 있음을 시사한다.
    또한 불균형 데이터 전처리는 예측 성능의 향상에는 기여할 수 있으나 변수 선택의 정확성과 안정성을 개선하는 것은 아니며, 전처리 기법의 특성에 따라 자료 구조가 변형되면서 잡음변수 선택이 오히려 크게 증가할 수 있음을 확인하였다. 특히 고차원 조건에서는 오버샘플링 기법의 적용이 검정력 향상과 동시에 거짓 발견율도 크게 증가시키는 경향이 나타났다. 이에 따라 높은 예측 성능뿐만 아니라 정확한 변수 선택과 해석 가능성이 중요한 연구 맥락에서는 불균형 데이터 전처리 기법을 신중하게 적용하고, 반복 분석을 통해 선택 빈도 임곗값을 높게 설정한 기준으로 변수 선택 결과를 해석할 필요가 있다.
    이 연구는 교육 자료의 특성을 고려한 모의실험 연구와 실제 자료 분석을 통해 벌점회귀모형과 오류 통제 기반 방법의 변수 선택 결과를 비교함으로써, 변수 선택 방법 적용 시 고려해야 할 분석 전략에 대한 실질적 가이드라인을 제시한다. 특히 예측 정확도 중심으로 활용되어 온 머신러닝 기반 예측모형의 한계를 지적하고, 선택된 변수가 정책적 의사결정의 근거로 활용된다는 점에서 예측 성능뿐만 아니라 변수 선택의 정확도를 높이려는 노력이 필요함을 강조한다.

    더보기

    목차 (Table of Contents)

    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 문제 8
    • Ⅱ. 이론적 배경 9
    • Ⅰ. 서론 1
    • 1. 연구의 필요성 및 목적 1
    • 2. 연구 문제 8
    • Ⅱ. 이론적 배경 9
    • 1. 벌점회귀를 통한 변수 선택 9
    • 가. 벌점회귀모형의 등장 배경 9
    • 나. LASSO 11
    • 다. LASSO의 변수 선택 일관성을 위한 조건 14
    • 라. Elastic Net과 Adaptive Lasso 17
    • 마. 벌점회귀모형의 한계 20
    • 2. 오류 통제를 통한 변수 선택 23
    • 가. 오류 통제의 필요성 23
    • 나. Knockoff와 Derandomized knockoff 26
    • 다. Data Splitting과 Multiple Data Splitting 39
    • 3 . 불균형 데이터 전처리 기법과 변수 선택 45
    • 가. 교육 데이터의 불균형 특성과 예측 성능 45
    • 나. 불균형 데이터 전처리 기법 48
    • 다. 불균형과 전처리 기법이 변수 선택에 미치는 영향 53
    • Ⅲ. 연구 방법 56
    • 1. 모의실험 연구 설계 56
    • 가. 모의실험 자료 57
    • 나. 변수 선택 및 불균형 데이터 전처리 기법 73
    • 다. 평가 준거 78
    • 2. 실제 자료 분석 설계 81
    • 가. 분석 자료 82
    • 나. 분석 방법 및 평가 준거 85
    • Ⅳ. 연구 결과 88
    • 1. 표본 수, 투입 변수 수, 불균형 비율에 따른 변수 선택 정확도 비교 88
    • 가. 표본 수가 500명일 때 88
    • 나. 표본 수가 2,000명일 때 98
    • 다. 표본 수가 5,000명일 때 109
    • 2. 불균형 데이터 전처리에 따른 변수 선택 정확도 비교 119
    • 가. 표본 수가 2,000명일 때 120
    • 나. 표본 수가 5,000명일 때 127
    • 3. 실제 자료 분석 결과 133
    • 가. 표본 수 및 불균형 비율에 따라 133
    • 나. 불균형 데이터 전처리 적용에 따라 138
    • Ⅴ. 요약 및 논의 141
    • 1. 요약 141
    • 2. 논의 148
    • 참고 문헌 155
    • 부록 170
    • Abstract 175
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼