학교 구성원들의 동기, 행동, 학업 성취, 학교 환경과 같은 핵심 변인들 간의 인과적 관계를 규명하는 것은 교육 연구자들의 주된 관심사이자 목표이다. 인과성에 대한 정보를 얻기 위해 가...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17450041
서울 : 서울대학교 대학원, 2026
학위논문(박사) -- 서울대학교 대학원 , 교육학과 교육측정평가 , 2026. 2
2026
영어
370
서울
xii, 153 ; 26 cm
지도교수: 김용남
I804:11032-000000196596
0
상세조회0
다운로드학교 구성원들의 동기, 행동, 학업 성취, 학교 환경과 같은 핵심 변인들 간의 인과적 관계를 규명하는 것은 교육 연구자들의 주된 관심사이자 목표이다. 인과성에 대한 정보를 얻기 위해 가...
학교 구성원들의 동기, 행동, 학업 성취, 학교 환경과 같은 핵심 변인들 간의 인과적 관계를 규명하는 것은 교육 연구자들의 주된 관심사이자 목표이다. 인과성에 대한 정보를 얻기 위해 가장 권장되는 연구방법은 무작위 대조 실험(Randomized Controlled Trial)이지만, 실제로 교육 연구 수행 환경에서 무작위 배정(random assignment)이 가능하거나 이러한 실험이 윤리적으로 허용되는 경우는 극히 제한적이다. 이에 연구자들은 회귀 조정, 성향점수(propensity score) 접근법 등 다양한 준실험적(quasi-experimental) 전략을 활용해 관측된 자료로부터 실험적 논리를 구현하고자 노력해왔다. 그러나 회귀 조정과 성향점수 접근법이 가정하고 있는 무교란성(unconfoundedness), 즉 모든 교란변수를 관찰할 수 있고 통제할 수 있다는 가정은 실제로 달성되기 매우 어려운 것이며, 측정된 공변량을 통제하는 과정 자체가 새로운 편향을 유발할 위험도 존재한다. 이러한 한계 속에서, 도구변수(instrumental variable)는 관찰하지 못한 교란변수가 존재하더라도 자연적 변동이나 제도적 요인을 활용해 인과 정보를 식별할 수 있다는 점에서 중요한 대안으로 자리해 왔다.
사회과학 및 보건 분야 전반에서 도구변수는 오랜 이론적 기반을 바탕으로 폭넓게 활용되어 왔지만, 교육학 분야에서의 방법론적 축적은 아직 충분하지 않은 상태이다. 도구변수 활용에 있어서 가장 큰 장애물은 도구변수 분석법의 핵심적인 가정인 외생성(exogeneity) 가정을 경험적으로 평가하기가 매우 어렵다는 점이다. Hausman 검정이나 Sargan, Hansen J Test 등의 표준 기법은, 도구변수가 이미 외생적이라는 가정을 바탕으로 하여 검정 대상이 되는 도구변수 추정치가 편향되었을 때 올바른 결론을 내리지 못할 가능성이 높고, 복수의 타당한 도구변수를 요구하기 때문에 현실적으로 해당 조건을 만족하기가 어렵다. 그 결과, 응용 연구자들은 선택 과정이 복잡하게 작용하는 실제 교육 자료에서 도구변수를 기반으로 한 인과적인 추론이 얼마나 신뢰할 만한지 체계적으로 점검하기 어려운 상황에 놓여 있다.
이러한 한계를 해소하기 위해, 본 논문에서는 교육 자료에서 빈번하게 관찰되는 전형적인 자료 구조인 사전-사후검사 설계(pretest-posttest design)를 활용하여, 도구변수의 외생성 가정을 검정 가능한 동등성 조건으로 치환하는 구조 기반 접근법을 제안한다. 이 접근법의 핵심은, 특정한 구조가 전제로 하는 변인들 간의 구조적인 관계를 활용하는 것이다. 도구변수 추정량과 구조적으로 대응되는 준거 추정량을 도출하고, 두 추정량이 동일해야 한다는 귀납적 함의를 바탕으로 외생성 가정을 경험적으로 검정 가능한 형태로 재구성한다. 이러한 원리는 준거치인 탐지변수 교정치와 도구변수 추정치를 비교하는 사전-사후검사 설계에서뿐 아니라, 집단변수를 활용해 문항 수준에서의 직접 효과가 존재하는지를 평가하는 측정 맥락에서도 동일하게 적용된다. 본 논문은 세 개의 상호 연계된 연구로 구성되며, 도구변수 분석법의 구조에 기반한 논리를 이론적으로 정식화하고 시뮬레이션과 실증 분석을 통해 검증함으로써, 교육 연구의 현실적인 제약 속에서도 도구변수 분석의 투명성과 타당성을 강화하는 새로운 접근을 제시한다.
연구 1은 도구변수 분석의 논리를 사전-사후검사 설계로 확장하여, 단일 도구변수 상황에서도 외생성 위배를 평가할 수 있는 방법을 제안한다. 사전점수와 사후점수가 공통된 잠재 교란변수의 영향을 받는다고 가정하면, 교란변수가 사전점수와 사후점수에 미치는 영향력의 크기가 동일하다는 공통추세(common trend) 가정을 만족하는 경우 차이점수 분석법(gain score analysis; Difference in Differences)을 활용해 인과효과를 식별할 수 있다. 만약 특정한 조건을 만족하는 탐지변수(compass variable)가 존재한다면, 공통추세 가정이 위배된 정도를 수량화하여 교정함으로써 타당한 인과효과 추정치를 구할 수 있다. 본 연구에서는 사전-사후검사 설계에서의 도구변수와 탐지변수의 구조적 유사성에 주목하여, 도구변수 추정치와 구조적으로 대응하는 탐지변수로 교정한 준거 추정량을 도출한다. 도구변수의 외생성이 성립한다면 두 추정량은 통계적으로 동등해야 하므로, 두 추정량의 차이는 외생성 가정을 귀납적으로 검정하는 통계량으로 기능한다. 이를 위해 연구 1은 적층 GMM(Stacked GMM)을 활용해 두 추정치의 차이를 검정하는 통계량을 유도하고, 몬테-카를로 시뮬레이션(Monte-Carlo Simulation)을 통해 이 진단의 검정력과 제1종 오류율을 검증하였다. 분석 결과, 도구변수가 충분히 강한 경우 외생성 위배를 효과적으로 탐지하였고, 도구변수와 교란변수 간 조건부 연관성(conditional relevance)과 같은 식별 조건이 약화되는 영역에서는 통계량의 분산이 팽창하여 추론이 구조적으로 취약해짐을 보였다.
연구 2는 연구 1에서 제안한 진단이 실제 교육 자료에서 어떻게 작동하는지를 검증한다. 이를 위해 미국 NELS:88 자료를 활용하여, 교육사회학 분야에서 가톨릭 학교 효과를 둘러싼 오랜 논쟁을 재조명한다. 기존 문헌은 종교적 배경을 도구변수로 활용한 도구변수 분석에서 추정치가 비정상적으로 크게 나타날 수 있음을 보고해 왔으며, 이러한 결과가 학교 유형의 인과효과인지 선택 편향의 반영인지에 대한 해석 불확실성을 남겨왔다. 본 연구는 전통적 2SLS 사양을 재현한 뒤, 사전–사후검사 구조에서 유도되는 탐지변수 준거 추정량과 전통적 IV 추정량을 후보 도구변수별로 비교하는 진단을 적용하여 외생성 가정의 개연성을 점검하였다. 그 결과 읽기에서는 어떤 후보 도구변수에서도 두 추정량의 불일치가 뚜렷하게 확인되지 않아, 특정 도구변수의 외생성 위배를 단정하기 어렵다. 반면 수학에서는 Urban$\times$Region 상호작용 도구변수에서만 구조 기반 준거치와의 체계적 불일치가 확인되어 외생성 가정 위배가 지목되었고, 종교 기반 도구변수들은 상대적으로 더 안정적인 패턴을 보였다. 이러한 결과는 수학에서의 다중 도구변수 2SLS 결론이 어떤 도구변수에 의존하는지에 따라 달라질 수 있음을 보여주며, 학교 선택이 사회경제적 배경과 얽혀 있는 맥락에서 구조 기반 진단이 IV 추정의 해석 투명성을 높이는 데 유용함을 시사한다.
연구 3에서는 차별기능문항(Differential Item Functioning, DIF)이 능력을 조건화(condition on)했을 때 나타나는 집단 간 문항반응 차이라는 통상적인 개념에서 벗어나, 도구변수 기반의 과잉식별 논리를 활용해 인과적 의미의 문항편향을 진단하는 새로운 구조 기반 접근을 제시한다. 기존의 DIF 분석법들은 잠재능력 추정치를 조건화하여 집단 간 문항반응에서의 차이를 평가한다. 그러나 잠재능력 모수와 문항반응에 영향을 미치는 관찰되지 않은 교란변수가 존재할 경우, 능력 모수가 충돌변수(collider)로 기능하면서 비인과적인 경로가 열릴 수 있다. 이 경우, 실제로 집단변수가 문항반응에 미치는 직접적인 효과인 문항편향이 없음에도 불구하고 비인과적인 관련성이 유입되면서 가짜 DIF(spurious DIF)가 탐지될 위험이 있다. 즉, 기존의 분석법들은 능력 모수를 조건화하기 때문에, 비인과적인 신로를 DIF로 오인할 가능성을 온전히 배제하기 어렵다. 이러한 한계를 보완하기 위해, 본 연구에서는 잠재능력 추정치를 조건화하지 않고, 잠재능력에 영향을 미치는 두 개 이상의 집단변수가 존재할 때 각 집단변수를 잠재능력에 대한 도구변수로 간주해서 도구별 추정치의 일관성을 검토하는 새로운 접근 방식을 제안한다. 만약 특정 문항이 편향되지 않았다면, 도구변수인 집단변수가 결과변수인 문항반응에 미치는 직접적인 효과가 없기 때문에 도구변수의 배제제한(exclusion restriction) 가정이 만족되어야 하고 집단변수별 도구변수 추정량은 서로 일치해야 한다. 반대로 특정 문항이 편향되었다면, 집단변수별 도구변수 추정량은 체계적으로 불일치하게 되고, 이를 문항편향의 존재에 대한 반증으로 해석할 수 있다. 시뮬레이션 결과, 무편향 상황에서의 기각률이 명목 수준에 근접하였고, 단일 집단특성만이 직접효과를 갖는 경우 의미 있는 검정력을 보였다. 관찰되지 않은 교란변수가 존재하더라도 잠재능력 추정치를 조건화하지 않고 문항편향을 점검할 수 있다는 점에서, 본 접근은 교육학에서 도구변수 활용의 범위를 자연스럽게 확장하는 데 방법적으로 기여할 수 있다.
이 세 연구는 관측자료에서의 인과적 해석을 가능하게 하는 도구변수 분석법의 구조를 교육 연구의 다양한 분석 맥락에 연결한다. 사전-사후검사 설계, 문항반응과 같은 교육 자료가 가지는 고유한 특성을 활용하여, 본 논문은 외생성 가정과 문항 공정성 가정을 추정량 간 동등성이라는 형태의 검정 가능한 귀납적 조건으로 정립한다. 이러한 접근은 단일 도구변수 상황에서도 외생성 위반을 경험적으로 탐지할 수 있도록 하고, 관찰되지 않은 교란변수가 존재하더라도 구조에서 유도되는 비교 기준을 통해 인과적인 정보를 식별할 수 있게 해준다는 점에서 기존 검정법을 보완한다. 또한 교육측정 및 평가 맥락에 도구변수의 과잉식별 논리를 도입함으로써, 능력 추정치를 조건화하지 않고도 문항편향 여부를 평가할 수 있는 새로운 접근을 제시한다. 종합적으로, 본 논문은 도구변수 활용이 교육 연구 전반에서 보다 투명하고 신뢰성 있게 이루어질 수 있도록 방법론적인 기반을 제공하고, 인과추론 이론과 측정 타당도 간의 연계 가능성을 확장함으로써 교육 연구의 실증적 엄밀성과 해석 가능성을 제고한다.
다국어 초록 (Multilingual Abstract)
Establishing credible causal inference regarding key constructs—such as academic achievement, student behavior, and school contexts—is a fundamental objective in education research. While randomized controlled trials remain the gold standard, they...
Establishing credible causal inference regarding key constructs—such as academic achievement, student behavior, and school contexts—is a fundamental objective in education research. While randomized controlled trials remain the gold standard, they are often limited by the ethical and practical constraints of applied educational settings. Consequently, traditional quasi-experimental strategies, including regression adjustment and propensity score methods, attempt to bridge this gap by emulating the logic of experimental designs with observational data. However, not all confounders are observable, and even careful conditioning on observed variables can induce new sources of bias, such as collider bias. To address these limitations, instrumental variable (IV) methods provide a useful alternative by isolating exogenous variation to estimate causal effects, even when unobserved confounding is a concern. By utilizing an instrument that affects the outcome only through the treatment, researchers can mitigate omitted variable bias and achieve more credible estimates of causal effect.
Despite its robust theoretical foundation and widespread adoption in the social sciences and epidemiology, the application and methodological development of IV methods in education research remain comparatively limited. A primary barrier is the difficulty of empirically assessing the exogeneity condition, which is essential for a causal interpretation of IV estimates. Standard diagnostics, such as endogeneity and overidentification tests, often fall short in applied research in education; they either require multiple instruments or rely on the very assumptions they are meant to test. In the presence of weak or potentially invalid instruments, these conventional diagnostics can be uninformative or even misleading, leaving researchers with limited practical guidance for evaluating the validity of their causal inferences.
To address these challenges, this dissertation introduces a structure-informed diagnostic framework that transforms the typically untestable assumption of exogeneity into empirically testable equality restrictions. The framework exploits causal structures inherent in commonly used educational datasets to derive an internal benchmark estimand. Under the null hypothesis of exogeneity, this benchmark should align with the conventional IV estimand. Thus, any systematic divergence between these estimands serves as falsification test indicating a violation of the maintained exogeneity assumption. This logic is formally developed for pretest–posttest designs—utilizing shared latent structures to adjust for confounding—and extended to measurement settings where multiple group indicators serve as competing instruments. The dissertation applies this framework to pretest–posttest settings and item-level measurement models, illustrating its utility in complex educational environments.
Study 1 develops the diagnostic for the pretest–posttest design, providing a practical strategy for interrogating exogeneity with a single instrument. Under a latent-factor structure linking pretest and posttest outcomes, the approach constructs a benchmark quantity that is theoretically comparable to the IV estimand under exogeneity. The discrepancy between the IV estimand and this benchmark serves as the diagnostic target, and a joint estimation strategy supports formal inference about whether the discrepancy is consistent with sampling variability. Simulation evidence suggests that the diagnostic is generally conservative when exogeneity holds, becomes increasingly informative as instrument relevance strengthens, and may become unstable when the identifying variation effectively collapses.
Study 2 applies the diagnostic to the National Education Longitudinal Study of 1988 (NELS:88) to reassess IV evidence in the long-standing debate over Catholic school effects. The study first reproduces a conventional two-stage least squares specification using multiple candidate instruments, and then applies the structure-informed diagnostic separately to each candidate, contrasting each conventional IV estimate with its corresponding benchmark implied by the pretest–posttest structure. The results show that conclusions can depend sharply on the particular source of identifying variation: for reading, discrepancies are not estimated precisely enough to isolate a single problematic candidate, whereas for mathematics the diagnostic is more discriminating and indicates that an urban-by-region interaction proxy contributes disproportionately to instability, while religion-based candidates appear more compatible with the pretest–posttest restriction. These findings highlight the value of structure-informed diagnostics for improving interpretive transparency in settings where school choice is closely intertwined with socioeconomic background.
Study 3 extends the structure-informed framework to measurement and introduces an overidentification-style IV strategy for diagnosing causal item bias in differential item functioning (DIF) analyses. Conventional DIF approaches commonly rely on conditioning on an estimated ability measure, which can be problematic when unobserved factors jointly influence the latent trait and item responses, potentially generating spurious DIF patterns. Study 3 instead treats multiple group indicators as instruments for latent ability and evaluates whether instrument-specific IV estimands agree. Under a causal fairness interpretation, valid group-based instruments should satisfy an exclusion-type requirement and yield consistent estimands; systematic disagreement provides falsification evidence that at least one group characteristic directly affects item responses in a manner consistent with item bias. Simulation results indicate that the diagnostic can maintain appropriate false-positive behavior in no-bias settings and can be meaningfully sensitive when item bias is driven by a single group characteristic, while sensitivity may diminish when multiple group characteristics simultaneously generate item bias. This study illustrates how IV reasoning can be adapted to measurement fairness questions and motivates further work on robustness to instrument strength and to practical details of ability estimation.
Together, these three studies demonstrate the value of structure-informed IV diagnostics for strengthening causal inference in education research. By leveraging structural features of educational data—such as pretest–posttest designs, school choice processes, and item response patterns---the dissertation transforms exogeneity and fairness assumptions into empirically checkable equality conditions. This unified framework helps researchers diagnose instrument invalidity in single-instrument settings and distinguish true item bias from spurious associations without conditioning on estimated ability, thereby enhancing the transparency, interpretability, and rigor of observational research on both program effects and measurement fairness.
목차 (Table of Contents)