본 논문은 기존의 조건화 기반 문항편향(item bias) 탐지 방법들이 지닌 한계를 비판적으로 분석하고, 인과추론 이론에 기반한 대안적 탐지 전략을 제시한다. 문항편향은 종종 DIF(differential item ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 논문은 기존의 조건화 기반 문항편향(item bias) 탐지 방법들이 지닌 한계를 비판적으로 분석하고, 인과추론 이론에 기반한 대안적 탐지 전략을 제시한다. 문항편향은 종종 DIF(differential item ...
본 논문은 기존의 조건화 기반 문항편향(item bias) 탐지 방법들이 지닌 한계를 비판적으로 분석하고, 인과추론 이론에 기반한 대안적 탐지 전략을 제시한다. 문항편향은 종종 DIF(differential item functioning)와 혼용되나, 동일한 잠재특성을 가진 응답자들이 집단 특성 때문에 특정 문항에 차별적으로 반응하는 현상을 의미한다. 기존 통계적 탐지 방법들은 문항편향 발생의 구조적 또는 인과적 원인을 충분히 반영하지 못하며, 가정 위반과 구조적 위험요인으로 인해 허위 탐지 오류가 빈번히 발생하는 문제점이 있다.
본 연구는 다음 네 가지 핵심 문제를 중심으로 진행된다: (a) 기존 조건화 기반 접근법이 실질적인 가정 위반 상황(타당하지 않은 대응 변수, 신뢰도 낮은 잠재 특성 추정치, 구조적 위험요인 등)에서 얼마나 높은 허위 탐지율을 보이는지, (b) DID (difference-in-differences) 접근법이 어떤 식별 조건 하에서 이진 반응(binary response)의 문항편향을 정확히 추정할 수 있는지, (c) SLFM (Structural Latent Factor Model) 및 aDID (adjusted DID) 접근법이 다양한 구조적 조건에서 허위 탐지율을 효과적으로 통제하면서도 높은 탐지력을 유지할 수 있는지, (d) 탐지 변수(compass variable)에 요구되는 지역독립성(local independence) 가정을 실증적으로 검증할 수 있는지에 중점을 두었다.
문헌 고찰 및 몬테카를로 시뮬레이션 분석을 통해 기존 조건화 기반 접근법의 논리적 구조와 실질적 한계를 파악하였으며, (a) 타당하지 않은 대응 변수, (b) 신뢰도 낮은 잠재 특성 추정치, (c) 구조적 위험요인(미관측 교란변수 및 충돌변수 구조)이라는 세 가지 주요 메커니즘이 허위 탐지율을 유의하게 증가시킴을 확인하였다. 이러한 한계를 극복하기 위한 인과추론 기반 대안으로 SLFM, DID, 그리고 탐지 변수를 활용한 aDID 접근법을 제안하고, 탐지 변수의 지역독립성 가정을 실증적으로 검증할 수 있는 방법도 함께 개발하였다.
SLFM 접근법은 집단 변수의 효과가 전적으로 잠재 특성을 통해 매개된다는 구조적 가정 하에 검증 가능한 함의를 도출하고, 부트스트랩 기반 검정 절차를 통해 이진 반응 문항에 적용되었다. DID 접근법은 문항 모수 동등성 가정 하에서 문항편향을 확률 척도로 추정할 수 있음을 보였으며, aDID 접근법은 탐지 변수를 활용하여 문항 모수 동등성 가정이 위반되는 경우에도 문항편향을 정확히 탐지할 수 있음을 보여주었다. 또한, 탐지 변수의 지역독립성 가정을 검증할 수 있는 방법도 함께 개발하였다.
몬테카를로 시뮬레이션 분석 결과, SLFM과 aDID 접근법은 다양한 가정 위반 상황에서도 낮은 허위 탐지율과 높은 탐지력을 유지하며, 기존 조건화 기반 방법(Mantel-Haenszel, 로지스틱 회귀, Lord 카이제곱 검정, SIBTEST)보다 안정적인 성능을 보였다. DID 역시 적절한 식별 조건 하에서 문항편향을 확률 척도에서 정확하게 해석할 수 있음이 확인되었다. 아울러, 탐지 변수 검정 또한 지역독립성 가정 위반 여부를 효과적으로 판별할 수 있음을 확인하였다. TIMSS 실증 자료 분석에서도 SLFM과 aDID 접근법의 실용적 유용성을 보여주었다. 이 적용 예시는 두 방법 모두 실제 데이터에 쉽게 적용할 수 있으며, 기존 문항편향 분석 절차에 무리 없이 통합될 수 있음을 보여준다.
종합적으로, 본 논문은 문항편향 분석에서 기존 조건화 기반 통계 접근법에서 인과추론적 접근으로의 전환 필요성을 제시하며, (a) 기존 접근법의 측정적 및 구조적 한계를 인과적 관점에서 비판하고, (b) SLFM 및 aDID 접근법이라는 대안적 방법을 개발하여 잠재 특성에 조건화없이 문항편향 탐지가 가능함을 보였으며, (c) 탐지 변수 가정 검증을 위한 실용적 진단 도구를 제시하였다. 본 연구는 교육 및 심리 평가에서 문항 공정성 확보를 위한 새로운 분석 틀을 제시한다.
다국어 초록 (Multilingual Abstract)
This dissertation investigates the limitations of conventional conditioning-based approaches to differential item functioning (DIF) analysis for detecting item bias and proposes alternative identification strategies grounded in causal inference. Item ...
This dissertation investigates the limitations of conventional conditioning-based approaches to differential item functioning (DIF) analysis for detecting item bias and proposes alternative identification strategies grounded in causal inference. Item bias—often mistakenly conflated with DIF—occurs when the different groups respond differently to an item, not because of differences in the underlying trait, but because of group membership. Although various conditioning-based statistical methods exist for detecting item bias, their underlying mechanisms are often taken for granted and rarely examined from a structural or causal perspective.
Central to this work is a critical examination of conditioning-based methods and the development of causal inference-based alternatives. This dissertation explores four key issues: (a) to what extent conventional methods exhibit inflated false detection rates under practical violations, including invalid matching, unreliable trait estimates, and structural risks; (b) under what identification conditions the difference-in-differences (DID) method can be applied to binary item bias detection, and whether the DID can accurately recover item bias on the probability scale; (c) whether causal inference-based approaches, specifically structural latent factor model (SLFM) and adjusted DID (aDID), can effectively control false detection rates while maintaining strong detection power; and (d) whether the local independence assumption required for compass variables can be empirically tested. Through formal methodological development and extensive Monte Carlo simulations, this dissertation establishes a principled and causally valid approach to improving item bias detection.
This literature review critically examines conditioning-based methods for detecting item bias, focusing on their conceptual foundations and practical limitations. It highlights three key mechanisms that may lead to false detection of item bias: (a) invalid matching variables; (b) unreliable trait estimates; and (c) structural risks such as unmeasured confounding and collider structures. In response to these limitations, recent causal inference-based approaches—specifically the SLFM and DID methods, which avoid conditioning on latent traits—are evaluated for their potential advantages.
Building on this foundation, this dissertation formally develops and extends causal inference-based methods for detecting item bias in binary response data. Specifically: (a) the SLFM approach is adapted by deriving testable implications under the structural assumption that group membership influences item responses solely through the latent trait, with inference procedures implemented using bootstrap methods; (b) the DID method is extended through the item parameter equivalence assumption, enabling consistent estimation of item bias on the probability scale when the assumption holds; and (c) an aDID method is developed using compass variables, along with a test for empirically evaluating the required local independence assumption. Formal identification results and testable implications for these methods are established through detailed derivations and theorems.
Monte Carlo simulations systematically evaluate the proposed methods across a range of realistic data-generating conditions where key assumptions of conventional methods are often violated. The simulations assess: (a) the limitations of conditioning-based methods, including inflated false detection rates under invalid matching variables, unreliable trait estimates, and structural risks; (b) the accuracy of the DID estimator in recovering item bias on the probability scale; (c) the performance of the SLFM and aDID approaches in controlling false detection rates while maintaining strong detection power; and (d) the performance of the proposed test for evaluating the local independence assumption of compass variables. The results show that SLFM and aDID consistently maintain low false detection rates in no item bias conditions and demonstrate strong power in detecting true item bias across diverse structural conditions, including a basic item bias, a hidden common cause, and a latent trait as a collider, substantially outperforming conventional conditioning-based methods.
Empirical illustrations using data from the Trends in International Mathematics and Science Study further demonstrate the practical utility of the SLFM and aDID approaches. These illustrations show that both methods can be readily applied to real-world data and integrated into existing item bias analysis workflows with relative ease.
Overall, this dissertation supports a shift from conventional statistical approaches toward a causal framework for detecting item bias. It contributes: (a) a causal critique of conditioning-based methods; (b) two alternative methods—SLFM and aDID—that avoid conditioning on latent traits; and (c) practical diagnostic tools, such as the compass validity test, to empirically verify the key assumption. By applying causal theory and testable implications to item bias analysis, this work offers an alternative perspective on evaluating item fairness in educational and psychological testing.
목차 (Table of Contents)