이 학위논문은 연속–범주형이 혼합된 자료(mixed–type, continuous–categorical) 에 대해 매개변수 없이 사용할 수 있는 비모수적 결합(association) 측도를 제안 하고, 그 대표본 이론을 다음의 세 축...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
이 학위논문은 연속–범주형이 혼합된 자료(mixed–type, continuous–categorical) 에 대해 매개변수 없이 사용할 수 있는 비모수적 결합(association) 측도를 제안 하고, 그 대표본 이론을 다음의 세 축...
이 학위논문은 연속–범주형이 혼합된 자료(mixed–type, continuous–categorical)
에 대해 매개변수 없이 사용할 수 있는 비모수적 결합(association) 측도를 제안
하고, 그 대표본 이론을 다음의 세 축에서 확립한다: (i) 일변량 연속–범주형 계수
𝜉′ (𝑋, 𝑌 ), (ii) 𝑋 ∈ R𝑝에 대한 다변량 확장 𝑇′ (𝑋, 𝑌 ), (iii) 조건부 버전 𝑇′ (𝑋, 𝑌 | 𝑍).
𝜉′에 대해서는, 조건부 클래스 확률과 𝑋의 지지집합에 대한 약한 정규성 하에
서 강한 일치성(strong consistency)을 보이며, 표본 통계량이 모수(target)로 거의
확실히 수렴하고 귀무가설 𝑋 ⊥ 𝑌 하에서의 영가설 중심 극한정규성을 이용한
검정을 제시한다. 다변량 𝑋의 경우, 최근접 이웃(Nearest–Neighbor) 라벨 일치율
에 기반한 𝑇′를 도입하고 𝑇′
𝑛
𝑎.𝑠.
−−−→ 𝑇′ (𝑋, 𝑌 )를 증명한다. 이때 𝑇′는 보정된 [0, 1]
스케일을 유지하며, 𝑇′ = 0 iff 𝑋 ⊥ 𝑌 , 𝑇′ = 1 iff 𝑌 = 𝑓 (𝑋) a.s.라는 성질을 갖는다.
조건부 통계량 𝑇′ (𝑋, 𝑌 | 𝑍)에 대해서는, 고차원에서 최근접 이웃 구성의 기하학
적 요인으로 인한 바이어스를 설명하고, (dim 𝑋, dim 𝑍) = (1, 1)인 특수한 경우
에 명시적 분산 분해를 갖는 중심극한정리를 수립한다. 또한 두 가지 실용적 분
산 추정기를 제안한다: 추정된 b𝑔에 조건한 라벨 재표집(conditional resampling)
기반 추정과, NN 그래프의 패턴에 대한 닫힌형 플러그인(closed–form pattern
plug–in) 추정이다. 아울러, 모집단 성질을 훼손하지 않으면서 유한표본 하강
편의를 실증적으로 줄이는 간단한 스케일 매칭 전처리를 권고한다.
광범위한 모의실험은 영가설 하 보정의 정확성(null calibration), 최신 대안
들과 견줄 만한 검정력, 그리고 우수한 실행 시간 확장성을 확인한다. 종합하면,
본 연구는 범주형 결과, 다변량 예측변수, 공변량 보정을 포함한 설정에서 비모
수적 결합과 조건부 독립성 평가를 위한 원리적이면서도 계산 친화적인 도구
상자를 제공한다.
다국어 초록 (Multilingual Abstract)
This dissertation develops a family of nonparametric association measures for mixed–type (continuous - categorical) data and establishes their large–sample the- ory along three axes: (i) a univariate continuous–categorical coefficient 𝜉...
This dissertation develops a family of nonparametric association measures for
mixed–type (continuous - categorical) data and establishes their large–sample the-
ory along three axes: (i) a univariate continuous–categorical coefficient 𝜉′ (𝑋, 𝑌 ),
(ii) a multivariate extension 𝑇′ (𝑋, 𝑌 ) for 𝑋 ∈ R𝑝, and (iii) a conditional version
𝑇′ (𝑋, 𝑌 | 𝑍).
For 𝜉′ (𝑋, 𝑌 ), we prove strong consistency under mild regularity on the condi-
tional class probabilities and the support of 𝑋, yielding almost–sure convergence of
the sample statistic to its population target and a null central limit theorem for test-
ing 𝑋 ⊥ 𝑌 . For multivariate 𝑋, we introduce 𝑇′ (𝑋, 𝑌 ) based on nearest–neighbor
label matches and show 𝑇′
𝑛
𝑎.𝑠.
−−−→ 𝑇′ (𝑋, 𝑌 ), preserving a calibrated [0, 1] scale
with the characterizations 𝑇′ = 0 iff 𝑋 ⊥ 𝑌 and 𝑇′ = 1 iff 𝑌 = 𝑓 (𝑋) a.s. We
further study the conditional statistic 𝑇′ (𝑋, 𝑌 | 𝑍), explaining a geometry–driven
i
bias of nearest–neighbor constructions in higher dimensions, establishing a spe-
cial–case CLT when the dimension of 𝑋 and 𝑍 are both 1. We calculate the
asymptotic variance, and propose two practical variance estimators. First, A direct
conditional plug-in that evaluates edgewise conditional covariances using the fitted
class probabilities and nearest-neighbor overlaps. Second, a closed-form, pattern-
pooled plug-in on the nearest-neighbor graph that aggregates anchor-based overlap
types and plugs their empirical frequencies. We also recommend a simple vari-
ance scale–matching preprocessing that empirically mitigates finite–sample drift
without altering population properties.
Comprehensive simulations demonstrate accurate null calibration, competi-
tive power against modern alternatives, and favorable run–time scaling. Together,
these results provide a principled, computation–friendly toolkit for nonparametric
association and conditional independence assessment with categorical outcomes
목차 (Table of Contents)