RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Nonparametric Association Measures for Mixed-Type Data with Multivariate and Conditional Extensions = 다변량 · 조건부 확장을 포함한 혼합형 데이터의 비모수 연관성

    한글로보기

    https://www.riss.kr/link?id=T17450711

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    이 학위논문은 연속–범주형이 혼합된 자료(mixed–type, continuous–categorical)
    에 대해 매개변수 없이 사용할 수 있는 비모수적 결합(association) 측도를 제안
    하고, 그 대표본 이론을 다음의 세 축에서 확립한다: (i) 일변량 연속–범주형 계수
    𝜉′ (𝑋, 𝑌 ), (ii) 𝑋 ∈ R𝑝에 대한 다변량 확장 𝑇′ (𝑋, 𝑌 ), (iii) 조건부 버전 𝑇′ (𝑋, 𝑌 | 𝑍).
    𝜉′에 대해서는, 조건부 클래스 확률과 𝑋의 지지집합에 대한 약한 정규성 하에
    서 강한 일치성(strong consistency)을 보이며, 표본 통계량이 모수(target)로 거의
    확실히 수렴하고 귀무가설 𝑋 ⊥ 𝑌 하에서의 영가설 중심 극한정규성을 이용한
    검정을 제시한다. 다변량 𝑋의 경우, 최근접 이웃(Nearest–Neighbor) 라벨 일치율
    에 기반한 𝑇′를 도입하고 𝑇′
    𝑛
    𝑎.𝑠.
    −−−→ 𝑇′ (𝑋, 𝑌 )를 증명한다. 이때 𝑇′는 보정된 [0, 1]
    스케일을 유지하며, 𝑇′ = 0 iff 𝑋 ⊥ 𝑌 , 𝑇′ = 1 iff 𝑌 = 𝑓 (𝑋) a.s.라는 성질을 갖는다.
    조건부 통계량 𝑇′ (𝑋, 𝑌 | 𝑍)에 대해서는, 고차원에서 최근접 이웃 구성의 기하학
    적 요인으로 인한 바이어스를 설명하고, (dim 𝑋, dim 𝑍) = (1, 1)인 특수한 경우
    에 명시적 분산 분해를 갖는 중심극한정리를 수립한다. 또한 두 가지 실용적 분
    산 추정기를 제안한다: 추정된 b𝑔에 조건한 라벨 재표집(conditional resampling)
    기반 추정과, NN 그래프의 패턴에 대한 닫힌형 플러그인(closed–form pattern
    plug–in) 추정이다. 아울러, 모집단 성질을 훼손하지 않으면서 유한표본 하강
    편의를 실증적으로 줄이는 간단한 스케일 매칭 전처리를 권고한다.
    광범위한 모의실험은 영가설 하 보정의 정확성(null calibration), 최신 대안
    들과 견줄 만한 검정력, 그리고 우수한 실행 시간 확장성을 확인한다. 종합하면,
    본 연구는 범주형 결과, 다변량 예측변수, 공변량 보정을 포함한 설정에서 비모
    수적 결합과 조건부 독립성 평가를 위한 원리적이면서도 계산 친화적인 도구
    상자를 제공한다.
    번역하기

    이 학위논문은 연속–범주형이 혼합된 자료(mixed–type, continuous–categorical) 에 대해 매개변수 없이 사용할 수 있는 비모수적 결합(association) 측도를 제안 하고, 그 대표본 이론을 다음의 세 축...

    이 학위논문은 연속–범주형이 혼합된 자료(mixed–type, continuous–categorical)
    에 대해 매개변수 없이 사용할 수 있는 비모수적 결합(association) 측도를 제안
    하고, 그 대표본 이론을 다음의 세 축에서 확립한다: (i) 일변량 연속–범주형 계수
    𝜉′ (𝑋, 𝑌 ), (ii) 𝑋 ∈ R𝑝에 대한 다변량 확장 𝑇′ (𝑋, 𝑌 ), (iii) 조건부 버전 𝑇′ (𝑋, 𝑌 | 𝑍).
    𝜉′에 대해서는, 조건부 클래스 확률과 𝑋의 지지집합에 대한 약한 정규성 하에
    서 강한 일치성(strong consistency)을 보이며, 표본 통계량이 모수(target)로 거의
    확실히 수렴하고 귀무가설 𝑋 ⊥ 𝑌 하에서의 영가설 중심 극한정규성을 이용한
    검정을 제시한다. 다변량 𝑋의 경우, 최근접 이웃(Nearest–Neighbor) 라벨 일치율
    에 기반한 𝑇′를 도입하고 𝑇′
    𝑛
    𝑎.𝑠.
    −−−→ 𝑇′ (𝑋, 𝑌 )를 증명한다. 이때 𝑇′는 보정된 [0, 1]
    스케일을 유지하며, 𝑇′ = 0 iff 𝑋 ⊥ 𝑌 , 𝑇′ = 1 iff 𝑌 = 𝑓 (𝑋) a.s.라는 성질을 갖는다.
    조건부 통계량 𝑇′ (𝑋, 𝑌 | 𝑍)에 대해서는, 고차원에서 최근접 이웃 구성의 기하학
    적 요인으로 인한 바이어스를 설명하고, (dim 𝑋, dim 𝑍) = (1, 1)인 특수한 경우
    에 명시적 분산 분해를 갖는 중심극한정리를 수립한다. 또한 두 가지 실용적 분
    산 추정기를 제안한다: 추정된 b𝑔에 조건한 라벨 재표집(conditional resampling)
    기반 추정과, NN 그래프의 패턴에 대한 닫힌형 플러그인(closed–form pattern
    plug–in) 추정이다. 아울러, 모집단 성질을 훼손하지 않으면서 유한표본 하강
    편의를 실증적으로 줄이는 간단한 스케일 매칭 전처리를 권고한다.
    광범위한 모의실험은 영가설 하 보정의 정확성(null calibration), 최신 대안
    들과 견줄 만한 검정력, 그리고 우수한 실행 시간 확장성을 확인한다. 종합하면,
    본 연구는 범주형 결과, 다변량 예측변수, 공변량 보정을 포함한 설정에서 비모
    수적 결합과 조건부 독립성 평가를 위한 원리적이면서도 계산 친화적인 도구
    상자를 제공한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This dissertation develops a family of nonparametric association measures for
    mixed–type (continuous - categorical) data and establishes their large–sample the-
    ory along three axes: (i) a univariate continuous–categorical coefficient 𝜉′ (𝑋, 𝑌 ),
    (ii) a multivariate extension 𝑇′ (𝑋, 𝑌 ) for 𝑋 ∈ R𝑝, and (iii) a conditional version
    𝑇′ (𝑋, 𝑌 | 𝑍).
    For 𝜉′ (𝑋, 𝑌 ), we prove strong consistency under mild regularity on the condi-
    tional class probabilities and the support of 𝑋, yielding almost–sure convergence of
    the sample statistic to its population target and a null central limit theorem for test-
    ing 𝑋 ⊥ 𝑌 . For multivariate 𝑋, we introduce 𝑇′ (𝑋, 𝑌 ) based on nearest–neighbor
    label matches and show 𝑇′
    𝑛
    𝑎.𝑠.
    −−−→ 𝑇′ (𝑋, 𝑌 ), preserving a calibrated [0, 1] scale
    with the characterizations 𝑇′ = 0 iff 𝑋 ⊥ 𝑌 and 𝑇′ = 1 iff 𝑌 = 𝑓 (𝑋) a.s. We
    further study the conditional statistic 𝑇′ (𝑋, 𝑌 | 𝑍), explaining a geometry–driven
    i
    bias of nearest–neighbor constructions in higher dimensions, establishing a spe-
    cial–case CLT when the dimension of 𝑋 and 𝑍 are both 1. We calculate the
    asymptotic variance, and propose two practical variance estimators. First, A direct
    conditional plug-in that evaluates edgewise conditional covariances using the fitted
    class probabilities and nearest-neighbor overlaps. Second, a closed-form, pattern-
    pooled plug-in on the nearest-neighbor graph that aggregates anchor-based overlap
    types and plugs their empirical frequencies. We also recommend a simple vari-
    ance scale–matching preprocessing that empirically mitigates finite–sample drift
    without altering population properties.
    Comprehensive simulations demonstrate accurate null calibration, competi-
    tive power against modern alternatives, and favorable run–time scaling. Together,
    these results provide a principled, computation–friendly toolkit for nonparametric
    association and conditional independence assessment with categorical outcomes
    번역하기

    This dissertation develops a family of nonparametric association measures for mixed–type (continuous - categorical) data and establishes their large–sample the- ory along three axes: (i) a univariate continuous–categorical coefficient 𝜉...

    This dissertation develops a family of nonparametric association measures for
    mixed–type (continuous - categorical) data and establishes their large–sample the-
    ory along three axes: (i) a univariate continuous–categorical coefficient 𝜉′ (𝑋, 𝑌 ),
    (ii) a multivariate extension 𝑇′ (𝑋, 𝑌 ) for 𝑋 ∈ R𝑝, and (iii) a conditional version
    𝑇′ (𝑋, 𝑌 | 𝑍).
    For 𝜉′ (𝑋, 𝑌 ), we prove strong consistency under mild regularity on the condi-
    tional class probabilities and the support of 𝑋, yielding almost–sure convergence of
    the sample statistic to its population target and a null central limit theorem for test-
    ing 𝑋 ⊥ 𝑌 . For multivariate 𝑋, we introduce 𝑇′ (𝑋, 𝑌 ) based on nearest–neighbor
    label matches and show 𝑇′
    𝑛
    𝑎.𝑠.
    −−−→ 𝑇′ (𝑋, 𝑌 ), preserving a calibrated [0, 1] scale
    with the characterizations 𝑇′ = 0 iff 𝑋 ⊥ 𝑌 and 𝑇′ = 1 iff 𝑌 = 𝑓 (𝑋) a.s. We
    further study the conditional statistic 𝑇′ (𝑋, 𝑌 | 𝑍), explaining a geometry–driven
    i
    bias of nearest–neighbor constructions in higher dimensions, establishing a spe-
    cial–case CLT when the dimension of 𝑋 and 𝑍 are both 1. We calculate the
    asymptotic variance, and propose two practical variance estimators. First, A direct
    conditional plug-in that evaluates edgewise conditional covariances using the fitted
    class probabilities and nearest-neighbor overlaps. Second, a closed-form, pattern-
    pooled plug-in on the nearest-neighbor graph that aggregates anchor-based overlap
    types and plugs their empirical frequencies. We also recommend a simple vari-
    ance scale–matching preprocessing that empirically mitigates finite–sample drift
    without altering population properties.
    Comprehensive simulations demonstrate accurate null calibration, competi-
    tive power against modern alternatives, and favorable run–time scaling. Together,
    these results provide a principled, computation–friendly toolkit for nonparametric
    association and conditional independence assessment with categorical outcomes

    더보기

    목차 (Table of Contents)

    • Contents
    • Abstract i
    • 1 Introduction 1
    • 2 An association measure for mixed-types variables 6
    • 2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
    • Contents
    • Abstract i
    • 1 Introduction 1
    • 2 An association measure for mixed-types variables 6
    • 2.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 6
    • 2.2 A coefficient for continuous-categorical association . . . . . . . . 10
    • 2.2.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . 10
    • 2.2.2 Basic properties . . . . . . . . . . . . . . . . . . . . . . . 15
    • 2.2.3 Connections to the runs statistic . . . . . . . . . . . . . . 18
    • 2.3 Asymptotic properties of 𝜉′
    • 𝑛 . . . . . . . . . . . . . . . . . . . . 21
    • 2.3.1 Assumptions . . . . . . . . . . . . . . . . . . . . . . . . 21
    • 2.3.2 Convergence to the population coefficient . . . . . . . . . 23
    • 2.3.3 Asymptotic normality under independence . . . . . . . . 36
    • 2.3.4 A test of independence for continuous-categorical variables 41
    • 2.4 Simulation studies . . . . . . . . . . . . . . . . . . . . . . . . . . 44
    • iii
    • 2.4.1 A calibration and variance estimation under independence 44
    • 2.4.2 Stability under label permutations . . . . . . . . . . . . . 46
    • 2.4.3 Power and Run Time Comparisons . . . . . . . . . . . . . 48
    • 2.5 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
    • 3 An extension to multivariate continuous variable 54
    • 3.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 54
    • 3.2 A coefficient for continuous-categorical association: multivariate
    • extension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 59
    • 3.2.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . 59
    • 3.2.2 Basic properties . . . . . . . . . . . . . . . . . . . . . . . 61
    • 3.3 Asymptotic properties of 𝑇′
    • 𝑛 . . . . . . . . . . . . . . . . . . . . 67
    • 3.3.1 Assumptions . . . . . . . . . . . . . . . . . . . . . . . . 67
    • 3.3.2 Convergence to the population coefficient . . . . . . . . . 69
    • 3.3.3 Asymptotic normality under independence . . . . . . . . 73
    • 3.4 Simulation Studies . . . . . . . . . . . . . . . . . . . . . . . . . 80
    • 3.4.1 Asymptotic distribution under independence . . . . . . . 80
    • 3.4.2 Test power comparisons . . . . . . . . . . . . . . . . . . 82
    • 3.5 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 93
    • 4 An extension to conditional setting 98
    • 4.1 Introduction . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 98
    • 4.2 A coefficient for continuous-categorical association: conditional
    • extension . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 104
    • 4.2.1 Definition . . . . . . . . . . . . . . . . . . . . . . . . . . 105
    • iv
    • 4.2.2 Basic properties . . . . . . . . . . . . . . . . . . . . . . . 107
    • 4.3 Asymptotic properties of 𝑇′
    • 𝑛 . . . . . . . . . . . . . . . . . . . . 112
    • 4.3.1 Assumptions . . . . . . . . . . . . . . . . . . . . . . . . 112
    • 4.3.2 Convergence to the population coefficient . . . . . . . . . 114
    • 4.4 Asymptotic behavior of 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍) under conditional indepen-
    • dence . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 119
    • 4.4.1 Bias from the nearest neighbor method . . . . . . . . . . 119
    • 4.4.2 Special asymptotic case: CLT for 𝑝 = 𝑞 = 1 . . . . . . . . 128
    • 4.4.3 Estimation of variance in the asymptotic normal distribution152
    • 4.5 Simulations . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 156
    • 4.5.1 Comparing two variance estimators: direct vs. pattern . . . 156
    • 4.5.2 Power comparisons . . . . . . . . . . . . . . . . . . . . . 158
    • 4.5.3 Real data . . . . . . . . . . . . . . . . . . . . . . . . . . 161
    • 4.6 Discussion . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 164
    • Abstract (in Korean) 178
    • List of Tables
    • 2.1 Empirical type I error (one-sided, 𝛼 = 0.05) under 𝐻0 (𝜃 = 0).
    • Each entry is the rejection rate over 1000 replicates; permutation
    • tests use 𝐵perm = 199. . . . . . . . . . . . . . . . . . . . . . . . . 44
    • 2.2 Sample means (standard deviations) of 𝜉′
    • 𝑛 and 𝜉𝑛 over 200 repli-
    • cates for (𝑛, 𝜃). . . . . . . . . . . . . . . . . . . . . . . . . . . . 46
    • 2.3 Null calibration of √𝑛 𝜉′
    • 𝑛 for 𝑘 = 6 under 𝐻0. Reported are the
    • mean plug-in variance b𝜅2, the empirical mean and variance of
    • √𝑛 𝜉′
    • 𝑛, and the Shapiro–Wilk 𝑝-value (based on 𝐵 = 2000 repli-
    • cates per 𝑛). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 48
    • 2.4 Average empirical power at 𝛼 = 0.05 across 200 replicates. . . . . 50
    • 2.5 Average runtimes (seconds) per replicate, 200 reps, 200 permuta-
    • tions. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 51
    • 3.1 Summary statistics of √𝑛 𝑇′
    • 𝑛 under independence. Reported are
    • the empirical mean, empirical standard deviation, 95% confidence
    • interval for the mean, plug–in standard deviation, and Shapiro–
    • Wilk normality 𝑝–value, based on 𝐵 = 1000 replicates. . . . . . . 81
    • vi
    • 3.2 Empirical Type I error (𝜃 = 0, 𝛼 = 0.05) from each runs. Dashes
    • indicate undefined MANOVA in 𝑝 ≫ 𝑛. . . . . . . . . . . . . . . 93
    • 3.3 Average time per replicate (seconds) for Design B1 (micro-clusters). 94
    • 3.4 Average time per replicate (seconds) for Design B2 (noisy circle
    • sectors). . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 95
    • 3.5 Average time per replicate (seconds) for Design B3 (high-dimensional
    • sparse prototypes). . . . . . . . . . . . . . . . . . . . . . . . . . 97
    • 4.1 Monte Carlo means and standard deviations of √𝑛 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍)
    • over 200 replications, under dependence (𝑌 depends on 𝑍) and
    • independence (𝑌 ⊥ 𝑍). . . . . . . . . . . . . . . . . . . . . . . . 124
    • 4.2 Monte Carlo means and variances of √𝑛 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍) with and
    • without variance scaling ( 𝑝 = 𝑞 = 1 ). . . . . . . . . . . . . . . . 128
    • 4.3 Monte Carlo means and variances of √𝑛 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍) over 1000
    • replications under 𝑋 ⊥ 𝑌
    • 𝑍. . . . . . . . . . . . . . . . . . . . . 157
    • 4.4 Empirical size (𝛾 = 0, 𝛼 = 0.05). . . . . . . . . . . . . . . . . . . 161
    • 4.5 Empirical power at 𝛾 = 0.2 (𝛼 = 0.05). . . . . . . . . . . . . . . . 161
    • 4.6 Empirical power at 𝛾 = 0.4 (𝛼 = 0.05). . . . . . . . . . . . . . . . 162
    • 4.7 Empirical power at 𝛾 = 0.6 (𝛼 = 0.05). . . . . . . . . . . . . . . . 162
    • 4.8 Wage data: 𝐻0 : 𝑋 ⊥ 𝑌
    • 𝑍 with 𝑋 = log(wage), 𝑌 = education,
    • 𝑍 = age. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 163
    • vii
    • List of Figures
    • 2.1 Comparison of two integer encodings (codeA and codeB) when
    • binning continuous 𝑋 into 15 categories 𝐴–𝑂. Chatterjee’s statistic
    • 𝜉𝑛 depends on the arbitrary coding (0.971 vs. 0.788), whereas the
    • proposed statistic 𝜉′
    • 𝑛 remains invariant (0.903). . . . . . . . . . . 9
    • 2.2 Scatterplot of 𝑋 versus 𝑌 for 𝑛 = 1000, 𝑘 = 5, 𝑟 = 0.8. Each
    • 𝑌 value is deterministically given by 𝑋, yet the rapid oscillation
    • produces apparent randomness. . . . . . . . . . . . . . . . . . . . 14
    • 2.3 Scatterplots of (𝑋, 𝑌 ) for the largest 𝑛 in each 𝜃 setting (categor-
    • ical 𝑌 shown as horizontal bands by class). As 𝜃 increases, label
    • homogeneity along 𝑋 becomes more noticeable. . . . . . . . . . . 43
    • 2.4 Power versus 𝑛 for 𝜃 ∈ {0.25, 0.5, 0.75, 1}, comparing 𝜉′
    • 𝑛 (asymp-
    • totic and permutation) and 𝜉𝑛 (asymptotic and permutation). . . . 45
    • 2.5 Null calibration for 𝑘 = 6. Histograms of √𝑛 𝜉′
    • 𝑛 with two overlays:
    • plug-in 𝑁 (0, b𝜅2) (solid blue) and empirical 𝑁 ( ¯𝑦, 𝑠2) (red dashed). 47
    • viii
    • 2.6 Distribution of 𝜉𝑛 across 6! random label permutations at 𝜃 = 0.5,
    • 𝑛 = 100, 𝑘 = 6. Solid lines: observed 𝜉𝑛 (black) and 𝜉′
    • 𝑛 (green).
    • Dotted lines: the critical value of 𝜉𝑛 and the critical value of 𝜉′
    • 𝑛 for
    • one-sided permutation tests at 𝛼 = 0.05, computed by shuffling 𝑌 . 49
    • 3.1 Illustration of a common setting where 𝑌 is categorical (disease
    • stage) and 𝑋 = (𝑋1, 𝑋2) are continuous biomarkers. Each color
    • corresponds to a stage, showing how multivariate biomarker panels
    • can predict categorical outcomes. . . . . . . . . . . . . . . . . . . 56
    • 3.2 Scatterplot of the simulated data under the two-group design.
    • Group 1 occupies diagonal clusters (±𝑎, ±𝑎), and group 2 oc-
    • cupies off-diagonal clusters (±𝑎, ∓𝑎). MANOVA fails to detect
    • dependence (𝑝 = 0.34), while 𝑇′
    • 𝑛 achieves the maximum value of 1. 58
    • 3.3 Checkerboard example with 𝑛 = 200 and 𝑚 ≍ 𝑛1/2. The two
    • colors indicate binary labels 𝑌 ∈ {1, 2}. Although 𝑌 = 𝑓 (𝑋)
    • almost surely, the rapid oscillations of the boundaries create many
    • nearest–neighbor pairs with different labels, preventing 𝑇′
    • 𝑛 from
    • converging to one. . . . . . . . . . . . . . . . . . . . . . . . . . . 67
    • 3.4 Empirical distribution of √𝑛 𝑇′
    • 𝑛 under 𝐻0 : 𝑋 ⊥ 𝑌 . Each panel
    • shows a histogram (black) overlaid with the empirical Gaussian fit
    • (blue) and the plug–in normal (red). Simulation settings: 𝑝 = 5,
    • 𝑌 ∼ Multinomial(0.20, 0.30, 0.25, 0.25), 𝐵 = 1000. . . . . . . . . 82
    • ix
    • 3.5 B1: micro–clusters (𝑝 = 2, 𝑘 = 4, 𝑚 = 24). Columns show
    • 𝜃 = 0, 0.5, 1. Colors indicate class labels 𝑌 . At 𝜃 = 0 the classes
    • share the same uniform mixture (so 𝑋 ⊥𝑌 ); larger 𝜃 concentrates
    • mass on class–specific blobs. . . . . . . . . . . . . . . . . . . . . 86
    • 3.6 B2: noisy circle sectors (𝑝 = 2, 𝑘 = 6). Columns show 𝜃 =
    • 0, 0.5, 1. Colors indicate class labels 𝑌 . At 𝜃 = 0 labels are
    • uniform in angle; as 𝜃 grows, labels align with angular sectors,
    • creating local class structure. . . . . . . . . . . . . . . . . . . . . 87
    • 3.7 B3: sparse prototypes (visualized in 2D; 𝑝 = 50, 𝑘 = 6). Columns
    • show 𝜃 = 0, 0.5, 1. Colors indicate class labels 𝑌 . Increasing 𝜃
    • moves samples toward class-specific prototype directions, yielding
    • sharper separation. . . . . . . . . . . . . . . . . . . . . . . . . . 88
    • 3.8 B1: micro–clusters (𝑝 = 2, 𝑘 = 4). Empirical power (𝛼 = 0.05)
    • versus 𝜃; facets show 𝑛 ∈ {50, 100, 200, 400}. Curves compare 𝑇′
    • 𝑛
    • (perm), HSIC (perm), dCor (perm), and MANOVA (perm). . . . . 90
    • 3.9 B2: noisy circle sectors (𝑝 = 2, 𝑘 = 6). Empirical power (𝛼 =
    • 0.05) versus 𝜃; facets show 𝑛 ∈ {50, 100, 200, 400}. Curves com-
    • pare 𝑇′
    • 𝑛 (perm), HSIC (perm), dCor (perm), and MANOVA (perm). 91
    • 3.10 B3: sparse prototypes (𝑝 = 50, 𝑘 = 6). Empirical power (𝛼 =
    • 0.05) versus 𝜃; facets show 𝑛 ∈ {50, 100, 200, 400}. Curves com-
    • pare 𝑇′
    • 𝑛 (perm), HSIC (perm), dCor (perm), and MANOVA (perm). 92
    • x
    • 4.1 Left (iris): 𝑋 = Petal.Length, 𝑍 = Sepal.Length, 𝑌 = Species
    • (color). Petal length separates species strongly even after account-
    • ing for sepal length. Right (Boston): 𝑋 = crim, 𝑍 = lstat,
    • 𝑌 = 1{medv > median} (color). The marginal association between
    • crim and the high–value label largely disappears once lstat is
    • taken into account. . . . . . . . . . . . . . . . . . . . . . . . . . 100
    • 4.2 Empirical bias of the numerator under 𝑋 ⊥ 𝑌
    • 𝑍. Top: when 𝑌
    • depends on 𝑍, the faster convergence rate of 𝑍–nearest neighbors
    • makes 𝐵𝑛 systematically larger than 𝐴𝑛, leading to a negative
    • numerator. Bottom: when 𝑌 ⊥ 𝑍, both terms coincide and the bias
    • disappears. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 167
    • 4.3 Monte Carlo means of √𝑛 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍) under dependence (left)
    • and independence (right). In the dependence case, √𝑛 𝑇′
    • 𝑛 diverges
    • negatively as predicted by the order analysis. In the independence
    • case, √𝑛 𝑇′
    • 𝑛 remains centered near zero, consistent with the absence
    • of bias. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 168
    • 4.4 Histogram of √𝑛 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍) at 𝑛 = 800. Left: dependence case,
    • mean ≈ −3.79, sd ≈ 0.75, Shapiro 𝑝 = 0.809. Right: independence
    • case, mean ≈ −0.03, sd ≈ 0.74, Shapiro 𝑝 = 0.136. Both exhibit
    • approximate normality despite different mean behavior. . . . . . . 168
    • 4.5 Empirical mean of √𝑛 𝑇′
    • 𝑛 (𝑋, 𝑌
    • 𝑍) under 𝑋 ⊥ 𝑌
    • 𝑍 versus 𝑛
    • (log2 scale), with and without variance scaling. . . . . . . . . . . 169
    • xi
    • 4.6 Empirical distribution of √𝑛𝑇′
    • 𝑛 across 𝑛 ∈ {100, 200, 400, 800}
    • with overlaid 𝑁 (0, b𝜎2) curves: empirical (red), pattern-based (blue),
    • and direct (green). . . . . . . . . . . . . . . . . . . . . . . . . . . 169
    • 4.7 Illustration of the data generation. Left: strong signal 𝛾 = 4; Right:
    • weaker signal 𝛾 = 0.4. Points show (𝑍𝑖 , 𝑋𝑖 ) colored by class
    • 𝑌𝑖 ∈ {1, 2, 3}. Dashed vertical lines mark the 𝑍-bins. . . . . . . . 170
    • 4.8 Wage data: 𝑋 vs. 𝑍 colored by 𝑌 . . . . . . . . . . . . . . . . . . . 170
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼