RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Conformalized Method for Empirical Bayes Normal Mean Inference Problem with Heteroscedastic Variance = 이분산성을 고려한 경험적 베이즈 정규 평균 추론을 위한 컨포멀 방법

    한글로보기

    https://www.riss.kr/link?id=T17314531

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문에서는 여러 정규분포의 평균에 대한 동시 가설 검정을 수행하는 정규평균 추론(normal mean inference) 문제를 다룬다. 이 문제는 경험적 베이즈(empirical Bayes) 방법론에서 오랫동안 연구되어 왔으나, 기존 방법들의 성능은 주로 (i) 사전분포의 정확한 설정과 (ii) 그 모수의 정밀한 추정에 크게 의존한다. 그러나 실제 데이터 분석에서는 이러한 조건을 충족하는 것이 어려울 뿐만 아니라, 해당 가정의 타당성을 검증하는 것 또한 쉽지 않다. 이러한 한계를 극복하기 위해, 본 연구에서는 사전분포의 정확한 설정이나 정밀한 추정을 요구하지 않는 새로운 알고리즘인 COIN (COnformal Inference for Normal mean problem)을 제안한다. 본 연구에서는 COIN 알고리즘으로부터 도출된 의사결정 규칙이 사전분포가 정확히 설정되거나 정밀하게 추정되지 않더라도 목표 수준의 허위발견률(false discovery rate, FDR)을 점근적으로 제어함을 이론적으로 증명하였다. 한편, COIN 알고리즘은 사전 분포와 정합도 점수 함수(conformity score function)의 추정을 위해 외부 훈련 데이터가 필요하기 때문에, 외부 훈련 데이터가 없는 경우에도 적용할 수 있도록 샘플 분할(sample-splitting) 및 특성 분할(feature-splitting) 기반의 두 가지 데이터 분할 방법을 추가로 제안하였다. 제안한 방법의 이론적 타당성과 수치적 성능은 다양한 시뮬레이션 연구를 통해 검증하였으며, 세 가지 실제 데이터 예제에 적용하여 실용성을 확인하였다.
    번역하기

    본 논문에서는 여러 정규분포의 평균에 대한 동시 가설 검정을 수행하는 정규평균 추론(normal mean inference) 문제를 다룬다. 이 문제는 경험적 베이즈(empirical Bayes) 방법론에서 오랫동안 연구되...

    본 논문에서는 여러 정규분포의 평균에 대한 동시 가설 검정을 수행하는 정규평균 추론(normal mean inference) 문제를 다룬다. 이 문제는 경험적 베이즈(empirical Bayes) 방법론에서 오랫동안 연구되어 왔으나, 기존 방법들의 성능은 주로 (i) 사전분포의 정확한 설정과 (ii) 그 모수의 정밀한 추정에 크게 의존한다. 그러나 실제 데이터 분석에서는 이러한 조건을 충족하는 것이 어려울 뿐만 아니라, 해당 가정의 타당성을 검증하는 것 또한 쉽지 않다. 이러한 한계를 극복하기 위해, 본 연구에서는 사전분포의 정확한 설정이나 정밀한 추정을 요구하지 않는 새로운 알고리즘인 COIN (COnformal Inference for Normal mean problem)을 제안한다. 본 연구에서는 COIN 알고리즘으로부터 도출된 의사결정 규칙이 사전분포가 정확히 설정되거나 정밀하게 추정되지 않더라도 목표 수준의 허위발견률(false discovery rate, FDR)을 점근적으로 제어함을 이론적으로 증명하였다. 한편, COIN 알고리즘은 사전 분포와 정합도 점수 함수(conformity score function)의 추정을 위해 외부 훈련 데이터가 필요하기 때문에, 외부 훈련 데이터가 없는 경우에도 적용할 수 있도록 샘플 분할(sample-splitting) 및 특성 분할(feature-splitting) 기반의 두 가지 데이터 분할 방법을 추가로 제안하였다. 제안한 방법의 이론적 타당성과 수치적 성능은 다양한 시뮬레이션 연구를 통해 검증하였으며, 세 가지 실제 데이터 예제에 적용하여 실용성을 확인하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods heavily depends on two key conditions: (i) the prior distribution is correctly specified, and (ii) it can be accurately estimated. In practice, both conditions are difficult to satisfy, and it is often unclear whether they hold in a given application. To overcome these limitations, we propose a new algorithm, called COIN (COnformal Inference for Normal mean inference problem). Unlike traditional empirical Bayes approaches, COIN produces decision rules whose validity does not depend on the correct specification or accurate estimation of the prior. We theoretically prove that COIN asymptotically controls the false discovery rate at the nominal level, even in the presence of prior misspecification or estimation errors. Since the COIN algorithm requires an external training dataset to estimate the prior distribution and conformity score function, we introduce two data-splitting strategies---sample-splitting and feature-splitting---for the case where such external data are unavailable. We provide theoretical guarantees for the data-splitting strategies and demonstrate their effectiveness through extensive numerical studies. Finally, we illustrate the practical utility of COIN through applications to three real data examples.
    번역하기

    We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods...

    We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods heavily depends on two key conditions: (i) the prior distribution is correctly specified, and (ii) it can be accurately estimated. In practice, both conditions are difficult to satisfy, and it is often unclear whether they hold in a given application. To overcome these limitations, we propose a new algorithm, called COIN (COnformal Inference for Normal mean inference problem). Unlike traditional empirical Bayes approaches, COIN produces decision rules whose validity does not depend on the correct specification or accurate estimation of the prior. We theoretically prove that COIN asymptotically controls the false discovery rate at the nominal level, even in the presence of prior misspecification or estimation errors. Since the COIN algorithm requires an external training dataset to estimate the prior distribution and conformity score function, we introduce two data-splitting strategies---sample-splitting and feature-splitting---for the case where such external data are unavailable. We provide theoretical guarantees for the data-splitting strategies and demonstrate their effectiveness through extensive numerical studies. Finally, we illustrate the practical utility of COIN through applications to three real data examples.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Tables v
    • List of Figures vi
    • Abstract i
    • Contents ii
    • List of Tables v
    • List of Figures vi
    • 1 Introduction 1
    • 1.1 Normal Mean Inference Problem 1
    • 1.2 Existing EBNMI Methods 2
    •  1.2.1 General Procedure 2
    •  1.2.2 An Illustrative Example: Zheng et al. (2021) 4
    • 1.3 Limitations of the Existing EBNMI Methods 5
    •  1.3.1 An Illustrative Numerical Example 6
    • 1.4 Our Contributions 9
    • 2 Problem Settings 12
    • 3 Methodology 15
    • 3.1 Construction of Calibration Variables 15
    •  3.1.1 Conditional Exchangeability 15
    •  3.1.2 Oracle Construction 16
    •  3.1.3 Data-Adaptive Construction 16
    • 3.2 COIN Algorithm 18
    • 3.3 Theory 20
    •  3.3.1 Assumption and Lemmas 20
    •  3.3.2 Main Results 27
    • 4 Data Splitting Methods 36
    • 4.1 Two Types of Test Dataset 36
    • 4.2 Sample-Splitting Method 37
    • 4.3 Feature-Splitting Method 39
    •  4.3.1 Theory 42
    •  4.3.2 A Uniform Improvement on Data-Adaptive Threshold 47
    • 5 Numerical Study 49
    • 5.1 Simulated Data Generation 49
    •  5.1.1 Construction of Individual-level Test Dataset, 𝒟ʳᵃʷ 50
    •  5.1.2 Derivation of the Summary-Level Test Dataset, 𝒟 50
    • 5.2 Performance Metric 51
    • 5.3 The Methods Considered in the Comparison 52
    • 5.4 Implementation Details of Data-Splitting Methods 53
    •  5.4.1 Working Prior 54
    •  5.4.2 Conformity Score Function 54
    •  5.4.3 Estimation Algorithm 54
    •  5.4.4 Additional Implementation Details for Conf-FS 57
    • 5.5 Simulation Setup 57
    • 5.6 Results 59
    •  5.6.1 Simulation Results under Scenario 1: σ²ᵢ ⫫ μᵢ 59
    •  5.6.2 Simulation Results under Scenario 2: σ²ᵢ ⫫ 𝟙(μᵢ ≠ 0) 63
    •  5.6.3 Simulation Results under Scenario 3: Dependent σ²ᵢ and 𝟙(μᵢ ≠ 0) 65
    • 5.7 Choice of Fold-specific Target FDR Levels α⁽ᵏ⁾ for Conf-FS 68
    • 6 Real Data Analysis 76
    • 6.1 Description of Datasets 76
    •  6.1.1 Example 1: DNA Methylation Dataset (Zhang et al., 2013) 76
    •  6.1.2 Example 2: CLL Dataset (Dietrich et al., 2018) 77
    •  6.1.3 Example 3: Breast Cancer Proteomics Dataset (Terkelsen et al., 2021) 78
    • 6.2 Results 79
    • 7 Discussion 81
    • Appendix 87
    • Abstract (In Korean) 94
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼