본 논문에서는 여러 정규분포의 평균에 대한 동시 가설 검정을 수행하는 정규평균 추론(normal mean inference) 문제를 다룬다. 이 문제는 경험적 베이즈(empirical Bayes) 방법론에서 오랫동안 연구되...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 논문에서는 여러 정규분포의 평균에 대한 동시 가설 검정을 수행하는 정규평균 추론(normal mean inference) 문제를 다룬다. 이 문제는 경험적 베이즈(empirical Bayes) 방법론에서 오랫동안 연구되...
본 논문에서는 여러 정규분포의 평균에 대한 동시 가설 검정을 수행하는 정규평균 추론(normal mean inference) 문제를 다룬다. 이 문제는 경험적 베이즈(empirical Bayes) 방법론에서 오랫동안 연구되어 왔으나, 기존 방법들의 성능은 주로 (i) 사전분포의 정확한 설정과 (ii) 그 모수의 정밀한 추정에 크게 의존한다. 그러나 실제 데이터 분석에서는 이러한 조건을 충족하는 것이 어려울 뿐만 아니라, 해당 가정의 타당성을 검증하는 것 또한 쉽지 않다. 이러한 한계를 극복하기 위해, 본 연구에서는 사전분포의 정확한 설정이나 정밀한 추정을 요구하지 않는 새로운 알고리즘인 COIN (COnformal Inference for Normal mean problem)을 제안한다. 본 연구에서는 COIN 알고리즘으로부터 도출된 의사결정 규칙이 사전분포가 정확히 설정되거나 정밀하게 추정되지 않더라도 목표 수준의 허위발견률(false discovery rate, FDR)을 점근적으로 제어함을 이론적으로 증명하였다. 한편, COIN 알고리즘은 사전 분포와 정합도 점수 함수(conformity score function)의 추정을 위해 외부 훈련 데이터가 필요하기 때문에, 외부 훈련 데이터가 없는 경우에도 적용할 수 있도록 샘플 분할(sample-splitting) 및 특성 분할(feature-splitting) 기반의 두 가지 데이터 분할 방법을 추가로 제안하였다. 제안한 방법의 이론적 타당성과 수치적 성능은 다양한 시뮬레이션 연구를 통해 검증하였으며, 세 가지 실제 데이터 예제에 적용하여 실용성을 확인하였다.
다국어 초록 (Multilingual Abstract)
We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods...
We study the normal mean inference problem, which involves simultaneous testing of the means of many normal distributions. This problem has been extensively studied within the empirical Bayes (EB) framework. However, the reliability of most EB methods heavily depends on two key conditions: (i) the prior distribution is correctly specified, and (ii) it can be accurately estimated. In practice, both conditions are difficult to satisfy, and it is often unclear whether they hold in a given application. To overcome these limitations, we propose a new algorithm, called COIN (COnformal Inference for Normal mean inference problem). Unlike traditional empirical Bayes approaches, COIN produces decision rules whose validity does not depend on the correct specification or accurate estimation of the prior. We theoretically prove that COIN asymptotically controls the false discovery rate at the nominal level, even in the presence of prior misspecification or estimation errors. Since the COIN algorithm requires an external training dataset to estimate the prior distribution and conformity score function, we introduce two data-splitting strategies---sample-splitting and feature-splitting---for the case where such external data are unavailable. We provide theoretical guarantees for the data-splitting strategies and demonstrate their effectiveness through extensive numerical studies. Finally, we illustrate the practical utility of COIN through applications to three real data examples.
목차 (Table of Contents)