RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Scalable and Efficient Multiple Imputation for Case-Cohort Studies via Influence Function-Based Supersampling = 영향함수 기반 초과표본추출을 활용한 사례-코호트 연구의 효율적 다중대체법

    한글로보기

    https://www.riss.kr/link?id=T17450908

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    높은 비용이 드는 바이오마커를 측정해야 하는 역학 연구에서는 비용 절감을 위해 2단계 표본설계가 널리 활용된다. 이러한 설계 하에서 연구자들은 흔히 콕스 비례위험모형을 이용해 생존결과와 위험요인의 관련성을 추정한다. 1단계에서 이미 확보된 저비용 공변량 정보를 최대한 활용하기 위해, 결측된 바이오마커를 다중대치법으로 보완하는 방법들이 꾸준히 개발되어 왔다. 그러나 코호트 규모가 커지고, 분석모형에 비선형 항이나 상호작용 항이 포함되어 수용–거절 단계가 필요한 경우, 계산 비용이 급격히 증가한다. 최근 Borgan et al. (2023)은 코호트 구성원 중 무작위로 선택한 일부 구성원에 대해서만 다중대치를 수행하는 무작위 초과표본추출 방법을 제안하여 계산량을 줄였으나, 추정 효율성이 감소한다는 한계가 있다. 이에 본 연구에서는 정보가 많은 대상자를 우선적으로 초과표본에 포함시키는 영향함수 기반 초과표본추출 방법과 가중치 보정 절차를 함께 제시한다. 제안 방법은 소규모 초과표본만으로도 결측을 모두 대치한 경우에 근접한 수준의 효율성을 달성하면서, 계산 부담을 크게 줄일 수 있다. 또한, 결측이 있는 변수가 고차원인 경우, 제안 방법의 이점이 더욱 두드러짐을 보인다. 다양한 모의실험을 통해 방법의 성능을 평가하고, NIH–AARP Diet and Health Study 자료에 적용하여 실제 데이터에서도 제안 방법의 실용성을 보인다.
    번역하기

    높은 비용이 드는 바이오마커를 측정해야 하는 역학 연구에서는 비용 절감을 위해 2단계 표본설계가 널리 활용된다. 이러한 설계 하에서 연구자들은 흔히 콕스 비례위험모형을 이용해 생존...

    높은 비용이 드는 바이오마커를 측정해야 하는 역학 연구에서는 비용 절감을 위해 2단계 표본설계가 널리 활용된다. 이러한 설계 하에서 연구자들은 흔히 콕스 비례위험모형을 이용해 생존결과와 위험요인의 관련성을 추정한다. 1단계에서 이미 확보된 저비용 공변량 정보를 최대한 활용하기 위해, 결측된 바이오마커를 다중대치법으로 보완하는 방법들이 꾸준히 개발되어 왔다. 그러나 코호트 규모가 커지고, 분석모형에 비선형 항이나 상호작용 항이 포함되어 수용–거절 단계가 필요한 경우, 계산 비용이 급격히 증가한다. 최근 Borgan et al. (2023)은 코호트 구성원 중 무작위로 선택한 일부 구성원에 대해서만 다중대치를 수행하는 무작위 초과표본추출 방법을 제안하여 계산량을 줄였으나, 추정 효율성이 감소한다는 한계가 있다. 이에 본 연구에서는 정보가 많은 대상자를 우선적으로 초과표본에 포함시키는 영향함수 기반 초과표본추출 방법과 가중치 보정 절차를 함께 제시한다. 제안 방법은 소규모 초과표본만으로도 결측을 모두 대치한 경우에 근접한 수준의 효율성을 달성하면서, 계산 부담을 크게 줄일 수 있다. 또한, 결측이 있는 변수가 고차원인 경우, 제안 방법의 이점이 더욱 두드러짐을 보인다. 다양한 모의실험을 통해 방법의 성능을 평가하고, NIH–AARP Diet and Health Study 자료에 적용하여 실제 데이터에서도 제안 방법의 실용성을 보인다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Two-phase sampling designs have been widely adopted in epidemiological studies to reduce costs when measuring certain biomarkers is prohibitively expensive. Under these designs, investigators commonly relate survival outcomes to risk factors using the Cox proportional hazards model. To fully utilize covariates collected in phase 1, multiple imputation (MI) methods have been developed to impute missing covariates for individuals not included in the phase 2 sample. However, MI becomes computationally intensive in large-scale cohorts, particularly when imputation relies on acceptance–rejection steps to handle nonlinear or interaction terms in the analysis model. To address this issue, Borgan et al. (2023) proposed a random supersampling (RSS) approach that randomly selects a subset of cohort members for imputation, albeit at the cost of reduced efficiency. In this study, we propose an influence-based supersampling (ISS) method with weight calibration. The method achieves efficiency comparable to imputing the entire cohort, even with a small supersample, while substantially reducing computational burden. We further demonstrate that the proposed method is advantageous when estimating hazard ratios for high-dimensional expensive biomarkers. Extensive simulation studies are conducted, and a real data application is provided using the National Institutes of Health--American Association of Retired Persons (NIH--AARP) Diet and Health Study.
    번역하기

    Two-phase sampling designs have been widely adopted in epidemiological studies to reduce costs when measuring certain biomarkers is prohibitively expensive. Under these designs, investigators commonly relate survival outcomes to risk factors using the...

    Two-phase sampling designs have been widely adopted in epidemiological studies to reduce costs when measuring certain biomarkers is prohibitively expensive. Under these designs, investigators commonly relate survival outcomes to risk factors using the Cox proportional hazards model. To fully utilize covariates collected in phase 1, multiple imputation (MI) methods have been developed to impute missing covariates for individuals not included in the phase 2 sample. However, MI becomes computationally intensive in large-scale cohorts, particularly when imputation relies on acceptance–rejection steps to handle nonlinear or interaction terms in the analysis model. To address this issue, Borgan et al. (2023) proposed a random supersampling (RSS) approach that randomly selects a subset of cohort members for imputation, albeit at the cost of reduced efficiency. In this study, we propose an influence-based supersampling (ISS) method with weight calibration. The method achieves efficiency comparable to imputing the entire cohort, even with a small supersample, while substantially reducing computational burden. We further demonstrate that the proposed method is advantageous when estimating hazard ratios for high-dimensional expensive biomarkers. Extensive simulation studies are conducted, and a real data application is provided using the National Institutes of Health--American Association of Retired Persons (NIH--AARP) Diet and Health Study.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Tables iv
    • List of Figures v
    • 1. Introduction 1
    • Abstract i
    • Contents ii
    • List of Tables iv
    • List of Figures v
    • 1. Introduction 1
    • 2. Background 6
    • 2.1 Notation 6
    • 2.2 Case-Cohort Studies 8
    • 2.3 Multiple Imputation for Case-Cohort Studies 10
    • 2.4 Randomly Supersampled Case-Cohort Studies 12
    • 3. Methodology 15
    • 3.1 Influence-Based Supersampling 15
    • 3.2 Influence Function of Log-Hazard Ratio Estimator 15
    • 3.3 Probability Proportional to Size (PPS) Sampling 16
    • 3.4 Balanced Sampling 18
    • 3.5 Influence-Based Supersampled Case-Cohort Analysis 19
    • 3.5.1 Weight Calibration 19
    • 3.5.2 Pseudo-Likelihood Estimation and Inference 22
    • 3.6 Remarks on Missing Scenarios 24
    • 3.6.1 Scenario 1: Single Expensive Covariate 24
    • 3.6.2 Scenario 2: Multi-Dimensional Expensive Covariates 25
    • 4 Simulation 27
    • 4.1 Setup 27
    • 4.2 Simulation Results 29
    • 5 Case Study 36
    • 6 Discussion 39
    • A Appendix 49
    • A.1 Stratified Case-Cohort Simulation Studies 49
    • A.2 Case Study 49
    • Abstract (In Korean) 52
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼