인간의 표현형 및 질병에 대한 유전적 결정 요인을 분석하는 연구는 대규모 바이오뱅크와 전장 유전체 연관 분석(GWAS)의 발전으로 급격히 가속화되었다. 특히 UK Biobank 및 FinnGen과 같은 바이...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
인간의 표현형 및 질병에 대한 유전적 결정 요인을 분석하는 연구는 대규모 바이오뱅크와 전장 유전체 연관 분석(GWAS)의 발전으로 급격히 가속화되었다. 특히 UK Biobank 및 FinnGen과 같은 바이...
인간의 표현형 및 질병에 대한 유전적 결정 요인을 분석하는 연구는 대규모 바이오뱅크와 전장 유전체 연관 분석(GWAS)의 발전으로 급격히 가속화되었다. 특히 UK Biobank 및 FinnGen과 같은 바이오뱅크는 다양한 인구집단을 대상으로 유전적 구조를 탐색하고 질병 위험을 예측할 수 있는 전례 없는 기회를 제공하였다. 그럼에도 불구하고, 대규모 바이오뱅크 데이터 분석의 광범위한 적용을 위해 여전히 해결해야 할 과제들이 존재하며, 대표적으로는 유럽 중심으로 구축된 GWAS 결과의 인종 간 일반화 문제, 흔한 변이에 의존하는 다유전자 위험 점수(PRS)의 성능 저하, 그리고 특히 단백질체(proteomics)를 포함한 다중 오믹스 정보의 제한적인 활용 등이 있다.
대부분의 GWAS는 유럽인을 중심으로 진행돼 왔기 때문에, 해당 결과를 다른 인종에 적용할 경우 인종에 의한 차이가 큰 표현형에 대해서는 예측 정확도가 크게 떨어지는 문제가 있다. 본 논문의 제2장에서는 이러한 한계를 극복하기 위해 한국인 바이오뱅크 데이터를 활용해 76개 표현형에 대한 GWAS를 수행했다. 그 결과, 총 2,242개의 유의한 유전자 위치(genomic loci)를 확인했고, 이 중 122개는 새롭게 발견된 연관성이었다. 또한 KoGES와 Biobank Japan 데이터를 통합한 메타 분석을 통해 추가로 379개의 새로운 유전자 위치를 찾아냈다. 이러한 분석은 동아시아인 특화 예측 모형의 성능을 크게 향상시킬 수 있음을 확인하였다.
다유전자 위험 점수는 복잡한 표현형과 질병에 대한 유전적 감수성(genetic susceptibility)을 예측하는 데 널리 사용되고 있다. 하지만 대부분 흔한 변이에만 기반하고 있어서, 희귀 변이의 효과는 제대로 반영되지 않는 한계가 있다. 제3장에서는 이러한 문제를 해결하기 위해 RareEffect라는 새로운 분석 기법을 제안한다. RareEffect는 분산 성분 모형과 경험적 베이지안 추정을 결합해 희귀 변이의 효과를 유전자 단위 및 개별 변이 단위로 정밀하게 추정할 수 있는 방법이다. 또한 표현형의 불균형에 대응할 수 있도록 Firth 편향 보정을 효율적으로 구현했다. UK Biobank의 전장 엑솜 시퀀싱 데이터를 활용한 시뮬레이션과 실제 바이오뱅크 데이터 분석을 통해 RareEffect는 정확도와 계산 효율 면에서 기존 방법보다 뛰어난 성능을 보였으며, 이로부터 얻은 희귀 변이 효과를 기존 PRS에 통합했을 때 예측력이 유의하게 향상됨을 확인하였다.
또한, 유전체 기반 분석을 넘어, 단백질체를 포함한 다중 오믹스 데이터의 통합은 정밀의료의 핵심 도구로 주목받고 있다. 제4장에서는 UK Biobank 데이터를 기반으로 희귀 변이와 일반 변이를 통합한 단백질 예측 모형을 구축하였다. 대다수 단백질에 대해서는 유전체 기반 예측력이 높지 않았지만, 희귀 변이를 포함함으로써 일부 단백질에서는 예측 정확도가 유의미하게 향상되었다. 이 모형을 활용한 단백질-표현형 연관성 분석(PWAS)은 linkage disequilibrium이나 유전자 간 물리적 거리로 인해 발생하는 1종 오류(type 1 error)를 줄이는 데 도움이 될 수 있다. 다만, 유전적으로 예측한 단백질 정보를 질병 예측에 활용할 경우 정확도는 다소 제한적이었으며, 실제 측정된 단백질 수치를 사용할 때 예측력이 크게 향상되는 것으로 나타났다. 이는 직접 측정된 단백질 정보의 임상적 가치를 강조하는 동시에, 유전 기반 단백질 예측 모형의 정밀도를 높이는 추가적인 연구가 필요함을 시사한다.
다국어 초록 (Multilingual Abstract)
Understanding genetic determinants of human traits and diseases has profoundly accelerated with the advent of genome-wide association studies (GWAS) and large-scale biobank resources. Biobanks such as the UK Biobank and FinnGen offer unprecedented opp...
Understanding genetic determinants of human traits and diseases has profoundly accelerated with the advent of genome-wide association studies (GWAS) and large-scale biobank resources. Biobanks such as the UK Biobank and FinnGen offer unprecedented opportunities to explore genetic architecture and predict disease risk across diverse populations. Nevertheless, key challenges remain, notably concerning the generalizability of genetic findings across ancestries, limitations of polygenic risk scores (PRS) that predominantly rely on common genetic variants, and the limited incorporation of multi-omics information, particularly proteomics.
European ancestry dominates most GWAS datasets, leading to a substantial performance drop when applying predictive models derived from European populations to non-European populations. In Chapter 2 of this dissertation, we address this critical gap by conducting comprehensive GWAS for 76 phenotypes using Korean biobank data (KoGES, n=72,298). Our analyses identified 2,242 significant genomic loci, including 122 novel associations. To enhance power, we employed state-of-the-art genetic association methodologies and demonstrated improved discovery and replication by integrating KoGES and Biobank Japan data via meta-analysis. This approach yielded an additional 379 novel loci, underscoring the benefits of ancestry-specific studies and demonstrating a significant enhancement in East Asian-specific predictive models. This work provides publicly available summary statistics to support further research on East Asian genetic architecture.
Polygenic risk scores have become integral to predicting genetic susceptibility to complex traits and diseases. However, their utility is constrained by their heavy reliance on common variants, typically neglecting rare variants due to methodological complexities. In Chapter 3, we address these limitations by introducing RareEffect, an innovative approach utilizing a variance component model integrated with empirical Bayesian estimation. RareEffect enables precise estimation of rare variant effects at both gene-region and individual variant levels, significantly enhancing interpretability. Moreover, this approach accounts for phenotype imbalance via a computationally efficient implementation of the Firth bias correction. Extensive simulation studies and analyses of 100 traits in the UK Biobank's whole-exome sequencing dataset validated RareEffect's accuracy and efficiency. Crucially, integrating rare variant effect estimates from RareEffect into traditional PRS substantially improved predictive performance, surpassing existing methods focused solely on rare variants.
Beyond genomics alone, integrating multi-omics data, particularly proteomics, presents a promising direction in precision medicine, providing deeper insights into biological mechanisms and potential therapeutic targets. Chapter 4 expands upon genome-based approaches by integrating rare and common genetic variants to develop genetically informed protein prediction models using UK Biobank data. Although genetic predictability for many proteins was modest, incorporating rare variants notably enhanced predictive accuracy for certain proteins. By incorporating rare variant information, proteome-wide association studies (PWAS) may help mitigate false-positive associations arising from local linkage disequilibrium and physical proximity of protein-coding genes, thereby facilitating the identification of novel protein-phenotype associations. However, despite methodological advancements, genetically predicted proteins demonstrated limited accuracy in direct disease risk prediction tasks. Conversely, using measured protein levels markedly enhanced predictive accuracy, emphasizing the inherent clinical value of direct proteomic measurements. At the same time, these findings highlight the need for improving the precision of genetically predicted protein models, which could enhance their utility in settings where direct measurement is unavailable. This work thus elucidates both the potential and limitations of genetically informed proteomics, pointing to promising directions for future improvements in protein prediction accuracy and broader multi-omics integration.
This dissertation advances methodologies for integrating diverse genetic data to improve disease prediction and genetic association studies. By addressing ancestry-specific generalizability, incorporating rare variants, and extending analyses to proteomics, it offers practical contributions toward precision medicine using large-scale biobank resources.
목차 (Table of Contents)