배경: 인공지능 (AI) 은 현대 과학적 탐구와 임상 진료의 전 영역에 걸쳐 패러다임 변화를 이끌고 있으며, 그중에서도 AI 기반 의료는 가장 혁신적이고 빠르게 발전하는 분야로 부상하고 있다....

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
배경: 인공지능 (AI) 은 현대 과학적 탐구와 임상 진료의 전 영역에 걸쳐 패러다임 변화를 이끌고 있으며, 그중에서도 AI 기반 의료는 가장 혁신적이고 빠르게 발전하는 분야로 부상하고 있다....
배경: 인공지능 (AI) 은 현대 과학적 탐구와 임상 진료의 전 영역에 걸쳐 패러다임 변화를 이끌고 있으며, 그중에서도 AI 기반 의료는 가장 혁신적이고 빠르게 발전하는 분야로 부상하고 있다. 이러한 흐름 속에서, 일상적으로 획득되는 임상 영상을 활용하여 추가적이고 기존에는 인지되지 않았던 건강 위험을 식별하는 기회 검진 (opportunistic screening) 은 1차의료 및 예방의학 측면에서 높은 확장성과 비용 효율성을 갖춘 전략으로 각광받고 있다. 대규모 비지도 데이터로 사전학습된 트랜스포머 기반의 거대 모델인 파운데이션 모델 (FM) 은 범용적이고 강력한 표현 학습 능력을 제공하며 다양한 의료 영상 분석 작업에 효율적으로 적응될 수 있다. 그러나 이러한 잠재력에도 불구하고 FM의 기회 검진 활용은 아직 제한적이며, 특히 해석 가능성과 투명성의 부족은 실제 임상 구현에 중요한 장애 요인으로 남아 있다.
목적: 본 학위논문은 FM 기반 기회 검진 모델을 임상적으로 신뢰할 수 있는 수준으로 개발·평가하기 위해, 진단 성능·설명 가능성·임상적 유용성을 체계적으로 평가하는 다차원적 평가 프레임워크 (multidimensional evaluation framework) 를 구축하고자 한다.
방법: 제안한 프레임워크의 강건성과 다양한 질환·영상 도메인 간 일반화 가능성을 평가하기 위해, 서울대학교병원 건강증진센터에서 수집한 네 가지 대표적 기회 검진 사례에 적용하였다: 흉부 X선 기반 골다공증, 안저영상 및 흉부 X선 기반 경동맥죽상경화증, 안저영상 기반 빈혈 (헤모글로빈 농도).
본 연구에서는 마스크드 이미지 모델링 (MIM), 대조학습, 텍스트-약지도 학습 등 다양한 학습 방식과 자연·의료 도메인을 아우르는 FM (OpenCLIP, DINOv2, MAE, CheXagent, RAD-DINO, RETFound, VisionFM) 을 사용하였다. 각 모델은 선형 프로빙, 부분 미세조정, LoRA 기반 저랭크 적응 등 파라미터 효율적 미세조정 (PEFT) 기법을 통해 특정 검진 작업에 최적화하였다. 모델 평가는 1.예측 성능, 2.설명 가능성 (해부학적·생리학적 구조와의 정렬 정도를 정량화한 saliency·perturbation 기반 지표), 3.임상적 유용성 (딥러닝 바이오마커 기반 Cox PH 생존 분석) 을 중심으로 이루어졌다.
결과: 흉부 X선 기반 골다공증 분류에서는 DINOv2–LoRA 모델이 최고 성능 (AUC 0.93) 을 보였으며, 척추·늑골 등 골다공증과 직접적으로 연관된 구조에 집중하는 명확한 해석 가능성을 나타냈다. 안저영상 기반 죽상경화증 예측에서는 DINOv2–LoRA가 가장 우수한 분류 성능(AUC 0.71)을 보였을 뿐 아니라, 연령·성별 보정 후에도 유의한 예후 예측력 (HR 1.37)을 유지하였고, 모델 주의 (attention) 가 동·정맥 분포와 강하게 정렬되어 C-index 0.79를 달성하였다. 흉부 X선 기반 죽상경화증 예측에서는 RAD-DINO–LoRA가 예측 성능과 설명 가능성의 균형 측면에서 최적 모델로 선정되었으며 (AUC 0.74), 모델 출력 (DL-CXAS) 삼분위수 간 약 11배의 심혈관 사망 위험도 차이가 나타났고 이는 프레이밍햄 위험점수 (FRS) 보정 후에도 유지되었다 (C-index 0.71). 안저영상 기반 빈혈 예측에서는 DINOv2–LoRA가 FM 중 최고 성능 (MAE 0.70; AUC 0.92)을 보였으며, 시신경유두 주변 및 황반부 등 생리적으로 타당한 영역에 일관된 saliency가 관찰되었다. 또한 동맥의 밝기·채도 증가가 예측 헤모글로빈 농도의 증가와 비례 관계를 보여, 혈색소의 광흡수 특성과 부합하였다.
결론: 본 박사학위 논문은 예측 성능과 설명가능성이 본질적으로 정렬되어 있지 않음을 보여준다. 즉, 유사한 정확도를 보이는 모델이라 하더라도 서로 다른--때로는 임상적으로 타당하지 않은--단서에 의존할 수 있으며, 이는 임상적 신뢰성을 확보하기 위해 다차원적 평가가 필수적임을 시사한다. 본 논문에 포함된 여러 연구 전반에서 DINOv2는 일관되게 최상위 수준의 성능을 달성하였으며, 이는 견고하고 미세한 시각적 표현을 학습할 수 있도록 설계된 자기지도학습 기반의 사전학습 전략에 기인한 것으로 해석된다. 이러한 결과는 자연 영상 도메인에서 사전학습된 파운데이션 모델이 의료 응용 분야에서도 강력하고 지속적으로 확장 가능한 잠재력을 지니고 있음을 보여준다. 또한 재파라미터화 기반의 매개변수 효율적 미세조정 기법인 LoRA는 성능과 설명가능성을 동시에 향상시키는 데 효과적이었으며, 모듈화 가능성이라는 특성으로 인해 향후 통합 임상 파운데이션 모델 생태계의 핵심 구성 요소로 활용될 수 있는 잠재력을 지닌다. 본 연구는 몇 가지 한계점을 갖는다: 첫째, 설명가능성을 기존 임상 지식과의 일치로 한정하여 정의하였다는 점이며, 둘째, 고위험군이 상대적으로 많이 포함된 단일 기관 코호트에서만 검증이 수행되었다는 점이다. 향후 연구에서는 설명가능성의 개념을 확장하여 임상적으로 의미 있는 새로운 패턴과 우연적 상관관계 또는 비인과적 신호를 구분할 수 있는 해석 프레임워크를 정립할 필요가 있다. 또한 일반적인 건강검진 인구를 보다 충실히 반영하는 다기관 기회 검진 데이터셋을 활용한 외부 검증이 요구되며, 설명가능성이 다른 데이터 분포 상황에서의 모델 강건성과 어떠한 관계를 갖는지에 대한 체계적인 분석도 필요하다. 추가적으로, 혈관 위험 평가와 같이 구조적·전신적 특성이 동시에 작용하는 질환을 대상으로 종합 건강검진 데이터를 공동으로 모델링할 수 있는 멀티모달 파운데이션 모델의 개발, 그리고 한국 또는 더 나아가 아시아 인구 집단에 특화된 의료 파운데이션 모델 구축의 타당성과 가치에 대한 탐구 역시 중요한 향후 연구 방향이다. 이는 최근 논의되고 있는 ''sovereign AI'' 주도권과도 맞닿아 있으며, 임상적·전략적 측면에서 의미 있는 기여를 할 수 있을 것으로 기대된다.
다국어 초록 (Multilingual Abstract)
Background: Artificial intelligence is reshaping scientific inquiry and clinical practice, with AI-enabled healthcare emerging as one of its most transformative frontiers. Within this landscape, opportunistic screening—leveraging routinely acquired ...
Background: Artificial intelligence is reshaping scientific inquiry and clinical practice, with AI-enabled healthcare emerging as one of its most transformative frontiers. Within this landscape, opportunistic screening—leveraging routinely acquired clinical data to uncover additional, previously unrecognized health risks—offers a scalable and cost-effective strategy with substantial implications for primary care and population-level prevention. Foundation models (FMs), large transformer-based architectures pretrained on massive unlabeled datasets, provide powerful, task-agnostic visual representations that can be efficiently adapted to downstream applications. However, their deployment in opportunistic screening remains limited, constrained by challenges in interpretability and transparency.
Purpose: This dissertation aims to establish a robust and generalizable multidimensional evaluation framework that bridges advances in foundation models with clinical requirements by systematically assessing diagnostic performance, model explainability, and clinical utility for opportunistic screening.
Methods: To examine the robustness and cross-domain generalizability of the proposed framework, I applied it to four representative case studies using data from the Health Promotion Center of Seoul National University Hospital: osteoporosis from chest X-rays, carotid atherosclerosis from retinal fundus images and chest X-rays, and anemia from fundus images. A diverse suite of vision FMs -- spanning masked image modeling, contrastive learning, and text-guided pretraining, and covering natural- and medical-domain models (OpenCLIP, DINOv2, MAE, CheXagent, RAD-DINO, RETFound, VisionFM) -- was adapted using parameter-efficient fine-tuning (PEFT) techniques including linear probing, partial fine-tuning, and low-rank adaptation (LoRA). Models were compared across three axes: predictive performance, explainability (quantified as the alignment between model reasoning and clinically relevant anatomical or physiological structures using saliency- and perturbation-based methods), and clinical utility (deep-learning biomarkers incorporated into Cox proportional hazards models).
Results: For osteoporosis screening from chest X-rays, DINOv2 fine-tuned with LoRA achieved the highest accuracy (AUC 0.93) while localizing attention to the spine, ribs, and other fracture-relevant structures. In fundus-based carotid atherosclerosis prediction, DINOv2 LoRA again performed best (AUC 0.71), generating deep-learning biomarkers with significant age- and sex-adjusted prognostic value (HR 1.37) and strong alignment with retinal vasculature, yielding a C-index of 0.79. For chest X-ray–based atherosclerosis prediction, RAD-DINO LoRA achieved the most favorable balance between performance and explainability (AUC 0.74), with the biomarkers remaining significantly associated with cardiovascular mortality after adjustment for the Framingham Risk Score (C-index 0.71). In anemia prediction from fundus images, DINOv2 LoRA achieved the best overall FM performance (mean absolute error 0.70; AUC 0.92), with saliency and perturbation analyses confirming physiologically plausible reliance on macular, peripapillary, and arterial features.
Conclusion: This dissertation shows that predictive performance and explainability are not intrinsically aligned; models with similar accuracy may rely on distinct and sometimes non-clinical cues, reinforcing the need for multidimensional evaluation to ensure clinical trustworthiness. Across studies, DINOv2 consistently achieved leading performance, which may be attributed to its self-supervised training strategy that produces robust and fine-grained visual representations. These findings illustrate the strong and expanding potential of natural-domain FMs for medical applications. Reparameterized PEFT, specifically LoRA, proved effective in enhancing both performance and interpretability, and its modularity positions it as a compelling component of future unified clinical FM ecosystems. This work has several limitations, including defining explainability solely as alignment with established clinical knowledge and conducting validation within a single-center cohort enriched for high-risk individuals. Future research should broaden the conceptualization of interpretability to distinguish clinically meaningful novel patterns from spurious correlations; pursue external validation using multi-institutional opportunistic screening datasets that better reflect general health-screening populations; and investigate how explainability relates to out-of-distribution robustness. Additional directions include developing multimodal FMs capable of jointly modeling comprehensive health-examination data, particularly for vascular risk assessment, and exploring the feasibility and value of constructing medical FMs tailored to Korean or broader Asian populations, which may contribute to emerging ''sovereign AI'' initiatives.
목차 (Table of Contents)