딥러닝은 의생명 영상의 정량적 분석에서 혁신적인 성과를 이루었으나, 실험실 환경에서의 우수한 성능이 실제 임상 환경에서의 신뢰성으로 이어지지 못하는 근본적인 한계가 여전히 존재...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17450868
서울 : 서울대학교 대학원, 2026
학위논문(박사) -- 서울대학교 대학원 , 융합과학부 방사선융합의생명전공 , 2026. 2
2026
영어
620.5
서울
xviii, 212 ; 26 cm
지도교수: 예성준
I804:11032-000000195244
0
상세조회0
다운로드딥러닝은 의생명 영상의 정량적 분석에서 혁신적인 성과를 이루었으나, 실험실 환경에서의 우수한 성능이 실제 임상 환경에서의 신뢰성으로 이어지지 못하는 근본적인 한계가 여전히 존재...
딥러닝은 의생명 영상의 정량적 분석에서 혁신적인 성과를 이루었으나, 실험실 환경에서의 우수한 성능이 실제 임상 환경에서의 신뢰성으로 이어지지 못하는 근본적인 한계가 여전히 존재한다. 깨끗한 테스트 데이터에서 높은 정확도를 보이는 모델이라 하더라도, 학습 데이터와 상이한 영상 조건이 빈번하게 발생하는 임상 현장에서는 성능이 급격히 저하되며, 이는 안전이 중요한 의료 인공지능 응용에서 심각한 위험 요소로 작용한다. 본 학위논문은 이러한 문제를 해결하기 위해, 임상적으로 신뢰 가능한 인공지능 시스템의 개발에 필수적인 두가지 요소인 강건성(robustness) 평가 프레임워크와 효율적인 불확실성 정량화(uncertainty quantification) 방법을 통합적으로 제안한다. 첫 번째 연구 기여는 의생명 영상 분야를 위한 체계적인 강건성 평가 프레임워크이다. 본 프레임워크는 Optical Diffraction Tomography(ODT)와 Computed Tomography(CT) 응용을 대상으로, 물리적으로 정의된 16가지 영상 왜곡(corruption)을 5단계 심각도 수준으로 구성한 도메인 특화 벤치마크를 제시한다. 더 나아가, 단순한 성능 평가를 넘어 강건성 향상을 위한 데이터 증강 기법으로서 ODT 세균 분류를 위한 CutPix와 CT 장기위험구조(organ-at-risk) 분할을 위한 CutPixMixup을 제안한다. 이 기법들은 프랙탈 패턴 혼합과 장기 인지 기반 패치 연산을 통해 형상(shape) 정보와 질감(texture) 정보의 균형을 명시적으로 학습하도록 설계되었다. 이는 질감 편향 모델이 깨끗한 데이터에서는 높은 정확도를 보이나 왜곡 환경에서는 취약한 반면, 형상 편향 모델은 강건성을 확보하는 대신 정확도를 희생하는 기존 딥러닝의 근본적인 trade-off를 해결하고자 한다. 실험 결과, CutPix는 ODT 분류에서 DenseNet-121 기준 평균 Corrupted Error(mCE) 42.97\%, ResNet-101 기준 56.03\%로 최저 성능 저하를 달성하였으며, CutPixMixup은 CT 분할에서 뇌간(brainstem) 62.25\%, 안구(eye) 69.90\%의 mCE로 기존 방법 대비 최고 수준의 강건성을 달성하였다. 이는 형상과 질감의 균형 잡힌 특징 학습이 높은 정확도와 왜곡 환경에 대한 강한 복원력을 동시에 가능하게 함을 입증한다.
두 번째 연구 기여는 Neural Bootstrapper(NeuBoots)라 명명된 효율적인 불확실성 정량화 기법이다. 기존의 앙상블 또는 베이지안 접근법이 다수의 모델 학습이나 반복적인 추론을 요구하는 반면,
NeuBoots는 고수준 특징 계층에 부트스트랩 가중치를 직접 적용함으로써 통계적 타당성을 유지하면서도 학습 및 추론 비용을 크게 절감한다. 예측 보정(calibration), 능동 학습(active learning), 분포 외 탐지(out-of-distribution detection), 의미론적 분할(semantic segmentation) 등 다양한 과제에서의 실험을 통해, NeuBoots는 표준 부트스트래핑과 유사한 95\% 명목 수준에 근접한 안정적인 빈도주의적 커버리지(frequentist coverage)를 달성함과 동시에, 능동 학습 및 OOD 탐지에서는 기존 기준 방법들을 유의미하게 상회하는 성능을 보였다.
본 논문의 핵심 주장은 강건성 평가와 불확실성 정량화가 상호 독립적인 문제가 아니라 상호 의존적인 필수 요소라는 점이다. 불확실성 정량화 없이 강건성만을 고려할 경우, 모델은 학습 분포를 벗어난 상황에서 실패를 인지하지 못한 채 침묵 오류(silent failure)를 일으킬 수 있으며, 반대로 강건성 없이 불확실성만을 고려할 경우 분포 변화 하에서 신뢰도 추정이 심각하게 왜곡된다. 본 논문에서 제안하는 두가지 기여는 이러한 문제를 통합적으로 해결함으로써, FDA 규제를 받는 의료기기 소프트웨어(AI Software as a Medical Device)에 적합한 인공지능 시스템 개발을 위한 핵심 도구를 제공한다. 제안된 방법들은 임상 검증 이전 단계에서의 체계적인 실패 모드 식별과, 임상 환경에서의 실시간 예측 신뢰도 전달을 가능하게 하여, 안전하고 책임 있는 임상 의사결정을 지원한다. 나아가, 본 연구는 FDA의 의료 인공지능 규제 가이드라인에서 요구하는 임상적 검증, 사용성 평가, 위험 관리 요구사항에 부합하는 구조화된 접근법을 제시함으로써, 안전 중심의 의생명 영상 인공지능 개발에 실질적인 기여를 한다.
다국어 초록 (Multilingual Abstract)
Deep learning has revolutionized quantitative analysis of biomedical images, yet a critical gap persists between laboratory performance and real-world reliability. Models that excel on clean test sets often fail catastrophically when deployed in clini...
Deep learning has revolutionized quantitative analysis of biomedical images, yet a critical gap persists between laboratory performance and real-world reliability. Models that excel on clean test sets often fail catastrophically when deployed in clinical settings where imaging conditions differ from training data, exposing fundamental vulnerabilities in safety-critical biomedical applications. This dissertation addresses this challenge through two complementary contributions that must be co-designed for trustworthy clinical deployment: comprehensive robustness evaluation frameworks and efficient uncertainty quantification methods.
The first contribution is a systematic robustness evaluation framework for biomedical imaging that defines domain-specific corruption benchmarks with sixteen physically defined corruption types at five severity levels for both Optical Diffraction Tomography (ODT) and Computed Tomography (CT) applications. Beyond evaluation, the framework introduces CutPix for ODT bacterial classification and CutPixMixup for CT organ-at-risk segmentation, two augmentation strategies that explicitly balance shape and texture information through fractal pattern mixing and organ-aware patch operations. These methods address the fundamental shape-texture trade-off in deep learning, where texture-biased models achieve high clean accuracy but fail under corruption, while shape-biased models sacrifice accuracy for robustness. CutPix achieves the lowest mean Corrupted Error (mCE) of 42.97\% for DenseNet-121 and 56.03\% for ResNet-101 in ODT classification, while CutPixMixup achieves state-of-the-art robustness with mCE of 62.25\% for brainstem and 69.90\% for eye segmentation, demonstrating that balanced feature learning enables both high accuracy and strong corruption resilience.
The second contribution is Neural Bootstrapper (NeuBoots), an efficient uncertainty quantification method that provides well-calibrated confidence estimates with minimal computational overhead. Unlike traditional ensemble or Bayesian methods that require multiple model training procedures or multiple forward passes, NeuBoots applies bootstrap weights directly to high-level feature layers, achieving significantly faster inference and training while maintaining statistical validity. Experimental validation across prediction calibration, active learning, out-of-distribution detection, and semantic segmentation tasks demonstrates that NeuBoots achieves stable frequentist coverage close to the nominal 95\% level, comparable to standard bootstrapping, while significantly outperforming baseline methods in active learning and OOD detection tasks.
The central argument of this dissertation is that robustness evaluation and uncertainty quantification are interdependent: robustness without uncertainty quantification leads to silent failures when models encounter conditions beyond their training distribution, while uncertainty quantification without robustness leads to miscalibrated confidence estimates under distribution shift. Together, these contributions provide essential tools for developing AI systems suitable for FDA-regulated medical devices, enabling systematic identification of failure modes before clinical validation and real-time communication of prediction reliability for appropriate clinical decision-making. The methods directly address FDA regulatory requirements for AI Software as a Medical Device, providing structured approaches to clinical validation, usability evaluation, and risk management that align with regulatory guidance for safety-critical biomedical imaging applications.
목차 (Table of Contents)