Medical Artificial Intelligence (AI) represents a pivotal technology for advancing patient care and diagnostic accuracy. However, substantial challenges remain in the successful deployment of models developed in controlled laboratory environments into...
Medical Artificial Intelligence (AI) represents a pivotal technology for advancing patient care and diagnostic accuracy. However, substantial challenges remain in the successful deployment of models developed in controlled laboratory environments into real-world clinical settings. This thesis introduces a novel framework to resolve two critical impediments confronting medical AI: `Domain Shift,' the degradation of model performance when applied to data from different environments, and `Data Scarcity,' the insufficiency of high-quality data optimized for specific clinical tasks. Transcending the limitations of conventional algorithm-centric approaches, this study establishes a new paradigm that systematically integrates the `clinical domain knowledge' of medical experts into the data augmentation process, thereby enhancing the generalization performance and reliability of the models.
First, regarding tabular clinical data, this study begins by delineating the inadequacies of conventional data-driven approaches through a foundational analysis of real-world pediatric clinical data. To overcome the identified limitations—specifically the challenge of domain shift—we propose a `Dual Knowledge-Guided Data Augmentation' framework. This includes `Similarity-guided Mixup,' which generates synthetic data by reflecting clinical similarity between patients, and `Group-based Masking,' which simulates realistic missing data patterns. This approach enables a model trained solely on single-institution data to maintain robust predictive performance on data from other medical institutions, providing a pragmatic solution that circumvents the realistic constraints of medical data sharing.
Second, extending this philosophy of `knowledge integration' to the domain of medical imaging, this study addresses the dual challenges of data scarcity and suboptimal quality in coronary angiography. We conceptualize a novel clinical problem: `minor coronary artery segmentation,' an area largely overlooked in prior research. To this end, we constructed the world's first high-quality benchmark dataset (Fine-ARCADE) featuring precise annotations down to the microvascular level, supervised by a cardiology expert. Concurrently, we developed a specialized data augmentation technique (CAG-specific Copy-paste) that reflects the complex anatomical characteristics of coronary arteries. This anatomy-guided approach dramatically improves segmentation performance in microvascular regions that were previously intractable for existing models, thereby broadening the applicability of AI models for precision diagnostics.
Comprehensive experimental validation demonstrates that the proposed methodologies significantly advance the state-of-the-art in both domain generalization for tabular data and medical image segmentation. These studies, spanning both tabular and imaging data, empirically substantiate that the systematic integration of clinical knowledge is a foundational strategy for ensuring the generalization performance and reliability of medical AI models. These contributions establish a critical foundation for the development of intelligent medical systems capable of stable operation across diverse clinical environments.