RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Advancing Stability and Validity in Machine Learning-Based Medical Data Analysis = 머신러닝 기반 의료 데이터 분석의 안정성 및 타당성 향상 연구

    한글로보기

    https://www.riss.kr/link?id=T17241972

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    인공지능과 데이터 저장/처리 기술의 발전은 다변수형 의료 데이터 분석에 큰 변화를 가져왔다. 해석 가능 인공지능(XAI)을 통해 인공지능이 중요한 변수들을 어떻게 해석하는지 이해할 수 있게 됨에 따라 기존에 통계학 중심으로 이루어지던 다변수형 의료 데이터 분석은 인공지능 분야로도 확장되기 시작했다. 그러나 의료 데이터는 흔히 데이터 수가 부족하고 정보의 차원이 커 과적합 문제에 취약하며, 통계학만큼의 검증과 신뢰도를 아직 의료 분야에서 얻지 못한 상황이다. 이러한 문제를 개선하기 위해 본 연구에서는 인공지능 모델의 전후에 적용할 수 있는 전처리 및 사후 처리 기법을 제안했다. 전처리로는 중요하지 않은 정보를 사전에 제거하는 알고리즘을 제안하였다. Bray-Curtis 유사성 기반의 매핑 변환을 통해 불필요한 변수를 안정적으로 제거하였으며 (Ch 2.2) SHAP 기반 이진화 기법을 통해 이진적인 특성을 가진 연속형 변수를 각 변수의 특성에 맞게 이진화하여 정보량을 직관적으로 압축하였다(Ch 2.3). 사후처리로는 인공지능 해석의 타당성을 검증하는 연구를 제안하였다. SHAP 값의 분포를 통계적으로 분석하고 유의미한 결과만을 압축하는 파이썬 패키지를 개발하였으며 (Ch 3.2) 모델 성능과 해석 타당성의 관계를 분석하여 실제로는 성능 이외의 요인이 타당성에 영향을 미침을 입증하였다(Ch 3.3). 본 연구는 ML 기반 해석의 안정성과 타당성 연구와 관련하여 새로운 관점들을 제시하였다.
    번역하기

    인공지능과 데이터 저장/처리 기술의 발전은 다변수형 의료 데이터 분석에 큰 변화를 가져왔다. 해석 가능 인공지능(XAI)을 통해 인공지능이 중요한 변수들을 어떻게 해석하는지 이해할 수 ...

    인공지능과 데이터 저장/처리 기술의 발전은 다변수형 의료 데이터 분석에 큰 변화를 가져왔다. 해석 가능 인공지능(XAI)을 통해 인공지능이 중요한 변수들을 어떻게 해석하는지 이해할 수 있게 됨에 따라 기존에 통계학 중심으로 이루어지던 다변수형 의료 데이터 분석은 인공지능 분야로도 확장되기 시작했다. 그러나 의료 데이터는 흔히 데이터 수가 부족하고 정보의 차원이 커 과적합 문제에 취약하며, 통계학만큼의 검증과 신뢰도를 아직 의료 분야에서 얻지 못한 상황이다. 이러한 문제를 개선하기 위해 본 연구에서는 인공지능 모델의 전후에 적용할 수 있는 전처리 및 사후 처리 기법을 제안했다. 전처리로는 중요하지 않은 정보를 사전에 제거하는 알고리즘을 제안하였다. Bray-Curtis 유사성 기반의 매핑 변환을 통해 불필요한 변수를 안정적으로 제거하였으며 (Ch 2.2) SHAP 기반 이진화 기법을 통해 이진적인 특성을 가진 연속형 변수를 각 변수의 특성에 맞게 이진화하여 정보량을 직관적으로 압축하였다(Ch 2.3). 사후처리로는 인공지능 해석의 타당성을 검증하는 연구를 제안하였다. SHAP 값의 분포를 통계적으로 분석하고 유의미한 결과만을 압축하는 파이썬 패키지를 개발하였으며 (Ch 3.2) 모델 성능과 해석 타당성의 관계를 분석하여 실제로는 성능 이외의 요인이 타당성에 영향을 미침을 입증하였다(Ch 3.3). 본 연구는 ML 기반 해석의 안정성과 타당성 연구와 관련하여 새로운 관점들을 제시하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The advancements in artificial intelligence and data storage/processing technologies have brought significant changes to multivariate medical data analysis. With feature importance analysis, it is now possible to understand how machine learning interprets key variables, leading to the expansion of multivariate medical data analysis from a traditionally statistics-centered field into the realm of machine learning. However, due to the limited sample size and high dimensionality of most medical data, AI models are often prone to overfitting and have not yet achieved the level of validation and reliability that statistics-based methods have in the medical field. To address these issues, this study proposed pre-processing and post-processing techniques applicable to AI models. For pre-processing, an algorithm was introduced to remove non-essential information in advance. Using a Bray-Curtis similarity-based mapping transformation, we reliably eliminated unnecessary variables (Ch 2.2), and with a SHAP-based binarization technique, we binarized continuous variables with binary characteristics according to each variable’s nature, allowing for intuitive data compression (Ch 2.3). For post-processing, we proposed a method to validate AI interpretations. We developed a Python package to statistically analyze SHAP value distributions and filter only significant results (Ch 3.2). Additionally, we examined the relationship between model performance and interpretability validity, demonstrating that factors beyond performance influence interpretability (Ch 3.3). Research to enhance the stability and validity of machine learning-based interpretations should continue, and this study offers new perspectives and directions in this field.
    번역하기

    The advancements in artificial intelligence and data storage/processing technologies have brought significant changes to multivariate medical data analysis. With feature importance analysis, it is now possible to understand how machine learning interp...

    The advancements in artificial intelligence and data storage/processing technologies have brought significant changes to multivariate medical data analysis. With feature importance analysis, it is now possible to understand how machine learning interprets key variables, leading to the expansion of multivariate medical data analysis from a traditionally statistics-centered field into the realm of machine learning. However, due to the limited sample size and high dimensionality of most medical data, AI models are often prone to overfitting and have not yet achieved the level of validation and reliability that statistics-based methods have in the medical field. To address these issues, this study proposed pre-processing and post-processing techniques applicable to AI models. For pre-processing, an algorithm was introduced to remove non-essential information in advance. Using a Bray-Curtis similarity-based mapping transformation, we reliably eliminated unnecessary variables (Ch 2.2), and with a SHAP-based binarization technique, we binarized continuous variables with binary characteristics according to each variable’s nature, allowing for intuitive data compression (Ch 2.3). For post-processing, we proposed a method to validate AI interpretations. We developed a Python package to statistically analyze SHAP value distributions and filter only significant results (Ch 3.2). Additionally, we examined the relationship between model performance and interpretability validity, demonstrating that factors beyond performance influence interpretability (Ch 3.3). Research to enhance the stability and validity of machine learning-based interpretations should continue, and this study offers new perspectives and directions in this field.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Tables v
    • List of Figures vi
    • 1 Introduction for medical data analysis 1
    • Abstract i
    • Contents ii
    • List of Tables v
    • List of Figures vi
    • 1 Introduction for medical data analysis 1
    • 1.1 Background 1
    • 1.1.1 Emergence of Big Data in Healthcare 1
    • 1.1.2 Limitations of Statistical Methods for Big Data 4
    • 1.1.3 Rise of ML based interpretation and challenges 5
    • 1.1.4 Problems of ML-Based Interpretation 6
    • 1.1.5 Proposal of Study 8
    • 1.2 Methods for tabular medical data analysis 9
    • 1.2.1 Conventional Statistics 9
    • 1.2.2 ML based Interpretation 13
    • 1.3 Application1-Watch-based blood pressure measurement 19
    • 1.3.1 Introduction 19
    • 1.3.2 Statistical analysis 23
    • 1.3.3 ML based analysis 25
    • 1.3.4 Discussion 27
    • 1.4 Application2-Decision Support System for influenza 28
    • 1.4.1 Introduction 28
    • 1.4.2 Statistical analysis 31
    • 1.4.3 ML based analysis 32
    • 1.4.4 Discussion 35
    • 2 Feature transformation for stable interpretation 37
    • 2.1 Background . 37
    • 2.1.1 Curse of dimensionality and Feature Engineering 37
    • 2.1.2 Project & Data Introduction(microbiome-IBD) 41
    • 2.2 Feature selection-Mapping transformation 43
    • 2.2.1 Highlight 43
    • 2.2.2 Background 43
    • 2.2.3 Methods 44
    • 2.2.4 Results . 50
    • 2.2.5 Discussion 54
    • 2.3 Feature Binarization-SHAP based Binarization 57
    • 2.3.1 Highlight 57
    • 2.3.2 Background 57
    • 2.3.3 Methods 58
    • 2.3.4 Results 61
    • 2.3.5 Discussion 65
    • 3 Validity of ML based interpretation 68
    • 3.1 Background 68
    • 3.1.1 Lack of validation step 68
    • 3.1.2 Debates on the validity 70
    • 3.1.3 Why is feature importance validity often overlooked in practice? 70
    • 3.2 Statistical validation on SHAP values 72
    • 3.2.1 Highlight 72
    • 3.2.2 Proposal 72
    • 3.2.3 Dataset Description 73
    • 3.2.4 Step1. Number of features to analyze 74
    • 3.2.5 Step2. Univariate analysis 76
    • 3.2.6 Step3. Interaction analysis 78
    • 3.2.7 Discussion 82
    • 3.3 Dependence of validity on performance 83
    • 3.3.1 Highlight 83
    • 3.3.2 Proposal 83
    • 3.3.3 Data and Methods 84
    • 3.3.4 Results . 88
    • 3.3.5 Probabilistic Approach 92
    • 3.3.6 Discussion 96
    • 4 Discussion 98
    • 4.1 Potential of Deep learning for interpretation 98
    • 4.2 Computational Cost 100
    • 4.3 Limitations & Future Works 101
    • 4.3.1 Application of statistical validation package to other studies . 101
    • 4.3.2 Unstandardized Pipeline for the analysis 103
    • 4.3.3 Distinctive and unique role from statistics 104
    • 4.3.4 Collaboration between engineering and medicine 106
    • 5 Conclusion 108
    • Abstract (In Korean) 132
    • Acknowlegement 133
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼