RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Comparative Analysis of Deep Learning Models in a Federated Learning Environment for Single-cell Cell Type Classification = Comparative Analysis of Deep Learning Models in a Federated Learning Environment for Single-cell Cell Type Classification

    한글로보기

    https://www.riss.kr/link?id=T17449942

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled high-resolution characterization of gene expression at the individual cell level. However, scRNA-seq data are inherently high-dimensional and often contain a limited number of samples, which leads to challenges such as overfitting and reduced generalization performance. Although integrating datasets from multiple institutions can help alleviate these issues, direct data sharing is restricted by stringent privacy regulations. Federated learning (FL) has therefore emerged as a promising paradigm for privacy-preserving collaborative model training. Nonetheless, most prior studies have primarily relied on intra-dataset simulations due to limited data accessibility, resulting in insufficient consideration of realistic inter-institutional heterogeneity.
    To address these limitations, this study presents a comprehensive experimental framework that simulates heterogeneous conditions through multiple partitioning strategies within single datasets, and incorporates inter-dataset experiments using diverse scRNA-seq collections to approximate real-world FL scenarios. We compare FedAvg and FedProx, two representative FL algorithms differing in their robustness to data heterogeneity, and evaluate scalability by varying the number of participating clients. In addition, we benchmark a traditional machine learning model and four state-of-the-art deep learning architectures designed for cell type classification to assess their suitability for FL in high-dimensional single-cell transcriptomics.
    Our empirical results show that FL can effectively mitigate instability and overfitting associated with scRNA-seq data while maintaining competitive performance across heterogeneous settings. The analyses further quantify the impact of data heterogeneity on model behavior and highlight the importance of algorithm selection in multi-institutional contexts. Overall, this study provides practical methodological guidance for applying FL to single-cell transcriptomic analysis and contributes to advancing privacy-preserving, multi-institutional research in the life sciences and biomedical fields.
    번역하기

    Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled high-resolution characterization of gene expression at the individual cell level. However, scRNA-seq data are inherently high-dimensional and often contain a limited number of samp...

    Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled high-resolution characterization of gene expression at the individual cell level. However, scRNA-seq data are inherently high-dimensional and often contain a limited number of samples, which leads to challenges such as overfitting and reduced generalization performance. Although integrating datasets from multiple institutions can help alleviate these issues, direct data sharing is restricted by stringent privacy regulations. Federated learning (FL) has therefore emerged as a promising paradigm for privacy-preserving collaborative model training. Nonetheless, most prior studies have primarily relied on intra-dataset simulations due to limited data accessibility, resulting in insufficient consideration of realistic inter-institutional heterogeneity.
    To address these limitations, this study presents a comprehensive experimental framework that simulates heterogeneous conditions through multiple partitioning strategies within single datasets, and incorporates inter-dataset experiments using diverse scRNA-seq collections to approximate real-world FL scenarios. We compare FedAvg and FedProx, two representative FL algorithms differing in their robustness to data heterogeneity, and evaluate scalability by varying the number of participating clients. In addition, we benchmark a traditional machine learning model and four state-of-the-art deep learning architectures designed for cell type classification to assess their suitability for FL in high-dimensional single-cell transcriptomics.
    Our empirical results show that FL can effectively mitigate instability and overfitting associated with scRNA-seq data while maintaining competitive performance across heterogeneous settings. The analyses further quantify the impact of data heterogeneity on model behavior and highlight the importance of algorithm selection in multi-institutional contexts. Overall, this study provides practical methodological guidance for applying FL to single-cell transcriptomic analysis and contributes to advancing privacy-preserving, multi-institutional research in the life sciences and biomedical fields.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    단일세포 RNA 시퀀싱(scRNA-seq)의 발전은 개별 세포 수준에서의 유전자 발현을 고해상도로 분석할 수 있게 하였으나, 데이터의 고차원성과 제한된 샘플 수로 인해 과적합 및 일반화 성능 저하와 같은 문제가 발생한다. 다기관 데이터를 통합하면 이러한 한계를 완화할 수 있지만, 여러 개인정보 보호 규제로 인해 직접적인 데이터 공유가 어려운 상황이다. 이에 따라 연합학습(Federated Learning, FL)은 개인정보를 보존하면서도 협력적 모델 학습을 가능하게 하는 접근법으로 주목받고 있다. 그러나 기존 연구는 데이터 접근성의 제약으로 인해 주로 단일 데이터셋 기반의 시뮬레이션에 의존해 왔으며, 실제 기관 간 이질성을 충분히 반영하지 못했다.
    본 연구는 이러한 한계를 보완하기 위해, (i) 단일 데이터셋 내에서 다양한 데이터 분할 방식을 적용해 이질적 조건을 시뮬레이션하고, (ii) 서로 다른 scRNA-seq 데이터셋을 활용한 inter-dataset 실험을 포함하여 실제 FL 환경을 근사하는 포괄적 실험 프레임워크를 제시한다. 본 연구에서는 데이터 이질성에 대한 강건성에서 차이를 보이는 대표적 FL 알고리즘 FedAvg와 FedProx를 비교하고, 참여 기관 수를 변화시키며 확장성을 평가하였다. 또한 전통적 머신러닝 모델과 세포 유형 분류에 특화된 딥러닝 모델 네 종을 포함하여 총 다섯 가지 모델의 성능을 비교하여, 고차원 단일세포 데이터에서의 연합학습 적용 가능성을 검증하였다.
    실험 결과, 연합학습은 scRNA-seq 데이터에서 나타나는 학습 불안정성과 과적합 문제를 효과적으로 완화하면서 이질적 환경에서도 경쟁력 있는 성능을 유지하는 것으로 나타났다. 더불어 데이터 이질성이 모델 학습에 미치는 영향을 정량화하고, 다기관 환경에서 알고리즘 선택의 중요성을 확인하였다. 본 연구는 단일세포 분석에 연합학습을 적용하기 위한 실질적 방법론적 가이드를 제공하며, 생명과학 및 의생명 분야에서 개인정보를 보호하면서도 다기관 협업을 가능하게 하는 연구 기반 마련에 기여하고자 한다.
    번역하기

    단일세포 RNA 시퀀싱(scRNA-seq)의 발전은 개별 세포 수준에서의 유전자 발현을 고해상도로 분석할 수 있게 하였으나, 데이터의 고차원성과 제한된 샘플 수로 인해 과적합 및 일반화 성능 저하...

    단일세포 RNA 시퀀싱(scRNA-seq)의 발전은 개별 세포 수준에서의 유전자 발현을 고해상도로 분석할 수 있게 하였으나, 데이터의 고차원성과 제한된 샘플 수로 인해 과적합 및 일반화 성능 저하와 같은 문제가 발생한다. 다기관 데이터를 통합하면 이러한 한계를 완화할 수 있지만, 여러 개인정보 보호 규제로 인해 직접적인 데이터 공유가 어려운 상황이다. 이에 따라 연합학습(Federated Learning, FL)은 개인정보를 보존하면서도 협력적 모델 학습을 가능하게 하는 접근법으로 주목받고 있다. 그러나 기존 연구는 데이터 접근성의 제약으로 인해 주로 단일 데이터셋 기반의 시뮬레이션에 의존해 왔으며, 실제 기관 간 이질성을 충분히 반영하지 못했다.
    본 연구는 이러한 한계를 보완하기 위해, (i) 단일 데이터셋 내에서 다양한 데이터 분할 방식을 적용해 이질적 조건을 시뮬레이션하고, (ii) 서로 다른 scRNA-seq 데이터셋을 활용한 inter-dataset 실험을 포함하여 실제 FL 환경을 근사하는 포괄적 실험 프레임워크를 제시한다. 본 연구에서는 데이터 이질성에 대한 강건성에서 차이를 보이는 대표적 FL 알고리즘 FedAvg와 FedProx를 비교하고, 참여 기관 수를 변화시키며 확장성을 평가하였다. 또한 전통적 머신러닝 모델과 세포 유형 분류에 특화된 딥러닝 모델 네 종을 포함하여 총 다섯 가지 모델의 성능을 비교하여, 고차원 단일세포 데이터에서의 연합학습 적용 가능성을 검증하였다.
    실험 결과, 연합학습은 scRNA-seq 데이터에서 나타나는 학습 불안정성과 과적합 문제를 효과적으로 완화하면서 이질적 환경에서도 경쟁력 있는 성능을 유지하는 것으로 나타났다. 더불어 데이터 이질성이 모델 학습에 미치는 영향을 정량화하고, 다기관 환경에서 알고리즘 선택의 중요성을 확인하였다. 본 연구는 단일세포 분석에 연합학습을 적용하기 위한 실질적 방법론적 가이드를 제공하며, 생명과학 및 의생명 분야에서 개인정보를 보호하면서도 다기관 협업을 가능하게 하는 연구 기반 마련에 기여하고자 한다.

    더보기

    목차 (Table of Contents)

    • ABSTRACT III
    • CONTENTS V
    • LIST OF TABLES VII
    • LIST OF FIGURES VIII
    • CHAPTER 1. INTRODUCTION 1
    • ABSTRACT III
    • CONTENTS V
    • LIST OF TABLES VII
    • LIST OF FIGURES VIII
    • CHAPTER 1. INTRODUCTION 1
    • 1.1 BACKGROUND 1
    • 1.2 CONTRIBUTIONS 3
    • CHAPTER 2. METHODOLOGY 5
    • 2.1 DATA 6
    • 2.1.1 Intra-dataset 6
    • 2.1.2 Inter-dataset 8
    • 2.2 MODEL 10
    • 2.2.1 Traditional Machine Learning Model (XGBoost) 11
    • 2.2.2 Deep Learning Model for Cell Type Classification 12
    • 2.3 FEDERATED LEARNING ALGORITHM 15
    • 2.3.1 FedAvg 16
    • 2.3.2 FedProx 17
    • CHAPTER 3. RESULTS 18
    • 3.1 INTRA-DATASET EXPERIMENT 18
    • 3.1.1 Statistical Analysis 19
    • 3.1.2 Evaluation across Different Datasets 20
    • 3.1.3 Evaluation over Different Numbers of Clients 24
    • 3.2 INTER-DATASET EXPERIMENT 28
    • 3.3 RUNTIME COMPARISON 30
    • CHAPTER 4. CONCLUSION 32
    • 4.1 CONTRIBUTIONS 32
    • 4.2 LIMITATIONS 33
    • BIBLIOGRAPHY 35
    • APPENDIX A: INTER-DATASET SELECTION PROCESS 38
    • APPENDIX B: TRAIN/TEST LOSS CURVE QUANTITY-SKEW (SCIAE) 39
    • APPENDIX C: INTER-DATASET RESULTS 40
    • ABSTRACT (KOREAN) 44
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼