Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled high-resolution characterization of gene expression at the individual cell level. However, scRNA-seq data are inherently high-dimensional and often contain a limited number of samp...
Recent advances in single-cell RNA sequencing (scRNA-seq) have enabled high-resolution characterization of gene expression at the individual cell level. However, scRNA-seq data are inherently high-dimensional and often contain a limited number of samples, which leads to challenges such as overfitting and reduced generalization performance. Although integrating datasets from multiple institutions can help alleviate these issues, direct data sharing is restricted by stringent privacy regulations. Federated learning (FL) has therefore emerged as a promising paradigm for privacy-preserving collaborative model training. Nonetheless, most prior studies have primarily relied on intra-dataset simulations due to limited data accessibility, resulting in insufficient consideration of realistic inter-institutional heterogeneity.
To address these limitations, this study presents a comprehensive experimental framework that simulates heterogeneous conditions through multiple partitioning strategies within single datasets, and incorporates inter-dataset experiments using diverse scRNA-seq collections to approximate real-world FL scenarios. We compare FedAvg and FedProx, two representative FL algorithms differing in their robustness to data heterogeneity, and evaluate scalability by varying the number of participating clients. In addition, we benchmark a traditional machine learning model and four state-of-the-art deep learning architectures designed for cell type classification to assess their suitability for FL in high-dimensional single-cell transcriptomics.
Our empirical results show that FL can effectively mitigate instability and overfitting associated with scRNA-seq data while maintaining competitive performance across heterogeneous settings. The analyses further quantify the impact of data heterogeneity on model behavior and highlight the importance of algorithm selection in multi-institutional contexts. Overall, this study provides practical methodological guidance for applying FL to single-cell transcriptomic analysis and contributes to advancing privacy-preserving, multi-institutional research in the life sciences and biomedical fields.