복잡한 데이터, 특히 비유클리디안 데이터의 차원 축소 기법은 현대 통계 분석과 데이터 과학 분야에서 점점 더 중요해지고 있다. 모든 가능한 저차원 부분공간을 동시에 탐색하는 전역적(gl...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17315188
서울 : 서울대학교 대학원, 2025
학위논문(박사) -- 서울대학교 대학원 , 통계학과 비유클리드 통계학 , 2025. 8
2025
영어
519.5
서울
xiv, 262 ; 26 cm
지도교수: 정성규
I804:11032-000000192472
0
상세조회0
다운로드복잡한 데이터, 특히 비유클리디안 데이터의 차원 축소 기법은 현대 통계 분석과 데이터 과학 분야에서 점점 더 중요해지고 있다. 모든 가능한 저차원 부분공간을 동시에 탐색하는 전역적(gl...
복잡한 데이터, 특히 비유클리디안 데이터의 차원 축소 기법은 현대 통계 분석과 데이터 과학 분야에서 점점 더 중요해지고 있다. 모든 가능한 저차원 부분공간을 동시에 탐색하는 전역적(global) 차원 축소 방법은 내재적 복잡성 및 높은 계산 비용으로 인해 비유클리디안 데이터에 적용하기 어려운 경우가 많다. 따라서 차원을 하나씩 증가시키거나 감소시키는 순차적(sequential) 차원 축소 기법이 현실적이고 유용한 대안으로 떠올랐다. 하지만 이러한 순차적 접근 방식은 이전 단계에서 결정된 저차원 구조에 의존하므로, 이론적인 분석과 해석이 쉽지 않다는 어려움이 있다. 실제 활용성에도 불구하고, 순차적 차원 축소 기법의 이론적 성질과 일관성(consistency)에 대한 연구는 여전히 부족한 실정이다.
본 학위논문에서는 generalized Fréchet mean 프레임워크를 기반으로 순차적 차원 축소 기법에 대한 엄밀한 이론적 토대를 제시한다. 2장에서는 순차적 차원 축소뿐만 아니라 기존의 M-추정량(M-estimator) 등 다양한 통계적 추정 방법을 포괄하는 generalized Fréchet mean 프레임워크를 소개하고, 이를 이용하여 순차적 차원 축소 맥락에서 generalized Fréchet mean의 강일치성(strong consistency)을 이론적으로 증명한다. generalized Fréchet mean의 응용 사례로서 Fréchet $p$-평균, 초구면(hypersphere) 위에서의 측지 PCA(geodesic PCA), 그리고 일반 거리 공간(general metric space)에서의 $k$-메도이드($k$-medoid) 문제의 강일치성을 추가적으로 입증한다.
나아가 본 논문에서는 두 가지 주요 비유클리디안 데이터 클래스—구성 데이터(compositional data)와 1차원 분포 데이터(one-dimensional distributional data)—에 적합한 실제적인 순차적 차원 축소 기법을 제안한다.
구성 데이터는 마이크로바이옴 분석, 지질학, 환경과학 등 다양한 분야에서 자주 등장하며, 모든 성분이 음이 아닌 값을 가지고, 성분들의 합이 1로 제한된 벡터 형태의 데이터를 의미한다. 이러한 내재적 제약으로 인해, 기존의 벡터 공간에서 사용되던 일반적인 차원 축소 방법은 구성 데이터에 직접 적용할 수 없다. 3장에서는 이러한 구성 데이터의 구조적 특성을 명시적으로 고려하는 새로운 순차적 차원 축소 기법인 Compositional PCA를 제안한다. generalized Fréchet mean 프레임워크를 통해 Compositional PCA의 이론적 성질, 특히 강일치성을 증명하며, 이를 실제 마이크로바이옴 데이터에 적용하여 그 실용성과 효율성을 입증한다. 이를 위한 구체적인 계산 알고리즘 또한 제시한다.
1차원 분포 데이터는 일일 전력 소비량 패턴, 사망률 분포, 인구 분포 등 밀도 함수 형태로 나타나는 중요한 비유클리디안 데이터의 한 예이다. 4장에서는 이러한 1차원 분포 데이터를 해당하는 분위수 함수(quantile function)로 변환한 후 이를 분석하는 새로운 차원 축소 기법인 Quantile PCA를 제안한다. 분위수 함수가 존재하는 공간은 무한차원이며, 또한 비감소성(monotonicity)을 유지해야 하는 강한 제약이 있어 이론적, 계산적으로 매우 까다로운 문제를 야기한다. 본 논문에서 제안한 Quantile PCA는 이러한 분위수 함수 공간의 기하학적 제약을 명시적으로 고려하여 효율적으로 차원을 축소할 수 있다. 우리는 Quantile PCA의 강일치성 등 주요 이론적 성질을 입증하고, 실제 계산을 위한 알고리즘을 개발한다. 실제 데이터 예시와 모의 실험(simulation studies)을 통해, Quantile PCA가 복잡한 분포 데이터의 주요 변동 패턴을 정확하게 포착하고 계산적으로 견고한 해법을 제공함을 확인한다.
다국어 초록 (Multilingual Abstract)
Dimension reduction techniques for complex data, particularly non-Euclidean data, have become increasingly critical in modern statistical analysis and data science. Global dimension reduction methods, which simultaneously explore all possible lower-di...
Dimension reduction techniques for complex data, particularly non-Euclidean data, have become increasingly critical in modern statistical analysis and data science. Global dimension reduction methods, which simultaneously explore all possible lower-dimensional subspaces, are computationally challenging and often impractical for non-Euclidean data due to intrinsic complexity and computational cost. Therefore, sequential dimension reduction techniques—incrementally increasing or decreasing dimensionality—have emerged as a practical alternative. However, these sequential approaches inherently depend on previously determined lower-dimensional structures, complicating their theoretical analysis and interpretation. Despite their practical utility, theoretical studies addressing the properties and consistency of sequential dimension reduction methods remain limited. This dissertation establishes a rigorous theoretical foundation for sequential dimension reduction techniques.
Chapter 2 introduces a comprehensive framework called the generalized Fréchet mean, which encompasses sequential dimension reduction methods and classical statistical estimators such as M-estimators by allowing random minimizing domains, thus extending existing methodologies. We provide a thorough theoretical analysis demonstrating the strong consistency of the generalized Fréchet mean, thereby laying a robust theoretical groundwork for developing and validating sequential dimension reduction techniques. As applications of the generalized Fréchet mean, we also establish the strong consistency of \fre $p$-means, geodesic PCA on hyperspheres, and the $k$-medoid problem in general metric spaces.
Furthermore, we propose practical sequential dimension reduction methods tailored to two prominent classes of non-Euclidean data: compositional data and one-dimensional distributional data.
Compositional data, which frequently arises in fields such as microbiome analysis, geology, and environmental sciences, refers to vector-valued data consisting of nonnegative components that sum to one. Due to these inherent compositional constraints, conventional dimension reduction methods for unrestricted vector spaces cannot be directly applied. In Chapter 3, we introduce a sequential dimension reduction method, Compositional PCA, explicitly designed to respect these constraints and efficiently capture the essential variation of compositional datasets. Theoretical properties of Compositional PCA, including strong consistency, have been shown by using the generalized \fre mean framework. We also develop practical computational algorithms and demonstrate the utility and effectiveness of this method through practical applications to real microbiome data.
One-dimensional distributional data, exemplified by density functions such as daily electricity consumption patterns, mortality rates, and population distributions, represent another critical class of non-Euclidean data. In chapter 4, we propose a dimension reduction approach, Quantile PCA, which transforms these distributions into their corresponding quantile functions. The space of quantile functions is infinite-dimensional and constrained by monotonicity (functions must be non-decreasing), creating significant theoretical and computational challenges. Quantile PCA addresses these challenges by incorporating geometric constraints inherent in the quantile-function space. We establish key theoretical properties, including strong consistency, and develop practical computational algorithms. Through real-world examples and simulation studies, we illustrate that Quantile PCA accurately identifies key modes of variation in complex distributional data and provides computationally robust solutions.
목차 (Table of Contents)