RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Sequential Dimension Reduction for Non-Euclidean data and its Strong Consistency = 비유클리디안 데이터에서의 순차적 차원축소기법과 강일치성

    한글로보기

    https://www.riss.kr/link?id=T17315188

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    복잡한 데이터, 특히 비유클리디안 데이터의 차원 축소 기법은 현대 통계 분석과 데이터 과학 분야에서 점점 더 중요해지고 있다. 모든 가능한 저차원 부분공간을 동시에 탐색하는 전역적(global) 차원 축소 방법은 내재적 복잡성 및 높은 계산 비용으로 인해 비유클리디안 데이터에 적용하기 어려운 경우가 많다. 따라서 차원을 하나씩 증가시키거나 감소시키는 순차적(sequential) 차원 축소 기법이 현실적이고 유용한 대안으로 떠올랐다. 하지만 이러한 순차적 접근 방식은 이전 단계에서 결정된 저차원 구조에 의존하므로, 이론적인 분석과 해석이 쉽지 않다는 어려움이 있다. 실제 활용성에도 불구하고, 순차적 차원 축소 기법의 이론적 성질과 일관성(consistency)에 대한 연구는 여전히 부족한 실정이다.

    본 학위논문에서는 generalized Fréchet mean 프레임워크를 기반으로 순차적 차원 축소 기법에 대한 엄밀한 이론적 토대를 제시한다. 2장에서는 순차적 차원 축소뿐만 아니라 기존의 M-추정량(M-estimator) 등 다양한 통계적 추정 방법을 포괄하는 generalized Fréchet mean 프레임워크를 소개하고, 이를 이용하여 순차적 차원 축소 맥락에서 generalized Fréchet mean의 강일치성(strong consistency)을 이론적으로 증명한다. generalized Fréchet mean의 응용 사례로서 Fréchet $p$-평균, 초구면(hypersphere) 위에서의 측지 PCA(geodesic PCA), 그리고 일반 거리 공간(general metric space)에서의 $k$-메도이드($k$-medoid) 문제의 강일치성을 추가적으로 입증한다.

    나아가 본 논문에서는 두 가지 주요 비유클리디안 데이터 클래스—구성 데이터(compositional data)와 1차원 분포 데이터(one-dimensional distributional data)—에 적합한 실제적인 순차적 차원 축소 기법을 제안한다.

    구성 데이터는 마이크로바이옴 분석, 지질학, 환경과학 등 다양한 분야에서 자주 등장하며, 모든 성분이 음이 아닌 값을 가지고, 성분들의 합이 1로 제한된 벡터 형태의 데이터를 의미한다. 이러한 내재적 제약으로 인해, 기존의 벡터 공간에서 사용되던 일반적인 차원 축소 방법은 구성 데이터에 직접 적용할 수 없다. 3장에서는 이러한 구성 데이터의 구조적 특성을 명시적으로 고려하는 새로운 순차적 차원 축소 기법인 Compositional PCA를 제안한다. generalized Fréchet mean 프레임워크를 통해 Compositional PCA의 이론적 성질, 특히 강일치성을 증명하며, 이를 실제 마이크로바이옴 데이터에 적용하여 그 실용성과 효율성을 입증한다. 이를 위한 구체적인 계산 알고리즘 또한 제시한다.

    1차원 분포 데이터는 일일 전력 소비량 패턴, 사망률 분포, 인구 분포 등 밀도 함수 형태로 나타나는 중요한 비유클리디안 데이터의 한 예이다. 4장에서는 이러한 1차원 분포 데이터를 해당하는 분위수 함수(quantile function)로 변환한 후 이를 분석하는 새로운 차원 축소 기법인 Quantile PCA를 제안한다. 분위수 함수가 존재하는 공간은 무한차원이며, 또한 비감소성(monotonicity)을 유지해야 하는 강한 제약이 있어 이론적, 계산적으로 매우 까다로운 문제를 야기한다. 본 논문에서 제안한 Quantile PCA는 이러한 분위수 함수 공간의 기하학적 제약을 명시적으로 고려하여 효율적으로 차원을 축소할 수 있다. 우리는 Quantile PCA의 강일치성 등 주요 이론적 성질을 입증하고, 실제 계산을 위한 알고리즘을 개발한다. 실제 데이터 예시와 모의 실험(simulation studies)을 통해, Quantile PCA가 복잡한 분포 데이터의 주요 변동 패턴을 정확하게 포착하고 계산적으로 견고한 해법을 제공함을 확인한다.
    번역하기

    복잡한 데이터, 특히 비유클리디안 데이터의 차원 축소 기법은 현대 통계 분석과 데이터 과학 분야에서 점점 더 중요해지고 있다. 모든 가능한 저차원 부분공간을 동시에 탐색하는 전역적(gl...

    복잡한 데이터, 특히 비유클리디안 데이터의 차원 축소 기법은 현대 통계 분석과 데이터 과학 분야에서 점점 더 중요해지고 있다. 모든 가능한 저차원 부분공간을 동시에 탐색하는 전역적(global) 차원 축소 방법은 내재적 복잡성 및 높은 계산 비용으로 인해 비유클리디안 데이터에 적용하기 어려운 경우가 많다. 따라서 차원을 하나씩 증가시키거나 감소시키는 순차적(sequential) 차원 축소 기법이 현실적이고 유용한 대안으로 떠올랐다. 하지만 이러한 순차적 접근 방식은 이전 단계에서 결정된 저차원 구조에 의존하므로, 이론적인 분석과 해석이 쉽지 않다는 어려움이 있다. 실제 활용성에도 불구하고, 순차적 차원 축소 기법의 이론적 성질과 일관성(consistency)에 대한 연구는 여전히 부족한 실정이다.

    본 학위논문에서는 generalized Fréchet mean 프레임워크를 기반으로 순차적 차원 축소 기법에 대한 엄밀한 이론적 토대를 제시한다. 2장에서는 순차적 차원 축소뿐만 아니라 기존의 M-추정량(M-estimator) 등 다양한 통계적 추정 방법을 포괄하는 generalized Fréchet mean 프레임워크를 소개하고, 이를 이용하여 순차적 차원 축소 맥락에서 generalized Fréchet mean의 강일치성(strong consistency)을 이론적으로 증명한다. generalized Fréchet mean의 응용 사례로서 Fréchet $p$-평균, 초구면(hypersphere) 위에서의 측지 PCA(geodesic PCA), 그리고 일반 거리 공간(general metric space)에서의 $k$-메도이드($k$-medoid) 문제의 강일치성을 추가적으로 입증한다.

    나아가 본 논문에서는 두 가지 주요 비유클리디안 데이터 클래스—구성 데이터(compositional data)와 1차원 분포 데이터(one-dimensional distributional data)—에 적합한 실제적인 순차적 차원 축소 기법을 제안한다.

    구성 데이터는 마이크로바이옴 분석, 지질학, 환경과학 등 다양한 분야에서 자주 등장하며, 모든 성분이 음이 아닌 값을 가지고, 성분들의 합이 1로 제한된 벡터 형태의 데이터를 의미한다. 이러한 내재적 제약으로 인해, 기존의 벡터 공간에서 사용되던 일반적인 차원 축소 방법은 구성 데이터에 직접 적용할 수 없다. 3장에서는 이러한 구성 데이터의 구조적 특성을 명시적으로 고려하는 새로운 순차적 차원 축소 기법인 Compositional PCA를 제안한다. generalized Fréchet mean 프레임워크를 통해 Compositional PCA의 이론적 성질, 특히 강일치성을 증명하며, 이를 실제 마이크로바이옴 데이터에 적용하여 그 실용성과 효율성을 입증한다. 이를 위한 구체적인 계산 알고리즘 또한 제시한다.

    1차원 분포 데이터는 일일 전력 소비량 패턴, 사망률 분포, 인구 분포 등 밀도 함수 형태로 나타나는 중요한 비유클리디안 데이터의 한 예이다. 4장에서는 이러한 1차원 분포 데이터를 해당하는 분위수 함수(quantile function)로 변환한 후 이를 분석하는 새로운 차원 축소 기법인 Quantile PCA를 제안한다. 분위수 함수가 존재하는 공간은 무한차원이며, 또한 비감소성(monotonicity)을 유지해야 하는 강한 제약이 있어 이론적, 계산적으로 매우 까다로운 문제를 야기한다. 본 논문에서 제안한 Quantile PCA는 이러한 분위수 함수 공간의 기하학적 제약을 명시적으로 고려하여 효율적으로 차원을 축소할 수 있다. 우리는 Quantile PCA의 강일치성 등 주요 이론적 성질을 입증하고, 실제 계산을 위한 알고리즘을 개발한다. 실제 데이터 예시와 모의 실험(simulation studies)을 통해, Quantile PCA가 복잡한 분포 데이터의 주요 변동 패턴을 정확하게 포착하고 계산적으로 견고한 해법을 제공함을 확인한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Dimension reduction techniques for complex data, particularly non-Euclidean data, have become increasingly critical in modern statistical analysis and data science. Global dimension reduction methods, which simultaneously explore all possible lower-dimensional subspaces, are computationally challenging and often impractical for non-Euclidean data due to intrinsic complexity and computational cost. Therefore, sequential dimension reduction techniques—incrementally increasing or decreasing dimensionality—have emerged as a practical alternative. However, these sequential approaches inherently depend on previously determined lower-dimensional structures, complicating their theoretical analysis and interpretation. Despite their practical utility, theoretical studies addressing the properties and consistency of sequential dimension reduction methods remain limited. This dissertation establishes a rigorous theoretical foundation for sequential dimension reduction techniques.

    Chapter 2 introduces a comprehensive framework called the generalized Fréchet mean, which encompasses sequential dimension reduction methods and classical statistical estimators such as M-estimators by allowing random minimizing domains, thus extending existing methodologies. We provide a thorough theoretical analysis demonstrating the strong consistency of the generalized Fréchet mean, thereby laying a robust theoretical groundwork for developing and validating sequential dimension reduction techniques. As applications of the generalized Fréchet mean, we also establish the strong consistency of \fre $p$-means, geodesic PCA on hyperspheres, and the $k$-medoid problem in general metric spaces.

    Furthermore, we propose practical sequential dimension reduction methods tailored to two prominent classes of non-Euclidean data: compositional data and one-dimensional distributional data.

    Compositional data, which frequently arises in fields such as microbiome analysis, geology, and environmental sciences, refers to vector-valued data consisting of nonnegative components that sum to one. Due to these inherent compositional constraints, conventional dimension reduction methods for unrestricted vector spaces cannot be directly applied. In Chapter 3, we introduce a sequential dimension reduction method, Compositional PCA, explicitly designed to respect these constraints and efficiently capture the essential variation of compositional datasets. Theoretical properties of Compositional PCA, including strong consistency, have been shown by using the generalized \fre mean framework. We also develop practical computational algorithms and demonstrate the utility and effectiveness of this method through practical applications to real microbiome data.

    One-dimensional distributional data, exemplified by density functions such as daily electricity consumption patterns, mortality rates, and population distributions, represent another critical class of non-Euclidean data. In chapter 4, we propose a dimension reduction approach, Quantile PCA, which transforms these distributions into their corresponding quantile functions. The space of quantile functions is infinite-dimensional and constrained by monotonicity (functions must be non-decreasing), creating significant theoretical and computational challenges. Quantile PCA addresses these challenges by incorporating geometric constraints inherent in the quantile-function space. We establish key theoretical properties, including strong consistency, and develop practical computational algorithms. Through real-world examples and simulation studies, we illustrate that Quantile PCA accurately identifies key modes of variation in complex distributional data and provides computationally robust solutions.
    번역하기

    Dimension reduction techniques for complex data, particularly non-Euclidean data, have become increasingly critical in modern statistical analysis and data science. Global dimension reduction methods, which simultaneously explore all possible lower-di...

    Dimension reduction techniques for complex data, particularly non-Euclidean data, have become increasingly critical in modern statistical analysis and data science. Global dimension reduction methods, which simultaneously explore all possible lower-dimensional subspaces, are computationally challenging and often impractical for non-Euclidean data due to intrinsic complexity and computational cost. Therefore, sequential dimension reduction techniques—incrementally increasing or decreasing dimensionality—have emerged as a practical alternative. However, these sequential approaches inherently depend on previously determined lower-dimensional structures, complicating their theoretical analysis and interpretation. Despite their practical utility, theoretical studies addressing the properties and consistency of sequential dimension reduction methods remain limited. This dissertation establishes a rigorous theoretical foundation for sequential dimension reduction techniques.

    Chapter 2 introduces a comprehensive framework called the generalized Fréchet mean, which encompasses sequential dimension reduction methods and classical statistical estimators such as M-estimators by allowing random minimizing domains, thus extending existing methodologies. We provide a thorough theoretical analysis demonstrating the strong consistency of the generalized Fréchet mean, thereby laying a robust theoretical groundwork for developing and validating sequential dimension reduction techniques. As applications of the generalized Fréchet mean, we also establish the strong consistency of \fre $p$-means, geodesic PCA on hyperspheres, and the $k$-medoid problem in general metric spaces.

    Furthermore, we propose practical sequential dimension reduction methods tailored to two prominent classes of non-Euclidean data: compositional data and one-dimensional distributional data.

    Compositional data, which frequently arises in fields such as microbiome analysis, geology, and environmental sciences, refers to vector-valued data consisting of nonnegative components that sum to one. Due to these inherent compositional constraints, conventional dimension reduction methods for unrestricted vector spaces cannot be directly applied. In Chapter 3, we introduce a sequential dimension reduction method, Compositional PCA, explicitly designed to respect these constraints and efficiently capture the essential variation of compositional datasets. Theoretical properties of Compositional PCA, including strong consistency, have been shown by using the generalized \fre mean framework. We also develop practical computational algorithms and demonstrate the utility and effectiveness of this method through practical applications to real microbiome data.

    One-dimensional distributional data, exemplified by density functions such as daily electricity consumption patterns, mortality rates, and population distributions, represent another critical class of non-Euclidean data. In chapter 4, we propose a dimension reduction approach, Quantile PCA, which transforms these distributions into their corresponding quantile functions. The space of quantile functions is infinite-dimensional and constrained by monotonicity (functions must be non-decreasing), creating significant theoretical and computational challenges. Quantile PCA addresses these challenges by incorporating geometric constraints inherent in the quantile-function space. We establish key theoretical properties, including strong consistency, and develop practical computational algorithms. Through real-world examples and simulation studies, we illustrate that Quantile PCA accurately identifies key modes of variation in complex distributional data and provides computationally robust solutions.

    더보기

    목차 (Table of Contents)

    • Abstract
    • 1 Introduction 1
    • 2 Generalized Fréchet mean 4
    • 2.1 Introduction 4
    • 2.2 Generalized Fréchet means 6
    • Abstract
    • 1 Introduction 1
    • 2 Generalized Fréchet mean 4
    • 2.1 Introduction 4
    • 2.2 Generalized Fréchet means 6
    • 2.2.1 Definition 6
    • 2.2.2 Relations to previous extensions of Fréchet means 8
    • 2.3 Consistency of Generalized Fréchet means 9
    • 2.3.1 Convergence of random sets 9
    • 2.3.2 Consistency of Generalized Fréchet means 12
    • 2.3.3 Examples of cost functions satisfying BP consistency 16
    • 2.4 Applications 19
    • 2.4.1 k-medoids clustering in metric spaces 19
    • 2.4.2 Principal geodesic analysis in hyperspheres 21
    • 3 Principal component analysis for zero-inflated compositional data 24
    • 3.1 Introduction 24
    • 3.1.1 PCA for compositional data 26
    • 3.1.2 Related literature 31
    • 3.1.3 Organization of the thesis 32
    • 3.2 PCA Approaches for Compositional Data 33
    • 3.2.1 Problem Formulation 33
    • 3.2.2 The Proposed Methods 35
    • 3.2.3 Comparison to log-ratio PCA 39
    • 3.3 Computational Algorithm 40
    • 3.4 Theoretical Properties 46
    • 3.4.1 Existence 49
    • 3.4.2 Consistency 51
    • 3.5 Simulation studies 54
    • 3.5.1 Data generation 55
    • 3.5.2 Simulation results 57
    • 3.6 Real Data Analysis 60
    • 3.7 Discussion 65
    • 4 Wasserstein-Quantile PCA 68
    • 4.1 Introduction 68
    • 4.2 The Wasserstein-Quantile PCA 71
    • 4.2.1 Isometry between Wasserstein space and quantile space 71
    • 4.2.2 Methodology: Quantile PCA 73
    • 4.2.3 Population Counterpart of Quantile PCA 78
    • 4.2.4 Existing Methods 80
    • 4.3 Strong Consistency of smooth-restricted principal quantile subspace 85
    • 4.3.1 Minimization over random domain: Generalized Fréchet mean 86
    • 4.3.2 Smooth restriction of quantile subspaces 89
    • 4.3.3 Consistency of the smooth-restricted global principal quantile subspaces 91
    • 4.3.4 Consistency of the smooth-restricted forward principal quantile subspaces 93
    • 4.4 Nonequivalence of various principal quantile directions 95
    • 4.4.1 Forward and Separated Principal Quantile Directions 96
    • 4.4.2 Examples 97
    • 4.5 Implementation of Quantile PCA 101
    • 4.5.1 Principal quantile direction as an optimization problem 101
    • 4.5.2 Pointwise Discretization and block coordinate descent algorithm 103
    • 4.6 Numerical Illustrations 108
    • 4.6.1 Simulation Analysis 108
    • 4.6.2 Real Data Analysis 116
    • A Supplementary Material for Chapter 2 121
    • I Proofs and technical details 121
    • I.1 Conditions ensuring BP consistency 121
    • I.2 Proofs of Lemmas 1, 2 and 3 121
    • I.3 Condition 1 holds when 𝑀𝑛 ≡ 𝑀0 123
    • I.4 Proofs of Lemmas 4 and 5 and Theorem 1 124
    • I.5 Additional technical results 127
    • I.6 Proofs and technical details for Section 2.3.3 128
    • I.7 Technical details for Section 2.4.1 142
    • I.8 Technical details for Section 2.4.2 158
    • II Comparison with previous results on the law of large numbers for various extensions of Fréchet means 167
    • II.1 Fréchet ρ-mean of Huckemann et al. (2011) 167
    • II.2 ℭ-Fréchet mean of Schötz (2022) 169
    • II.3 H-Fréchet mean and Fréchet-p means of Schötz (2022) 170
    • II.4 C-restricted Fréchet p-mean of Evans and Jaffe (2024) 171
    • B Supplementary Material for Chapter 3 173
    • I Preliminary results 173
    • I.1 Equicontinuity of convex function 173
    • I.2 Convergence of subsequence 174
    • I.3 Geometric results for compositional subspaces 174
    • II Technical details and proofs 183
    • II.1 Proof of Lemma 6 183
    • II.2 Proof of Lemma 7 184
    • II.3 Proof of Theorem 5, and Corollary 4 185
    • II.4 Proof of Proposition 1 188
    • II.5 Proof of Proposition 2 189
    • II.6 Proof of Proposition 3 190
    • II.7 Consistency without Assumption 1 190
    • C Supplementary Material for Chapter 4 192
    • I Preliminary results 192
    • I.1 Geometry of quantile subspaces 192
    • II Technical Details 216
    • II.1 Technical details in Section 4.2 216
    • II.2 Technical details in Section 4.3 227
    • II.3 Technical details in Section 4.5 244
    • Abstract (in Korean) 260
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼