Graph-valued data represent data residing on the vertices of the graphs, which are commonly referred to as graph signals. With advancements in monitoring tools and information technology, the availability of graph-valued data from diverse sources, suc...
Graph-valued data represent data residing on the vertices of the graphs, which are commonly referred to as graph signals. With advancements in monitoring tools and information technology, the availability of graph-valued data from diverse sources, such as social media, sensor networks, and brain activity, has increased. This growth has driven the demand for processing such data and extracting meaningful insights across various disciplines, including economics, epidemiology, and environmental science, leading to the development of graph signal processing in the past two decades (Shuman et al., 2013).
A key distinction in analyzing graph signals, compared to classical statistical methods, is the need to account for the underlying graph structure, which is discrete, irregular, and finite. In this dissertation, we present several novel methodologies for the statistical analysis of graph signals, broadly focusing on dimension reduction and feature extraction, developed both in the vertex domain (Chapter 3) and the graph spectral domain (Chapters 4-6). Specifically, we focus on four key topics: data fitting, cross-spectral analysis, dimension reduction, and factor modeling on graphs, which are explored in detail throughout the dissertation.
In Chapter 3, we propose a quantile-based fitting method for a noisy graph signal, consisting of both the underlying signal and noise on the graph. Unlike traditional data fitting methods, such as smoothing splines or quantile smoothing splines in Euclidean space, the proposed method is designed for the graph domain, considering the inherent structure of graphs. Prevalent graph signal fitting methods rely on optimization problems with $L_2$-norm fidelity, which primarily extract mean-based features of the signal (Stankovic et al., 2019). Moreover, they cannot effectively handle various signal distributions beyond the normal distribution, such as asymmetric or heavy-tailed distributions. In contrast, the proposed method efficiently represents the underlying structure of graph signals in the presence of outliers. More importantly, it identifies various distributional structures of graph signals beyond the mean feature. We further investigate the theoretical properties of the proposed solution, including its existence and uniqueness. Through comprehensive simulation studies and real data analysis using Manhattan Taxi data and United States temperature data, we demonstrate the promising performance of the proposed method.
Chapter 4 presents a cross-spectral analysis tool for bivariate graph signals, aimed at revealing the relationship between their constituent quantities, which is not readily apparent from individual spectral analyses (Segarra et al., 2018). We define joint weak stationarity and introduce graph cross-spectral density and coherence for bivariate graph processes. Several estimators for the cross-spectral density are proposed, and we investigate their theoretical properties. Furthermore, the effectiveness of these estimators is demonstrated through numerical experiments, including simulation studies and applications to French meteorological data. As an interesting extension, we also discuss robust spectral analysis of graph signals in the presence of outliers. This chapter broadens the scope of spectral analysis for graph signals to bivariate graph signals, which plays a crucial role in developing the analysis of multivariate graph signals in the subsequent chapters.
In Chapter 5, a novel principal component analysis method in the graph frequency domain is proposed for dimension reduction of multivariate data residing on graphs. The proposed method not only effectively reduces the dimensionality of multivariate graph signals, but also provides principal components that are interpretable with respect to the graph and a closed-form reconstruction of the original data. In addition, we investigate several propositions related to principal components and the reconstruction errors. We also introduce a graph spectral envelope and optimal graph frequency scaling, which aid in identifying common graph frequencies in multivariate graph signals and estimating the normalized amplitudes of the signal bases in each dimension, respectively. We demonstrate the validity of the proposed method through a simulation study and further apply it to the analysis of G20 economic data and Seoul Metropolitan Subway passenger data.
In Chapter 6, we introduce a new factor model in the graph frequency domain for multivariate graph signals. Utilizing graph filters, the proposed model extends the frequency-domain approach of the dynamic factor model from time series to graphs, enabling a graph-aware and multiscale interpretation of factors across graph frequencies. This approach reduces the dimensionality of graph signals and improves the understanding of their structure. It also allows the use of the extracted factors as the basis for subsequent analyses, such as clustering. We describe the estimation of factors and loadings and investigate the consistency of the factor estimator. In addition, we propose two consistent estimators for determining the number of factors. The finite sample performance of the proposed method is demonstrated through simulation studies under different graph structures, including a comparison with classical factor analysis and an exploration of how the graph structure affects the results. Furthermore, we show its effectiveness by applying it to the G20 economic data, water quality parameter data from the Miho-Cheon catchment on the Geum River network, and passenger data from the Seoul Metropolitan Subway.