
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
A Data-Adaptive Principal Component Analysis: Use of Composite Asymmetric Huber Function
Lim, Yaeji,Oh, Hee-Seok Informa UK (TaylorFrancis) 2016 Journal of computational and graphical statistics Vol.25 No.4
<P>This article considers a new type of principal component analysis (PCA) that adaptively reflects the information of data. The ordinary PCA is useful for dimension reduction and identifying important features of multivariate data. However, it uses the second moment of data only, and consequently, it is not efficient for analyzing real observations in the case that these are skewed or asymmetric data. To extend the scope of PCA to non-Gaussian distributed data that cannot be well represented by the second moment, a new approach for PCA is proposed. The core of the methodology is to use a composite asymmetric Huber function defined as a weighted linear combination of modified Huber loss functions, which replaces the conventional square loss function. A practical algorithm to implement the data-adaptive PCA is discussed. Results from numerical studies including simulation study and real data analysis demonstrate the promising empirical properties of the proposed approach. Supplementary materials for this article are available online.</P>


Multimodel ensemble forecasting of rainfall over East Asia: regularized regression approach
Lim, Yaeji,Jo, Seongil,Lee, Jaeyong,Oh, Hee‐,Seok,Lee, Sang‐,Goo,Park, Yongtae,Kang, Hyun‐,Suk John Wiley Sons, Ltd 2014 International Journal of Climatology Vol.34 No.14
<P><B>ABSTRACT</B></P><P>This paper considers the problem of predicting the rainfall over East Asia from multimode outputs. For this purpose, we propose a new multimode ensemble method based on regularized regression approach, which consists of two steps, the pre‐processing step and the ensemble step. In the pre‐processing step, we improve prediction from each model output using regularized regression, and in the ensemble step, we apply regularization‐based regression method to combine the result from the pre‐processing step. The main benefits of the proposed method are that it improves prediction accuracy, and it is capable of solving the singularity problem so that it can integrate many climate variables from multimode outputs for a better prediction. The proposed method is applied to monthly outputs from nine general circulation models (GCMs) on boreal summer (June, July, and August) over 20 years (1983–2002). The prediction ability of the proposed ensemble forecast is compared with the observations and the outputs (prediction) from each GCM. The results show that the proposed method is capable of improving forecast accuracy by adjusting each model before combining.</P>


An improvement of seasonal climate prediction by regularized canonical correlation analysis
Lim, Yaeji,Jo, Seongil,Lee, Jaeyong,Oh, Hee‐,Seok,Kang, Hyun‐,Suk John Wiley Sons, Ltd. 2012 International Journal of Climatology Vol.32 No.10
<P><B>Abstract</B></P><P>This article proposes a statistical method based on the regularized canonical correlation analysis (RCCA) to improve on the conventional canonical correlation analysis (CCA) method for seasonal climate prediction. The fundamental idea of this method is to combine the regularization principle with the classical CCA to handle high‐dimensional data in which the number of variables is larger than the number of observations. This study focuses on prediction of future precipitation for the boreal summer (June‐July‐August, JJA) on both global and regional scales. We apply the RCCA method to the JJA hindcast/forecast archives for 29 years (1979–2007) obtained from the operational seasonal prediction system at Korea Meteorological Administration (KMA) in order to correct the model biases and provide a more accurate climate prediction. It is observed that the results from the RCCA method demonstrate a more accurate seasonal climate prediction as compared to the results from the general circulation model (GCM) and the CCA method coupled with empirical orthogonal functions (EOF) analysis, which is the modified CCA technique widely used, in terms of both correlation and mean square error. Copyright © 2011 Royal Meteorological Society</P>
M-estimation of the long-memory parameter by Laplace periodogram
Yaeji Lim 한국데이터정보과학회 2018 한국데이터정보과학회지 Vol.29 No.2
The estimation of the long-memory parameter is a crucial issue in the long-range dependent process. The log-regression method proposed by Geweke and Porter-Hudak (1983) is one of the popular semi-parametric approach to estimate the long-memory parameter. However, the conventional method is highly influenced by the presence of outliers or heavy-tailed distributed errors. This paper investigates the possibility of using Laplace periodogram to analyze long-memory processes. Laplace periodogram derived by the least absolute deviations in the harmonic regression procedure is a robust alternative to the ordinary periodogram for spectral analysis. Numerical studies including simulation study and real data analysis are presented for the comparison.
임예지,Lim, Yaeji 한국통계학회 2017 응용통계연구 Vol.30 No.3
Principal component analysis is a popular statistical method to reduce the dimension of the high dimensional climate data and to extract meaningful climate patterns. Based on the principal component analysis, we can further apply a regression approach for the linear prediction of future climate, termed as principal component regression (PCR). In this paper, we develop a new PCR method based on the regularized principal component analysis for spatial data proposed by Wang and Huang (2016) to account spatial feature of the climate data. We apply the proposed method to temperature prediction in the East Asia region and compare the result with conventional PCR results.
Clustering non-stationary advanced metering infrastructure data
Kang, Donghyun,Lim, Yaeji The Korean Statistical Society 2022 Communications for statistical applications and me Vol.29 No.2
In this paper, we propose a clustering method for advanced metering infrastructure (AMI) data in Korea. As AMI data presents non-stationarity, we consider time-dependent frequency domain principal components analysis, which is a proper method for locally stationary time series data. We develop a new clustering method based on time-varying eigenvectors, and our method provides a meaningful result that is different from the clustering results obtained by employing conventional methods, such as K-means and K-centres functional clustering. Simulation study demonstrates the superiority of the proposed approach. We further apply the clustering results to the evaluation of the electricity price system in South Korea, and validate the reform of the progressive electricity tariff system.
Classification via principal differential analysis
Jang, Eunseong,Lim, Yaeji The Korean Statistical Society 2021 Communications for statistical applications and me Vol.28 No.2
We propose principal differential analysis based classification methods. Computations of squared multiple correlation function (RSQ) and principal differential analysis (PDA) scores are reviewed; in addition, we combine principal differential analysis results with the logistic regression for binary classification. In the numerical study, we compare the principal differential analysis based classification methods with functional principal component analysis based classification. Various scenarios are considered in a simulation study, and principal differential analysis based classification methods classify the functional data well. Gene expression data is considered for real data analysis. We observe that the PDA score based method also performs well.

Aerosol optical depth prediction based on dimension reduction methods
Jungkyun Lee,Yaeji Lim 한국통계학회 2024 Communications for statistical applications and me Vol.31 No.5
As the concentration of fine dust has recently increased, numerous related studies are being conducted to address this issue. Aerosol optical depth (AOD) is a vital atmospheric parameter for measuring the optical properties of aerosols in the atmosphere, providing crucial information related to fine dust. In this paper, we apply three dimension reduction methods, nonnegative matrix factorization (NMF), empirical orthogonal functions (EOF) analysis and independent component analysis (ICA), to AOD data to analyze the patterns of fine dust in the East Asia region. Through a comparison of three dimension reduction methods, we observe that some patterns are observed in all three method, while some information are only extracted in a specific method. Additionally, we forecast AOD levels based on three methods, and compare the predictive performance of the three methodologies.
Forecasting GDP time series via the K-means based factor model
Jo Yejin,Lim Yaeji 한국통계학회 2025 Communications for statistical applications and me Vol.32 No.2
This article investigates a forecasting method for the European Union’s gross domestic product (GDP) based on a dynamic factor model and the K-means clustering method. The dynamic factor models are commonly used in analyzing macroeconomic data. Since the idiosyncratic terms of the factor model contain individual information about the time series, we apply K-means to these terms and classify multiple series into groups that reflect the series’ individual features. Then, we generate forecasting models for each clustering group, so that we use only the relevant variables in each model. By segmenting the data in this way, we aim to capture heterogeneous patterns that may not be well-represented in traditional forecasting approaches. Compared to other classical forecasting methods, the one herein proposed showed the best prediction results with the lowest prediction error values. The findings highlight the advantage of incorporating clustering techniques into factor models, offering a more tailored and accurate forecasting framework for economic data.


Park, Hyunjin,Lim, Yaeji,Ko, Eun Sook,Cho, Hwan-ho,Lee, Jeong Eon,Han, Boo-Kyung,Ko, Eun Young,Choi, Ji Soo,Park, Ko Woon American Association for Cancer Research 2018 Clinical Cancer Research Vol.24 No.19
<P><B>Purpose:</B> To develop a radiomics signature based on preoperative MRI to estimate disease-free survival (DFS) in patients with invasive breast cancer and to establish a radiomics nomogram that incorporates the radiomics signature and MRI and clinicopathological findings.</P><P><B>Experimental Design:</B> We identified 294 patients with invasive breast cancer who underwent preoperative MRI. Patients were randomly divided into training (<I>n</I> = 194) and validation (<I>n</I> = 100) sets. A radiomics signature (Rad-score) was generated using an elastic net in the training set, and the cutoff point of the radiomics signature to divide the patients into high- and low-risk groups was determined using receiver-operating characteristic curve analysis. Univariate and multivariate Cox proportional hazards model and Kaplan–Meier analysis were used to determine the association of the radiomics signature, MRI findings, and clinicopathological variables with DFS. A radiomics nomogram combining the Rad-score and MRI and clinicopathological findings was constructed to validate the radiomic signatures for individualized DFS estimation.</P><P><B>Results:</B> Higher Rad-scores were significantly associated with worse DFS in both the training and validation sets (<I>P</I> = 0.002 and 0.036, respectively). The radiomics nomogram estimated DFS [C-index, 0.76; 95% confidence interval (CI); 0.74–0.77] better than the clinicopathological (C-index, 0.72; 95% CI, 0.70–0.74) or Rad-score–only nomograms (C-index, 0.67; 95% CI, 0.65–0.69).</P><P><B>Conclusions:</B> The radiomics signature is an independent biomarker for the estimation of DFS in patients with invasive breast cancer. Combining the radiomics nomogram improved individualized DFS estimation. <I>Clin Cancer Res; 24(19); 4705–14. ©2018 AACR</I>.</P>