
http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
Testing normality based on cumulative residual Kullback-Leibler information
백보현 Graduate School, Yonsei University 2015 국내석사
This paper proposes a test based on cumulative residual Kullback-Leibler (CRKL) information for testing normality. Previously, Rao et al. (2004) suggested the cumulative residual entropy (CRE) which is the extension of Shannon entropy. Cumulative residual entropy (CRE) supplements the shortcomings of Shannon entropy. Using this cumulative residual entropy (CRE), Baratpour and Rad (2012) proposed cumulative residual Kullback-Leibler (CRKL) information and test an exponentiality. In the present study, we expand the support of cumulative residual Kullback-Leibler information to the unrestricted support. Furthermore, the power of the expanded CRKL is compared to that of other typical normality tests.
농축산물의 가격은 변동성이 크기 때문에 생산자는 정확한 가격예측을 통해 출하시점을 선택해야만 하고 변동성의 변화 문제를 제외하더라도 정확한 가격예측이 이루어지지 않을 수도 있다. 본 연구에서는 주요 축산물 가격의 변동성의 변화가 가격에 어떤 영향을 미치는 지에 대해 논의한다. 본 연구에서는 변동성이 시간에 따라 변화하고 있는지 여부를 살펴보기 위해서 Engle의 조건부이분산모형을 이용하였다. 먼저 자료들은 일간 평균가격들로 구성되어 있었으나 상호 관찰치 숫자가 상이하였고, 한우지육을 제외한 두 품목에서는 미 거래일이 상당수 존재하여 상호간의 비교 가능과 관찰치 생략의 방지를 위해 월평균가격을 산술평균으로 계산하였다. 각 품목들의 이분산모형의 적용가능성을 살펴보기 위해 정규성 검정과 자기 상관 검정, 단위근 검정 등을 거쳤다. 정규성 검정시는 한우지육만이 정규성을 만족하는 것으로 나타나 조건부이분산(autoregressive conditional heteroscedasticity (ARCH) model) 모형은 한우지육에만 적용하고 다른 두 품목에 대해서는 EGARCH (expotential generalized ARCH)모형을 적용함이 타당함을 보였다. 자기 상관 검정에서는 모든 품목이 자기상관이 있음을 보여, 시계열의 불안정하여 조건부이분산성이 있음을 발견하였다. 그리고 단위근 검정시에는 돼지지육만이 단위근을 가지는 것으로 나타나 ARCH모형 적용시에 1차 차분하였다. 다음으로 ARCH모형의 추정을 위해 SBC를 계산하여 회귀식의 AR차수를 결정하였고 ARCH-LM 검정을 통해 ARCH 효과가 몇 번째 래그까지 존재하는지를 살펴보았다. 이 과정을 통해 모든 품목이 2차에서 8차까지 ARCH효과가 있음을 보였고, 실제 추정과정에서는 한우지육은 AR(2)-ARCH(1) 모형을 적용하였고 돼지지육은 AR(1) - EGARCH(1, 1) 모형을, 육우지육은 AR(4) - EGARCH(1, 1) 모형을 적용하였다. 변동성 변화가 있음을 ARCH과정을 통해 확인한 후 역사적 변동성을 계산하여 연속되는 두 달간의 변동성 비율을 비교하여 변동성 변화가 실제로 일어나는 시점을 조사하였다. 그리고 가격변화율과 변동성의 변화율과 상관계수를 계산하여 세 품목 모두 음의 상관관계가 있음을 보였다. 본 연구에서는 변동성 변화를 계산하여 변동성 변화와 가격 변화율간의 관계를 살펴보았다. 그러나 축산물 도매시장에서 가격이 결정될 때는 상품의 품질을 보고 가격이 결정되므로 이러한 질적 변수가 가격과 가격의 변동성에 미치는 영향을 간과하였다. 보다 엄밀한 추정을 위해서는 질적 영향이 제거된 가격자료를 이용해야할 것이고, 이에 대한 보완 연구가 필요할 것이다. The fact of the high price volatility of agricultural products confronts with the accurate price forecast on the problem of choosing selling time. But without considering the change of price volatility, it may not be correct forecasts. The study is to examine the effects of volatility change to price. Especially the prices of livestock products such as cow, pork, beef cattle are analysed. Engle's Autoregressive Conditional Heteroscedasticy(ARCH) model and Nelson's Expotential GARCH(EGARCH) model was used to examine time varying volatility. First, the data of monthly average prices are used. In the case of missing observations for the daily average prices, monthly data are used by calculating arithmetical average. Second, normality test, autocorrelation test and unit root test are conducted for the time series of three products. In normality test, only the distribution of cow price time series shows normality. So cow price time series is applied to ARCH model, the others is applied to EGARCH model. And all products turn out to have autocorrelations as a result of Ljung-Box's Q test and pork price is differenced according to Augmented Dickey-Fuller test. Next, model is decided and estimated. The lag period for the AR is decided by SBC. And with ARCH-LM test, they are tested whether to exist ARCH effects. As a result, Cow price time series fits AR(2)-ARCH(1) model, pork price time seires fits AR(1)-EGARCH(1,1) model and beef cattle price time series fits AR(4)-EGARCH(1,1) model. After finding price volatility changes, historical volatility of price series are examined considering volatility changing point of time. The correlation coefficient between price change rate and volatility change rate shows negative relations for all commodities. However, this study misses the quality of livestock products on pricing. For more future study, the quality variables should be considered for both the price change rate and volatility change rate.
핵의학검사실 내부정도관리 데이터의 정규성 가정 평가 : 보수적 측정불확도 산출 프레임워크의 구축
강현규 울산대학교 일반대학원 2026 국내석사
Purpose: To evaluate the normality of internal quality control (IQC) data used for measurement uncertainty (MU) estimation in a nuclear medicine laboratory and to implement a structured framework for adaptive and conservative MU model selection. Methods: IQC datasets from 45 assays performed between May and July 2025 at Asan Medical Center were segmented according to reagent kit lot continuity, yielding 129 pooled datasets. Normality was assessed using the Shapiro–Wilk test and visualized with Q–Q plots. Segments were classified as normal (p ≥ 0.05), mildly non-normal (0.01 ≤ p < 0.05), non-normal (p < 0.01), or non-testable (n < 15). Measurement uncertainty was calculated using four models: CV-based, Algorithm A, robust (MAD/IQR), and rectangular (Type B), selected by an integrated decision algorithm incorporating sample size, normality, and Δ-based stability indices. Results: Among 129 pooled datasets, 123 (95.3%) had n ≥ 15 and underwent normality testing; 59 (48.0%) were normal, 12 (9.8%) mildly non-normal, and 52 (42.3%) non-normal. Dataset size was the only significant predictor of normality (p = 0.003). Of the 127 analyzable segments (n ≥ 2), 100 (78.7%) required non-Gaussian or conservative MU models. MU from the proposed framework strongly correlated with CV-based MU (Spearman r = 0.97, p < 0.001) without systematic bias (Wilcoxon p = 0.503). Relative differences in MU was near zero across normality groups, but model-specific differed: Algorithm A showed minimal deviation (−0.03%), robust MAD and IQR produced higher MU (+22.3% and +12.0%), and the rectangular model showed wide variability. Overall, 62% of datasets yielded higher MU under the proposed framework. Conclusion: More than half of IQC datasets with sufficient sample size deviated from Gaussian assumptions, indicating that routine CV-based MU estimation may be inappropriate for many assays. The proposed framework maintains consistency with conventional estimation under normal conditions while adaptively providing more conservative uncertainty estimates when distributional assumptions are violated. These results support its applicability as a reproducible approach consistent with current international guidelines for MU evaluation in nuclear medicine laboratories. 목적: 핵의학검사실에서 측정불확도 산출에 활용되는 내부정도관리 데이터의 정규성 가정을 평가하고, 데이터 분포 특성에 따라 적응적으로(adaptive) 적용 가능한 보수적인 측정불확도 모델 선택 프레임워크를 구축하고자 하였다. 방법: 2025년 5월부터 7월까지 서울아산병원 핵의학검사실에서 시행된 45개 검사의 내부정도관리 데이터를 시약 키트 로트 연속성에 따라 구분하여 총 129개의 통합 데이터 세그먼트를 구성하였다. 각 세그먼트의 정규성은 Shapiro–Wilk 검정을 통해 평가하였고, Q–Q plot을 이용해 시각적으로 확인하였다. 세그먼트는 p값에 따라 정규(p ≥ 0.05), 경도 비정규(0.01 ≤ p < 0.05), 비정규(p < 0.01)로 분류하였고, 표본 크기가 15 미만인 세그먼트는 통계적 검정의 검정력이 부족하여 비검정 가능(non-testable) 세그먼트로 분리하였다. 측정불확도는 네 가지 모델(변동계수 기반, Algorithm A, 강건 추정[MAD/IQR], 직사각형[Type B])을 이용하여 산출하였으며, 표본 크기, 정규성, 그리고 측정불확도 차이에 기반한 안정성 지표를 통합한 의사결정 알고리즘에 따라 최종 모델을 선택하였다. 결과: 총 129개 세그먼트 중 6개(4.7%)는 표본 수가 15 미만이었으며, 이 중 2개(n = 1)는 측정불확도 산출에서 제외되었다. 정규성 검정이 가능한 123개(n ≥ 15) 세그먼트 중 59개(48.0%)는 정규, 12개(9.8%)는 경도 비정규, 52개(42.3%)는 비정규로 분류되었으며, 정규성에 유의하게 영향을 미친 요인은 표본 크기뿐이었다(p = 0.003). 분석 가능한 127개 세그먼트(n ≥ 2) 중 100개(78.7%)는 비정규 또는 보수적 측정불확도 모델을 필요로 하였다. 제안된 프레임워크로 산출된 측정불확도는 기존 변동계수 기반 추정치와 높은 상관관계(Spearman r = 0.97, p < 0.001)를 보였으며, 체계적인 편향은 관찰되지 않았다(Wilcoxon p = 0.503). 결론: 내부정도관리 데이터 중 정규성 검정이 가능한 세그먼트의 절반 이상이 가우시안 가정을 충족하지 않아 일률적인 변동계수 기반 불확도 산출은 다수의 검사항목에서 부적절할 수 있음을 확인하였다. 본 연구에서 제안한 프레임워크는 정규 조건에서는 기존 추정치와의 일관성을 유지하면서도, 분포 가정이 위배될 경우 더 보수적인 측정불확도 값을 산출하는 적응형 접근법을 제시한다. 이러한 결과는 본 프레임워크가 핵의학검사실에서 재현 가능하며 국제 지침과도 부합하는 신뢰도 높은 측정불확도 평가 방법으로 활용 가능함을 시사한다.
구조방정식 모형에서 정규성 가정 위배 시 ML의 대안 탐색
신은경 이화여자대학교 대학원 2025 국내석사
최대우도 방법은 구조방정식 모형을 추정할 때 가장 일반적으로 사용되는 방법 중 하나로, 추정 결과가 점근적으로 불편향적이고 일관적이며 효율적이라는 특징을 가진다. 하지만 이러한 특성은 방법이 전제하고 있는 주요 가정, 즉 자료가 정규분포를 따른다는 가정을 만족할 때에만 보장되는데, 심리학을 포함한 사회과학 분야에서는 정규성 가정이 위배되는 사례가 빈번히 보고되고 있다. 이는 추정 결과에 편향을 초래하여 통계적 추론의 타당성을 저하시키는 주요 요인으로 작용하기 때문에, 정규성 가정이 위배된 상황에서도 신뢰할 수 있는 결과를 주는 여러 대안적인 방법들이 탐색되어 왔다. 그러나 대안적 방법들의 수행도에 대한 연구가 지속되어 왔음에도 불구하고, 방법별 수행도가 연구마다 일관적이지 않아 적절한 추정 방법을 선택하기 위한 기준이 명확하지 않은 상황이다. 따라서 본 연구는 정규성 가정 위배 시 발생하는 문제에 대응할 수 있는 대안적인 방법을 정리하고, 이들 방법의 수행도를 비교·제시함으로써 연구자들에게 실질적인 지침을 제안하는 것을 목적으로 한다. 이를 위해 지난 30여 년간의 관련 연구를 통합하여 각 방법의 수행도를 체계적으로 비교하고 분석한다. 특히 기존 연구에서 공통적으로 설정된 조건을 중심으로 각 방법의 전반적인 경향과 특성을 분석하여, 연구자들이 상황에 맞는 추정 방법을 선택할 수 있는 근거를 제공하고자 한다. 먼저, 최대우도 방법에서 정규성 가정의 의미와 가정 위배가 추정 결과에 미치는 영향을 설명한다. 이를 통해 정규성 가정의 중요성을 명확히 하고, 정규성 가정이 충족되지 않은 경우에서 최대우도 방법을 이용했을 때의 문제점을 논의한다. 다음으로, 정규성 가정이 위배되었을 때에도 활용가능한 다양한 방법들을 소개하고, 이들 방법이 비정규성에 대응하는 원리를 논의한다. 구체적으로, 수정된 최대우도와 부트스트랩 및 베이지안 방법의 이론적 원리와 적용 가능성을 검토하여, 대안적 방법들의 특징을 이해할 수 있는 기초를 마련한다. 나아가, 각 방법의 수행도 비교라는 주요 목적을 달성하기 위해 기존 연구들을 체계적으로 탐색한 후 연구 결과를 조건별로 분류하고, 이를 표와 그림으로 시각화한다. 이를 통해 각 방법이 특정 조건에서 보이는 패턴을 정리하여 방법별 수행도를 보다 체계적이고 구체적으로 비교할 수 있도록 한다. 마지막으로, 위에서 논의된 이론적 논의와 수행도 비교 결과를 종합한 가이드라인을 제공하면서 본 연구의 의의와 한계에 관해 논한다. Maximum likelihood (ML), which is commonly used to estimate structural equation models, is based on the assumption of normality in the data. However, violations of the normality assumption are frequently reported in psychology and the social sciences, which can lead to biased estimation results and undermine the validity of statistical inferences. Although alternative methods that can provide reliable results under non-normal conditions have been explored, the performance of these methods has shown inconsistent patterns across studies, making it difficult to establish clear criteria for selecting appropriate methods. This study aims to address the problems posed by violations of the normality assumption and to explore alternative methods for dealing effectively with such violations. By integrating studies from the last 30 years of research, the study attempts to provide practical guidelines for researchers confronted with non-normality in their data. It first discusses the importance of the normality assumption in ML and examines the impact of its violation on estimation results. It then presents several alternative methods that are applicable under non-normal conditions and analyses the principles by which these methods deal with non-normality. Furthermore, previously published studies are systematically reviewed and categorized according to specific conditions, with the results visualized through tables and figures to compare the performance of different methods. Finally, the study integrates these discussions to propose guidelines for researchers and highlight their implications and limitations.
잠재분포를 활용한 IRT 분류정확도 추정 : 2-요소 정규혼합분포를 중심으로
Classification accuracy, which can be considered as the validity of criterion-referenced test, is a useful and essential information in test development and utilization of results, as it represents the accuracy of classifying examinees based on their achievement levels. As a result, several methods have been proposed to estimate classification accuracy and particularly, research on classification accuracy estimation using IRT has been actively conducted since the 2000s. According to previous studies, it has been observed that the estimation of IRT classification accuracy is influenced by several factors associated with the test. Additionally, examinees’ ability distribution has also been identified as another influencing factor in this regard. In practice, the assumption of normality for the examinees’ ability distribution is often inappropriate in educational evaluation or psychometric data. Despite the frequent occurrence of violations of the normality assumption regarding the examinees’ ability distribution, which can impact the estimation of classification accuracy, there has been a lack of systematic attempts to investigate the influence of various examinees’ ability distributions on the IRT classification accuracy estimation methods. In particular, the Guo method, which is difficult to find previous studies on, is also susceptible to the non-normality of examinees’ ability distribution. The Guo method assumes a non-informative prior distribution when estimating the posterior distribution, which represents the latent distribution of examinee. This makes it difficult to reflect the non-normality of examinees’ ability distribution. Furthermore, IRT classification accuracy estimation methods, including the Guo method, require accurate parameter estimation. To achieve this the estimation of parameters can be performed using the MMLE-EM method, and in MMLE, the analysis is typically conducted under the assumption that the latent distribution follows a normal distribution. However, the assumption of normality for the examinees’ ability distribution is often inappropriate. In such cases, assuming that the latent distribution follows a normal distribution can increase bias in parameter estimation and hinder accurate estimation of classification accuracy. In other words, it was necessary to verify whether accurate classification accuracy estimation is possible through the IRT classification accuracy estimation methods in cases where the assumption of normality is inappropriate. Moreover, in MMLE, it is possible to estimate the latent distribution along with parameter estimation in order to reduce the bias in parameter estimation. It was expected that by using the estimated latent distribution to estimate classification accuracy, it would better reflect the non-normality of the examinees’ ability distribution compared to the Guo method, which assumes a non-informative prior distribution. In this study, a new method (referred to as the 2NM method) for estimating classification accuracy was proposed by using the the estimated latent distribution obtained through the latent distribution estimation method assuming the two-component normal mixture distribution. The purpose of the simulation study was to examine the accuracy of the 2NM-latent distribution estimation, which is a prerequisite for utilizing the 2NM method. Additionally, the performance of the 2NM method was compared to that of the existing IRT classification accuracy estimation methods under the simulation conditions. To this end, the simulation study was designed with various conditions including examinees’ ability distribution, test length, sample size, ability estimation method, and cut-score location. The goal was to examine the accuracy of the 2NM-latent distribution estimation method based on different examinees' ability distributions and test lengths. Additionally, the study aimed to investigate the differences in classification accuracy estimates obtained from three methods based on variations in the examinees’ ability distribution, cut-score location, test length, and ability estimation method. Furthermore, to assess the stability and accuracy of the three classification accuracy estimation methods, the study calculated the standard error of estimate, bias, and root mean square error based on the true classification accuracy index under different simulation conditions. The results and conclusions of this study are as follows : First, 2NM-latent distribution estimation method can handle not only normal distribution but also various non-normal distributions with flexibility. It has demonstrated stable and accurate estimation of the latent distribution. Secondly, in non-normal situations where skewness and bimodality are present, the 2NM method has been shown to outperform other methods. Thirdly, the difference between the 2NM method and other methods became even more pronounced when the test length was shorter. Lastly, the 2NM method did not show any significant difference in ability estimation method, whereas the choice of ability estimation method should be carefully considered for the other methods depending on the situation. 준거참조검사의 타당도라 할 수 있는 분류정확도는 피험자 성취 수준 분류 결정의 정확성을 나타내는 지표로 검사 개발 및 결과 활용에 있어 유용하며 필수적으로 파악해야 하는 정보이다. 이에 따라 여러 분류정확도 추정 방법이 제안되었고 특히, IRT 분류정확도 추정 방법에 관한 연구가 2000년대 이후 활발하게 진행되었다. 선행 연구에 따르면 IRT 분류정확도 추정 방법은 검사와 관련된 여러 요인에 영향을 받으며, 피험자 능력 분포 또한 영향을 미치는 요인인 것으로 나타났다. 그리고 교육평가 및 심리측정 분야에서 피험자 능력 분포에 대한 정규성 가정이 부적절한 경우가 많았다. 이처럼 분류정확도 추정에 영향을 미칠 수 있는 피험자 능력 분포에 대한 정규성 가정의 위배 상황이 빈번하게 나타남에도 불구하고, 다양한 피험자 능력 분포가 IRT 분류정확도 추정 방법에 미치는 영향력을 체계적으로 규명하려는 시도는 부족하였다. 특히, 선행 연구를 찾아보기 어려운 Guo 방법 역시 피험자 능력 분포의 비정규성에 영향을 받을 가능성이 있었다. Guo 방법에서는 피험자의 잠재분포 즉, 사후분포를 구할 때 무정보적 사전분포를 가정하기에 피험자 능력 분포의 비정규성을 반영하기 어려울 수 있기 때문이다. 또한, Guo 방법을 포함한 IRT 분류정확도 추정 방법에는 정확한 모수 추정이 필요한데. 이를 위해 MMLE-EM 방법이 사용될 수 있고 MMLE에서는 기본적으로 잠재분포가 정규분포를 따른다는 가정하에 분석이 이루어진다. 하지만, 정규성 가정이 부적절한 경우가 많고 이때, 잠재분포가 정규분포를 따른다고 가정하면 편의가 증가해 분류정확도의 정확한 추정을 저해할 수 있는 것으로 나타났다. 다시 말해, 정규성 가정이 부적절한 상황에서 IRT 분류정확도 추정 방법을 통해 정확한 분류정확도 추정이 가능한지 확인할 필요가 있었다. 또한, MMLE에서는 모수 추정의 편의를 줄이기 위해 잠재분포를 함께 추정할 수 있는데 추정된 잠재분포를 활용하여 분류정확도를 추정한다면 무정보적 사전분포를 가정하는 Guo 방법에 비해 피험자 능력 분포의 비정규성을 더욱 제대로 반영할 수 있을 것으로 기대되었다. 이에 본 연구에서는 2요소-정규혼합분포를 가정한 잠재분포 추정 방법을 통해 추정된 잠재분포를 활용한 IRT 분류정확도 추정 방법(2NM 방법)을 새롭게 제안하였다. 모의실험의 목적은 2NM 방법을 활용하기 위한 전제 조건인 2NM-잠재분포 추정 방법의 정확성을 확인하고, 2NM 방법의 성능을 모의실험 조건에서 기존의 IRT 분류정확도 추정 방법과 비교하는 것이다. 이를 위해 잠재분포 및 분류정확도 추정에 영향을 미칠 수 있는 요인들로 모의실험 조건을 구성하여 피험자 능력 분포와 검사 길이에 따른 2NM-잠재분포 추정 방법의 정확성을 확인하고, 세 방법으로 산출한 분류정확도 추정치가 피험자 능력 분포와 분할점수 위치, 검사 길이 그리고 능력추정법에 따라 어떠한 차이 및 상호작용 효과를 보이는지 살펴보았다. 또한, 세 방법의 분류정확도 지수 추정에 대한 안정성과 정확성을 판단하기 위해 진 분류정확도 지수를 활용하여 모의실험 조건에 따른 추정의 표준오차, 편의, 평균 제곱근 오차를 계산하였다. 본 연구의 결과 및 결론은 다음과 같다. 첫째, 2NM-잠재분포 추정 방법은 정규분포뿐만 아니라 여러 비정규분포에 유연하게 대처해 안정적이고 정확하게 잠재분포를 추정하였다. 둘째, 편포 및 이봉성이 나타나는 비정규 상황에서는 2NM 방법이 다른 두 방법에 비해 성능이 우수한 것으로 나타났다. 셋째, 검사 길이가 짧을수록 2NM 방법과 다른 방법 간 차이는 더욱 크게 나타났다. 넷째, 2NM 방법은 능력추정법 간 차이가 없었으나 다른 두 방법은 상황에 따라 능력추정법 선택에 유의하여야 한다.
(The) goodness-of-fit test for normality based on general cumulative entropy
Jang, Min-ju Graduate School, Yonsei University 2017 국내석사
통계적 추론이나 검정에 있어서 정규성 가정을 만족하는지 검정하는 과정은 매우 중요하다. 본 연구에서는 V.Zardasht et al. (2015)의 연구를 확장하여 새로운 정규분포의 적합도 검정 방법을 제안하고자 한다. V.Zardasht et al. 은 분포 비교 함수로서 누적 잔차 엔트로피를 사용하여 지수 분포의 적합성 검정을 제시하였다. 그러나 누적 잔차 엔트로피는 항상 양수 값을 가지므로 정규 분포의 검정에는 적합하지 않다. 따라서, 본 연구에서는 누적 잔차 엔트로피 대신에 Park and Kim(2016)에 의해 제안된 일반화 누적 엔트로피를 분포 비교 함수로 이용하여 정규성 검정 통계량을 제안하고자 한다. 또한, 검정통계량의 근사적 정규 분포와 표본 수가 10,20,30,40,50일 때의 임계값을 제시하였다. 몬테카를로 시뮬레이션을 통해 많이 쓰이는 분포들에 대해 기존의 정규성 검정방법들과 검정력을 비교해 보았다. Testing normality is one of the most important processes for statistical analysis. We suggest a new goodness-of-fit test for normality by extending Zardasht et al. (2015). Zardasht et al. (2015) proposed a goodness-of-fit test for exponentiality by using cumulative residual entropy(CRE) as a comparison distribution function. However, it is not a appropriate comparison distribution function for testing normality, because CRE always takes non-negative values. Thus, we utilize general residual entropy(GRE) proposed by Kim and Park (2016) instead of CRE and suggest a normality test statistic. We propose an asymptotic normal distribution of test statistic and critical values for sample size n=10,20,30,40,50 through Monte-Calro simulations. A power study for comparison with several common tests for normality is presented with regard to popular alternatives.
Merging of radar and rain gauge data considering the rainfall characteristics
노용훈 Graduate School, Korea University 2017 국내박사
Improving the quality of radar rain rate data is important to maximize the application of data. Generally, high-quality rainfall information is used as a key input for hydrologic analysis. Quantitative precipitation estimation (QPE) has been performed to minimize the various errors of radar rain rate data. Recently, merging of the radar data and rain gauge data is important as it is expected to get higher-quality rain rate field. However, the quality of the radar rain rate data still does not meet the expected level. To make better quality of rain rate field, this study evaluated the effect of rainfall characteristics in the application of a data merging technique. Specifically, this study focused on (1) analysis of zero measurements and distribution function of rainfall data on the correlation coefficient, (2) evaluation of rainfall characteristics on the simple Kriging, and (3) evaluation of rainfall characteristics on the merging of radar and rain gauge rain rate data. First, the effect of zero measurements and distribution function on the correlation coefficient was analyzed in the application of the Kriging method. Generally, the correlation coefficient is found without considering the zero measurements of data and with the assumption that data follow the Gaussian distribution. However, rain rate data is generally positively skewed, also showed a strong spatial and temporal intermittency. These characteristics change the correlation coefficient which also affect the shape of the variogram. As a result, the shape of the variogram decides the covariance and finally the Kriging weighted values. Second, the effect of rainfall intermittency and log-normality on the simple Kriging was evaluated, in the estimation of rainfall spatial coverage. In this study, the artificial data and radar and rain gauge rain rate data were applied to the simple Kriging. The results showed that the zero values has the tendency to decrease the sample variance but to increase the correlation coefficient among data. When the data do not follow the Gaussian distribution, the correlation length could be longer when taking the natural logarithm to the original data. It was confirmed that the data intermittency and data log-normality should be considered to derive a proper variogram. Overall, it was found that the consideration of the data intermittency and data log-normality can improve the simple Kriging result. Especially, the effect of considering the data intermittency was found very significant. However, it was also found that several abnormally high values can be generated or the area of no rain can be decreased due to the longer correlation length. Third, the effect of rainfall intermittency and log-normality on the merging of radar and rain gauge rain rate data using the co-kriging was evaluated. As a result, for both variogram and cross-variogram, the correlation length was found to be longer, but sill height was smaller in the cases with consideration for data characteristics. The longest correlation length was derived when considering both the data intermittency and log-normality. Additionally, the co-kriging result can be better in case of considering the data characteristics. Especially, the data log-normality was found to be higher effect on the spatial coverage and quality of the rain rate field than the data intermittency. Furthermore, when considering the data characteristics, the mean of the merged rain rate field became more or less the same as that of the ground data.
Autoencoder-Based Robust Anomaly Detection for Multivariate Time Series with Various Normalities
시계열 데이터의 차원이 증가함에 따라 데이터 속에 정상과 비정상의 다양성이 증가하여, 정상 상황과 비정상 상황을 구분하는 것이 어려워졌다. 오토인코더는 고차원의 데이터를 처리하고 직관적으로 이상을 탐지하기 때문에 이상 탐지에 많이 사용된다. 그러나, 단일 인코더 구조는 정상 특징 추출 능력이 제한되어 재구성의 질이 최적화되지 못한다. 또한, 오토인코더의 우수한 일반화 능력과 노이즈에 대한 강건성으로 인해 정상과 비정상의 구별이 어렵다. 이러한 한계를 극복하기 위해, 비지도 학습 방식의 새로운 방법을 제안한다. 제안 방법은 듀얼 어텐션 메커니즘을 포함한 다중 인코더 구조를 통해 다양한 관점에서 정상 패턴을 학습한다. 또한, 메모리 모듈을 사용하여 정상에 대한 원형 정보를 기록한다. 이를 통해 정상 데이터에 대한 재구성의 질이 향상되고 이상을 안정적으로 탐지할 수 있다. 3개의 다변량 시계열 데이터셋에 대한 비교 실험을 통해, 다양한 정상성이 포함된 시나리오와 간단한 이진 분류 시나리오에서 제안 방법이 가장 우수한 성능을 보이는 것을 확인했다. 또한, 제안 방법의 구성 요소가 성능 향상에 미치는 영향을 검증했다. 심화 실험을 통해, 제안 방법이 정상성의 다양한 정도가 점진적으로 증가할 때에도 안정적으로 높은 성능을 보이는 것을 확인했다. As the dimensionality of time series data increases, distinguishing between normal and abnormal situations becomes challenging owing to the growing number of cases that can be classified as either. Autoencoders are widely used in anomaly detection because of their ability to handle high-dimensional data and detect anomalies intuitively. However, their single-encoder structure limits their ability to extract normal features, resulting in a suboptimal reconstruction quality. Furthermore, their strong generalization and robustness to noise make it difficult to differentiate between normality and abnormality. Motivated by these challenges, we proposed a novel method called multiple encoders with dual-attention memory network (MEDAM), which utilizes a multiple-encoders with dual-attention structure to learn normal patterns from various perspectives and memory modules to record normal profiles. This results in improved reconstruction quality for normal data and stable detection of anomalies. Extensive experiments conducted on three datasets (DSADS, PAMAP2, and PDgait) revealed the superiority of MEDAM over baseline methods in two distinct scenarios: detecting anomalies in data with varying normalities and performing straightforward binary classification. We also verified the positive effects of the detailed elements of MEDAM on anomaly detection. Further experiments demonstrated that MEDAM consistently detected anomalies, even when the diversity within normality systematically increased.
강우량 자료와 SPEI를 이용한 우리나라 습윤 및 건조 조건의 장기변동 경향성 분석
권기량 국립경국대학교 일반대학원 2026 국내박사
Long-term Trend Analysis of Wet and Dry Conditions Using Rainfall Data and SPEI in South Korea Kwon, Gi Ryang Department of Civil & Environmental Engineering Graduate School Gyeongkuk National University Abstract Global climate change has led to increasingly frequent extreme precipitation events and prolonged droughts worldwide. In particular, changes in the hydrological cycle have intensified the spatial and temporal variability of precipitation, simultaneously aggravating the contrasting phenomena of floods and droughts. In this era of climate change, comprehensive analyses from long-term and large-scale perspectives are required. Therefore, this study aims to identify the overall variation in the hydrological cycle by integrating analyses of wet and dry conditions, thereby enhancing the understanding of the complex characteristics of extreme weather events. The normality of precipitation data and the Standardized Precipitation Evapotranspiration Index (SPEI) was examined to verify whether the datasets satisfy normal distribution according to their periods and characteristics. The D'Agostino–Pearson and Shapiro–Wilk tests were applied for normality assessment. Once normality was determined, various statistical methods were used to detect long-term trends in the data. These methods were categorized into parametric and non-parametric tests. A linear regression model was used as the parametric approach, which assumes independence and normality of observations. However, because hydrological time series data often violate normality assumptions, non-parametric tests are frequently applied. Among them, the Mann–Kendall and Spearman’s rho (ρ) tests are the most widely used for detecting trends in hydrological time series. In this study, long-term trends of precipitation and SPEI were analyzed. Both annual and rainy seasonal precipitation showed slightly increasing trends, though not statistically significant at the 10% level. Similarly, most monthly precipitation data failed to reach statistical significance; however, a general decreasing trend was observed in June, while July, August, and September exhibited increasing tendencies. The long-term trend analysis of SPEI also revealed distinct regional characteristics. Except for a slight worsening trend in short-term (3–6 months) minimum SPEI, no significant changes were found in mean, maximum, or long-term (9–12 months) SPEI. Nonetheless, some localized areas exhibited statistically significant variations, and their spatial patterns were identified. Although extreme rainfall and heatwave events associated with climate change are increasingly frequent worldwide, only a few regions in Korea demonstrated statistically significant long-term trends. This suggests that Korea has not yet experienced changes strong enough to affect long-term climatic trends. However, as climate change continues to progress, ongoing monitoring and research will be essential to reflect these evolving patterns. Keywords: precipitation, rainfall, SPEI, trend, normality, Mann-Kendall test.