공간 정보를 포함하는 데이터셋을 회귀 분석에 적용할 경우, 반응 변수와 설명 변수 간의 관계가 전체 공간 영역에서 동일하게 유지된다는 전통적인 가정이 현실적으로 부적절하거나 과도...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
공간 정보를 포함하는 데이터셋을 회귀 분석에 적용할 경우, 반응 변수와 설명 변수 간의 관계가 전체 공간 영역에서 동일하게 유지된다는 전통적인 가정이 현실적으로 부적절하거나 과도...
공간 정보를 포함하는 데이터셋을 회귀 분석에 적용할 경우, 반응 변수와 설명 변수 간의 관계가 전체 공간 영역에서 동일하게 유지된다는 전통적인 가정이 현실적으로 부적절하거나 과도하게 제한적일 수 있다. 이러한 문제를 해결하고 보다 유연한 분석을 가능하게 하기 위해, 본 연구에서는 변수 선택과 공간적으로 군집화된 계수 추정을 동시에 수행할 수 있는 베이지안 정규화 공간 클러스터 계수 모형(Bayesian Regularized Spatially Clustered Coefficient model, BRSCC)을 제안한다. 제안된 모형은 변수 선택을 위한 spike-and-slab 사전분포를 도입하여 반응 변수에 영향을 미치는 핵심 공변량을 식별하고, Voronoi tessellation 클러스터링 사전분포를 통해 해당 공변량의 계수가 공간적으로 유사한 지역 간에 군집화되는 구조를 반영한다.
모형 추론은 가변적인 차원을 가지는 파라미터 공간에서의 탐색을 가능하게 하는 리버서블 점프 마코프 체인 몬테카를로(Reversible Jump Markov chain Monte Carlo, RJMCMC) 기법을 통해 수행되며, 두 가지 형태의 점프(move)를 설계하여 효율적인 샘플링을 구현하였다. 또한, 알고리즘의 혼합 성능을 개선하기 위해 collapsed posterior와 parallel tempering 기법을 함께 적용하였다. 모델 구조의 유연성으로 인해 사후분포 분석이 어려운 점을 고려하여, 효과적인 사후 해석을 위한 전략적 접근도 함께 제안한다.
제안된 모형의 성능은 다양한 시뮬레이션 환경에서 체계적으로 검증되었다. 구체적으로는, 계수의 공간 분포가 Voronoi tessellation 기반의 클러스터 구조에 적합한 경우와 그렇지 않은 경우(Voronoi setting vs. non-Voronoi setting)로 나누어 모형의 표현력과 유연성을 평가하였다. 또한, 공분산 구조의 변화(예: 공간 상관 계수와 설명 변수 간 상관 정도)를 통해 모형의 민감도와 견고성을 분석하였으며, 노이즈 수준($\sigma^2$)의 변화에 따른 변수 선택 정확도와 계수 추정 정확도의 변화도 함께 관찰하였다. 비교 대상 모형으로는 기존의 Spatially Clustered Coefficient (SCC), Regularized Spatially Clustered Coefficient (RSCC) 방법을 포함하였고, 주요 성능 지표로는 변수 선택의 정·오 탐지율, 계수 추정의 평균 오차 및 랜드 지수(Rand Index) 등을 활용하였다.
실증 분석으로는 한국의 제19대 및 제20대 대통령 선거 데이터를 활용하였다. 해당 분석을 통해 당선자의 지역별 득표율에 영향을 미치는 주요 요인을 규명하고, 이들 요인의 효과가 지역적 클러스터를 따라 어떻게 달라지는지를 정량적으로 파악하였다. 선거 데이터의 강한 공간 종속성과 지역적 정치 성향 차이를 효과적으로 반영한 본 모형은, 공간 군집 기반의 회귀 분석에서 유용한 도구로 활용될 수 있음을 시사한다.
다국어 초록 (Multilingual Abstract)
Analyzing datasets with spatial information in regression setting often challenges the assumption that the relationship between the response variable and explanatory variables remains homogeneous across the spatial domain. To relax such assumption, we...
Analyzing datasets with spatial information in regression setting often challenges the assumption that the relationship between the response variable and explanatory variables remains homogeneous across the spatial domain. To relax such assumption, we propose a Bayesian Regularized Spatially Clustered Coefficient model, which detects cluster-wise varying effects and performs variable selection simultaneously. The proposed model identifies key covariates influencing the response variable by introducing a selection prior and uncovers their spatially clustered coefficient effects using a clustering prior. Bayesian inference is implemented via the Reversible Jump Markov chain Monte Carlo (RJMCMC) method, utilizing two types of reversible jump moves that enable efficient exploration of the parameter space. The collapsed posterior and parallel tampering are also considered for better mixing of the RJMCMC samples. We introduce strategic approaches for posterior analysis, as the varying dimension of the model space makes it challenging. A simulation study was conducted under various settings to demonstrate the effectiveness of the proposed approach. Finally, our model was applied to real data from the 19th and 20th presidential elections in South Korea to identify factors influencing the vote share of the elected president and to examine their spatial cluster-specific effects, as the data exhibit strong spatial dependence driven by the pronounced regional characteristics of political preferences.
목차 (Table of Contents)