의학 및 사회과학 분야에서는 같은 개체를 시간에 따라 반복하여 측정하는 경시적 데이터(longitudinal data)가 흔히 관찰된다. 최근 예측 모형에 대한 수요가 커지면서, 혼합효과 랜덤포레스트(M...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17388033
서울 : 성균관대학교 일반대학원, 2026
학위논문(석사) -- 성균관대학교 일반대학원 , 통계학과 , 2026. 2
2026
한국어
서울
Longitudinal boosting with mixed effects
84 p. : 삽화 ; 30 cm
지도교수: 김재직
참고문헌: p. 67-69
I804:11040-000000189951
0
상세조회0
다운로드의학 및 사회과학 분야에서는 같은 개체를 시간에 따라 반복하여 측정하는 경시적 데이터(longitudinal data)가 흔히 관찰된다. 최근 예측 모형에 대한 수요가 커지면서, 혼합효과 랜덤포레스트(M...
의학 및 사회과학 분야에서는 같은 개체를 시간에 따라 반복하여 측정하는 경시적 데이터(longitudinal data)가 흔히 관찰된다. 최근 예측 모형에 대한 수요가 커지면서, 혼합효과 랜덤포레스트(Mixed Effects Random Forest; MERF)나 혼합효과 경사 부스팅(Mixed Effects Gradient Boosting; MEGB)과 같이 경시적 데이터를 대상으로 한 기계학습(machine learning) 예측 모형들이 개발되고 있다. 그러나 기존 혼합효과 기반 트리 모형들은 임의효과가 개체 내 상관을 전부 설명한다고 가정하여 오차항들이 독립이라고 간주한다. 이러한 강력한 가정은 경시적 데이터에서 자주 발견되는 개체 내 시간 상관(serial correlation)이나 이분산(heteroscedasticity)이 존재하는 경우 예측 성능을 저하시킬 수 있다. 이에 본 연구에서는 경사 부스팅 알고리즘에 개체별 공분산행렬을 반영한 혼합효과 기반 트리 모형과 경시적 부스팅(longitudinal boosting)을 제안한다. 제안하는 모형은 개체-특정적(subject-specific) 및 시간 공변량을 활용해 오차 공분산행렬의 구조를 명시적으로 학습함으로써, 오차항의 독립성 가정에 의존하지 않는 유연한 예측을 가능하게 한다. 제안하는 모형의 성능은 선형 및 비선형 구조 모두에서 다양한 결측 데이터 메커니즘 및 임의효과 크기를 적용한 광범위한 모의실험을 통해 검증된다. 더 나아가, 본 연구에서는 실제 데이터를 활용한 성능 평가를 통해 경시적 부스팅의 현실 적용 가능성 또한 입증한다.
다국어 초록 (Multilingual Abstract)
Longitudinal data, which measures the same subjects repeatedly over time, is commonly used in medical and social sciences. Machine learning models for correlated data, such as mixed effects random forest and mixed effects gradient boosting, have been ...
Longitudinal data, which measures the same subjects repeatedly over time, is commonly used in medical and social sciences. Machine learning models for correlated data, such as mixed effects random forest and mixed effects gradient boosting, have been proposed as the demand for prediction models grows. However, existing mixed effect-based tree models assume residuals are independent, implying that random effects account for all within-subject correlations. This assumption may degrade predictive performance in the presence of within-subject serial correlation or heteroscedasticity, which are inherent characteristics of longitudinal data. To address this limitation, we propose a longitudinal boosting model with mixed effects that incorporates subject-specific covariance matrices into the gradient boosting algorithm. The proposed model enables flexible prediction by explicitly modeling the error covariance matrices through subject-specific and time covariates. The performance of the model is verified through extensive simulations by applying various missing data mechanisms and random effect sizes in both linear and nonlinear structures. We further demonstrate the practical applicability of the proposed model via performance evaluation on real-world datasets.
목차 (Table of Contents)