This study investigates how regression estimation using auxiliary information improves finite-population mean estimation when observations exhibit temporal dependence. A finite population with panel structure is first generated by random-number simula...
This study investigates how regression estimation using auxiliary information improves finite-population mean estimation when observations exhibit temporal dependence. A finite population with panel structure is first generated by random-number simulation, assigning each unit a time series of fixed length. The parameter of interest is the population mean of a survey variable. For auxiliary variables, the corresponding population means are treated as known, so that regression estimator can be constructed in the usual way.
From this finite population, repeated simple random samples without replacement (SRSWOR) of fixed size are drawn, and the same samples are a set of competing regression-based estimators. In this way, differences in performance can be attributed to the estimation methods themselves rather than to changes in the sampling design.
All comparisons are made at the level of regression estimators for the finite-population mean. As a baseline, a pooled ordinary least squares (OLS) regression estimator is considered, which ignores the time structure and treats the stacked panel as a single cross-section.
Next, Cochrane–Orcutt and Prais–Winsten transformations are applied using the estimated first-order autocorrelation of the OLS residuals, and the same regression specification is refitted on the transformed data to obtain transformation-based regression estimators.
Finally, GLS estimators are fitted that explicitly model the within-unit error covariance structure: one specification assumes an AR(1) process GLS_AR1, and another allows for a more flexible ARMA structure GLS_ARMA, For each method, the estimated regression coefficients are combined with the known auxiliary-information totals to produce a regression estimator of , which is then compared to the true finite-population mean.
Performance is evaluated by Monte Carlo simulation using standard design-based criteria. For each iteration the Durbin–Watson statistic is computed to summarize the overall level of first-order serial correlation in the sample. Across repetitions, the regression estimators are compared in terms of bias, root mean squared error (RMSE), mean reported standard error (Mean_SE), and empirical coverage of nominal 95% confidence intervals (95% 포함률). These metrics jointly describe how accurately and how reliably each method recovers the finite-population mean under the same sampling design and the same auxiliary information.
The simulation results show that the OLS regression estimator, which ignores temporal dependence, tends to have the largest bias and RMSE and often fails to achieve the nominal coverage level. The CO and PW estimators substantially reduce bias and RMSE relative to OLS, but they also exhibit distortions in the intercept and, in some cases, severe undercoverage of confidence intervals, reflecting the side effects of transformation-based correction. In contrast, GLS estimators that incorporate the covariance structure directly—especially GLS_ARMA—achieve the smallest bias and RMSE while maintaining mean standard errors and empirical coverage that are closest to their theoretical targets. Overall, the study indicates that in finite-population settings with autocorrelated errors, regression estimation with auxiliary information performs best when the error covariance is explicitly modeled, and that GLS_ARMA provides the most balanced improvement across bias, RMSE, Mean SE, and Coverage 95% .