This study performed a multivariate statistical analysis of NPS pollution data of Saemangeum watershed to understand spatial and temporal patterns and to identify potential pollution sources. Multivariate statistical methods, such as cluster analysis ...
This study performed a multivariate statistical analysis of NPS pollution data of Saemangeum watershed to understand spatial and temporal patterns and to identify potential pollution sources. Multivariate statistical methods, such as cluster analysis (CA), discriminant analysis (DA) and factor analysis/principal component analysis (FA/PCA), were used to analyze the water quality dataset surveyed on monthly and 8-day basis including 6 parameters namely; DO, BOD, COD, SS, TN and TP from 22 monitoring sites within the study area for period of 2001 to 2013. Correlation between pollution items of the water quality data was carried using spearman’s correlation method. Results revealed strong positive correlation between BOD and COD, and also between BOD and TP indicating presence of biologically active organic matter. CA classified 12 months into 4 seasons and the 22 monitoring sites into two groups based on the two estuaries of Dongjin and Mangyeong. With the use of DA, 3 water quality parameters from the monthly dataset were identified namely DO, TN and TP and similarly for 8-day dataset were DO, BOD and COD were identified for distinguishing seasons, with 51.3% and 50.3% correct assignment respectively for temporal variation analysis. While 5 parameters that is BOD, COD, SS, TN and TP were revealed to correct assignment of 77.1% of monthly data and 4 parameters namely DO, BOD, SS and TN to 71.7% of 8-day data for the spatial variation analysis. FA/ PCA were applied to standardized log-transformed data sets of the 2 spatial groups for both monthly and 8-day datasets. This resulted in 2 latent factors each for both monthly and 8-day data explaining 70.6% and 60.9% of the total variance respectively in the dataset of Dongjin and similarly one latent factor each for both monthly and 8-day data explaining 76.9% and 63% of the total variance respectively of Mangyeong dataset. These results provide a basis for NPS pollution management and the extracted cluster analysis grouping information could be used in reducing the number of monitoring stations in the study area without missing much information.