RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    결측치를 포함한 데이터의 k-평균 군집분석 방법 비교 = Comparison of k-mean clustering with missing data

    한글로보기

    https://www.riss.kr/link?id=A108897565

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Cluster analysis is an unsupervised learning method to find heterogeneous clusters that capture similarities among items and separate different items into different clusters. Various cluster analysis techniques have been proposed, and the k-means clustering method, which minimizes the sum of Euclidean distances between cluster centroids and individual entities, is widely recognized as a standard cluster analysis method. When data include missing values, it is challenging to conduct cluster analysis, because it is impossible to calculate distances between centroids of clusters and incomplete items, resulting in excluding classification of these items. Techniques have been suggested to handle missing values in k-means clustering, including conducting cluster analysis after imputation of missing values or cluster analysis based on available information. In this study, we explore methods to perform k-means cluster analysis on data with missing values and evaluate performance of these methods using a simulation. The results of simulation studies indicate that conducting k-means cluster analysis after imputation yields the better performance than the one based on available information. Among the various imputation methods, k-nearest neighbors imputation performed the best.
    번역하기

    Cluster analysis is an unsupervised learning method to find heterogeneous clusters that capture similarities among items and separate different items into different clusters. Various cluster analysis techniques have been proposed, and the k-means clus...

    Cluster analysis is an unsupervised learning method to find heterogeneous clusters that capture similarities among items and separate different items into different clusters. Various cluster analysis techniques have been proposed, and the k-means clustering method, which minimizes the sum of Euclidean distances between cluster centroids and individual entities, is widely recognized as a standard cluster analysis method. When data include missing values, it is challenging to conduct cluster analysis, because it is impossible to calculate distances between centroids of clusters and incomplete items, resulting in excluding classification of these items. Techniques have been suggested to handle missing values in k-means clustering, including conducting cluster analysis after imputation of missing values or cluster analysis based on available information. In this study, we explore methods to perform k-means cluster analysis on data with missing values and evaluate performance of these methods using a simulation. The results of simulation studies indicate that conducting k-means cluster analysis after imputation yields the better performance than the one based on available information. Among the various imputation methods, k-nearest neighbors imputation performed the best.

    더보기

    참고문헌 (Reference)

    1 윤혜경 ; 김태영 ; 김용하, "군집분석을 활용한 포스트휴먼 교양의 범주화에 관한 연구" 한국자료분석학회 23 (23): 209-217, 2021

    2 송주원, "결측자료의 k-평균 군집분석" 한국자료분석학회 19 (19): 689-697, 2017

    3 송주원, "결측자료 분석에서 결측 비율이 결측자료 k-평균 군집분석에 미치는 영향" 한국자료분석학회 19 (19): 1273-1282, 2017

    4 Van Buuren, S., "mice : Multivariate imputation by chained equations in R" 45 : 1-67, 2011

    5 Chi, J. T., "k-pod : A method for k-means clustering of missing data" 70 (70): 91-99, 2016

    6 Hastie T, "impute: impute: Imputation for microarray data. R packageversion 1.74.1"

    7 Little, R. J. A., "Statistical Analysis With Missing Data" John Wiley 2002

    8 Doove, L. L., "Recursive partitioning for missing data imputation in thepresence of interaction effects" 72 : 92-104, 2014

    9 이은진 ; 김영서 ; 홍세희, "Network and Cluster Analysis of the Funding Relationship among the U.N. Agencies and State Donors" 한국자료분석학회 23 (23): 2547-2562, 2021

    10 Rubin, D. B., "Multiple Imputation for Nonresponse in Surveys" John Wiley 1987

    1 윤혜경 ; 김태영 ; 김용하, "군집분석을 활용한 포스트휴먼 교양의 범주화에 관한 연구" 한국자료분석학회 23 (23): 209-217, 2021

    2 송주원, "결측자료의 k-평균 군집분석" 한국자료분석학회 19 (19): 689-697, 2017

    3 송주원, "결측자료 분석에서 결측 비율이 결측자료 k-평균 군집분석에 미치는 영향" 한국자료분석학회 19 (19): 1273-1282, 2017

    4 Van Buuren, S., "mice : Multivariate imputation by chained equations in R" 45 : 1-67, 2011

    5 Chi, J. T., "k-pod : A method for k-means clustering of missing data" 70 (70): 91-99, 2016

    6 Hastie T, "impute: impute: Imputation for microarray data. R packageversion 1.74.1"

    7 Little, R. J. A., "Statistical Analysis With Missing Data" John Wiley 2002

    8 Doove, L. L., "Recursive partitioning for missing data imputation in thepresence of interaction effects" 72 : 92-104, 2014

    9 이은진 ; 김영서 ; 홍세희, "Network and Cluster Analysis of the Funding Relationship among the U.N. Agencies and State Donors" 한국자료분석학회 23 (23): 2547-2562, 2021

    10 Rubin, D. B., "Multiple Imputation for Nonresponse in Surveys" John Wiley 1987

    11 Hunt, L., "Mixture model clustering for mixed data with missing information" 41 (41): 429-440, 2003

    12 Little, R. J., "Missing-data adjustments in large surveys" 6 (6): 287-296, 1988

    13 Du, M., "Grid-Based Clustering Using Boundary Detection" 24 (24): 2022

    14 Hubert, L., "Comparing partitions" 2 : 193-218, 1985

    15 Wagstaff, K., "Clustering with missing values: No imputation required" Illinois Institute of Technology 649-658, 2004

    16 Strehl, A., "Cluster ensembles-a knowledge reuse framework for combining multiple partitions" 3 : 583-617, 2002

    17 Pfaffel, O, "ClustImpute: An R package for k-means clustering with build-in missing data imputation, Rpackage version 0.2.4"

    18 Moorthy, KA, "A review on missing value imputation algorithms for microarraygene expression data" 9 : 18-22, 2014

    19 Dunn, J. C., "A Fuzzy Relative of the ISODATA Process and Its Use in Detecting Compact Well-SeparatedClusters" 3 (3): 32-57, 1973

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼