RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    워드 임베딩 기법을 활용한 범주형 데이터의 군집분석 : 고차원 데이터를 중심으로

    한글로보기

    https://www.riss.kr/link?id=T15372496

    • 저자
    • 발행사항

      서울 : 국민대학교 일반대학원, 2019

    • 학위논문사항
    • 발행연도

      2019

    • 작성언어

      한국어

    • DDC

      658.4038 판사항(23)

    • 발행국(도시)

      서울

    • 형태사항

      vi, 32 p. : 삽화 ; 26 cm

    • 일반주기명

      Cluster Analysis on Categorical Data using Word Embedding Method : Focused on a High Dimensional Data
      지도교수 : 정여진
      참고문헌 : p. 28-30

    • UCI식별코드

      I804:11014-200000222847

    • 소장기관
      • 국민대학교 성곡도서관 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    머신 러닝 분야의 대표적인 비지도 학습 방법 중 하나인 군집분석은 데이터를 서로 유사한 집단끼리 묶어주는 분석으로, 마케팅, 공학, 의학 등 다양한 분야에서 활용되고 있다(Wilson et al, 2011). 오늘날 군집 분석과 관련된 연구는 꾸준히 진행되고 있는데, 데이터가 수치형 일 때 활용될 수 있는 연구가 주를 이루고 있으며 범주형 데이터의 군집분석과 관련된 연구는 활발하게 진행되고 있지 않다(Mingoti et al, 2012).
    이에 본 연구에서는 범주형 데이터의 군집분석 시, 텍스트 분석에서 주로 사용되고 있는 워드 임베딩 기법을 활용하여 데이터를 수치형으로 변환을 한 뒤 수치형 데이터에 대한 군집분석 방법을 적용하는 방법을 제시하고자 한다.
    워드 임베딩은 현재 가장 많이 사용 되고 있는 기법인 Word2vec, FastText, Glove 기법을 각각 적용하였고, 각 기법 적용시 어떠한 성능의 차이를 보이는지를 비교분석 하였다.
    또한 제시하는 모형의 성능을 기존의 범주형 데이터의 군집분석 모형과 비교해 보면서 모형의 우수성을 검증하였고, 이때 가장 많이 알려진 방법인 K-mode, ROCK 등 의 방법과 비교분석 하였다.
    데이터의 구조에 따른 모형의 성능의 변화를 파악하기 위해 시뮬레이션을 통해 다양한 조건의 데이터를 생성한 뒤 각 데이터 조건별 모형의 성능을 비교 하였고, 나아가 실제 데이터에서도 모형이 잘 군집하는지를 평가하기 위하여 실제 데이터를 통해 모형의 성능을 평가 하였다.
    번역하기

    머신 러닝 분야의 대표적인 비지도 학습 방법 중 하나인 군집분석은 데이터를 서로 유사한 집단끼리 묶어주는 분석으로, 마케팅, 공학, 의학 등 다양한 분야에서 활용되고 있다(Wilson et al, 201...

    머신 러닝 분야의 대표적인 비지도 학습 방법 중 하나인 군집분석은 데이터를 서로 유사한 집단끼리 묶어주는 분석으로, 마케팅, 공학, 의학 등 다양한 분야에서 활용되고 있다(Wilson et al, 2011). 오늘날 군집 분석과 관련된 연구는 꾸준히 진행되고 있는데, 데이터가 수치형 일 때 활용될 수 있는 연구가 주를 이루고 있으며 범주형 데이터의 군집분석과 관련된 연구는 활발하게 진행되고 있지 않다(Mingoti et al, 2012).
    이에 본 연구에서는 범주형 데이터의 군집분석 시, 텍스트 분석에서 주로 사용되고 있는 워드 임베딩 기법을 활용하여 데이터를 수치형으로 변환을 한 뒤 수치형 데이터에 대한 군집분석 방법을 적용하는 방법을 제시하고자 한다.
    워드 임베딩은 현재 가장 많이 사용 되고 있는 기법인 Word2vec, FastText, Glove 기법을 각각 적용하였고, 각 기법 적용시 어떠한 성능의 차이를 보이는지를 비교분석 하였다.
    또한 제시하는 모형의 성능을 기존의 범주형 데이터의 군집분석 모형과 비교해 보면서 모형의 우수성을 검증하였고, 이때 가장 많이 알려진 방법인 K-mode, ROCK 등 의 방법과 비교분석 하였다.
    데이터의 구조에 따른 모형의 성능의 변화를 파악하기 위해 시뮬레이션을 통해 다양한 조건의 데이터를 생성한 뒤 각 데이터 조건별 모형의 성능을 비교 하였고, 나아가 실제 데이터에서도 모형이 잘 군집하는지를 평가하기 위하여 실제 데이터를 통해 모형의 성능을 평가 하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Clustering algorithms is technique for grouping similar data and have been used in a variety of fileds such as engineering, medicine, marketing, etc. There are lot of study about clustering analysis but, majority of studies are about algorithms for nimerical data.
    In this study, we propose a method that transform categorical data to numerical data using word embedding. We used three word embedding model(Skip-gram, FastText, Glove) and compared the performance with algorithms for categorical data(K-mode, ROCK).
    To determine the performance of the model depending on the structure of the data, we generated data with different conditions and furthermore, evaluated performance of the model through real data. we used Silhouette score and Adjusted Rand score for performance evaluation.
    By the Simulation, We Compared performance of the model by the number of categories and the number of data and as a result, embedding using the glove shown the best performance except where the number of categories is high and the number of data is low.
    We compared performance of the model through the real hospital care data and performance was good in order of K-means using glove embedding, K-mode, K-means using word2vec, K-means using FastText.
    번역하기

    Clustering algorithms is technique for grouping similar data and have been used in a variety of fileds such as engineering, medicine, marketing, etc. There are lot of study about clustering analysis but, majority of studies are about algorithms for ni...

    Clustering algorithms is technique for grouping similar data and have been used in a variety of fileds such as engineering, medicine, marketing, etc. There are lot of study about clustering analysis but, majority of studies are about algorithms for nimerical data.
    In this study, we propose a method that transform categorical data to numerical data using word embedding. We used three word embedding model(Skip-gram, FastText, Glove) and compared the performance with algorithms for categorical data(K-mode, ROCK).
    To determine the performance of the model depending on the structure of the data, we generated data with different conditions and furthermore, evaluated performance of the model through real data. we used Silhouette score and Adjusted Rand score for performance evaluation.
    By the Simulation, We Compared performance of the model by the number of categories and the number of data and as a result, embedding using the glove shown the best performance except where the number of categories is high and the number of data is low.
    We compared performance of the model through the real hospital care data and performance was good in order of K-means using glove embedding, K-mode, K-means using word2vec, K-means using FastText.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 1
    • 제2장 관련연구 3
    • 2.1 수치형 데이터의 군집분석 3
    • 2.1.1 K-means 3
    • 2.1.2 Mixture model 4
    • 제1장 서론 1
    • 제2장 관련연구 3
    • 2.1 수치형 데이터의 군집분석 3
    • 2.1.1 K-means 3
    • 2.1.2 Mixture model 4
    • 2.2 범주형 데이터의 군집분석 4
    • 2.2.1 K-mode 4
    • 2.2.2 ROCK 5
    • 2.3 워드 임베딩 6
    • 2.3.1 Word2vec 7
    • 2.3.2 FastText 8
    • 2.3.3 Glove 8
    • 제3장 워드 임베딩 기법을 활용한 범주형 변수의 군집분석 10
    • 3.1 제안 기법 10
    • 3.1.1 Category Embedding 10
    • 3.1.2 Observation Embedding 11
    • 3.1.3 Cluster Analysis 13
    • 3.2 성능 평가 13
    • 3.2.1 Adjusted Rand index 14
    • 3.2.2 Silhouette coefficient 15
    • 제4장 시뮬레이션 16
    • 4.1 Data Generation 16
    • 4.2 PCA 19
    • 4.3 분석 결과 20
    • 제5장 실제 데이터 24
    • 5.1 데이터 설명 24
    • 5.2 분석 결과 24
    • 제6장 결론 및 향후 연구계획 26
    • 참 고 문 헌 28
    • Abstract 31
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼