RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    효율적인 문헌 분류를 위한 시계열 기반 데이터 집합 선정 기법

    한글로보기

    https://www.riss.kr/link?id=A102679709

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    인터넷 기술이 발전함에 따라 온라인상의 데이터는 급격하게 증가하고 있고, 증가하는 데이터에 대해 점진적인 기계학습 기법을 통해 효율적으로 학습하기 위한 연구가 진행되고 있다. 온라인상의 문서는 대부분 게시일, 출판일과 같은 시계열적 정보를 포함하고 있고, 이를 분류에 반영한다면 효율적인 분류가 가능할 것이다. 본 연구에서는 웹 문서상에서 나타나는 어휘의 시계열적 변화를 분석하였고, 분석한 시계열 정보를 기반으로 데이터 집합을 분할하여 효율적인 분류 학습 기법을 제안한다. 실험 및 검증을 위해 온라인상의 뉴스 기사 100만 건을 시계열 정보를 포함하여 수집하였다. 수집된 데이터를 바탕으로 데이터 집합을 분할하여 Naïve Bayes 및 SVM 분류기를 사용하여 실험을 진행하였고, 각 모델에서 전체 데이터 집합 학습 대비 최대 2.02% 포인트, 2.32% 포인트의 성능 향상을 확인하였다. 본 연구를 통해 시계열적 어휘의 변화를 분류에 반영하여 분류의 성능을 향상시킬 수 있음을 확인하였다.
    번역하기

    인터넷 기술이 발전함에 따라 온라인상의 데이터는 급격하게 증가하고 있고, 증가하는 데이터에 대해 점진적인 기계학습 기법을 통해 효율적으로 학습하기 위한 연구가 진행되고 있다. 온...

    인터넷 기술이 발전함에 따라 온라인상의 데이터는 급격하게 증가하고 있고, 증가하는 데이터에 대해 점진적인 기계학습 기법을 통해 효율적으로 학습하기 위한 연구가 진행되고 있다. 온라인상의 문서는 대부분 게시일, 출판일과 같은 시계열적 정보를 포함하고 있고, 이를 분류에 반영한다면 효율적인 분류가 가능할 것이다. 본 연구에서는 웹 문서상에서 나타나는 어휘의 시계열적 변화를 분석하였고, 분석한 시계열 정보를 기반으로 데이터 집합을 분할하여 효율적인 분류 학습 기법을 제안한다. 실험 및 검증을 위해 온라인상의 뉴스 기사 100만 건을 시계열 정보를 포함하여 수집하였다. 수집된 데이터를 바탕으로 데이터 집합을 분할하여 Naïve Bayes 및 SVM 분류기를 사용하여 실험을 진행하였고, 각 모델에서 전체 데이터 집합 학습 대비 최대 2.02% 포인트, 2.32% 포인트의 성능 향상을 확인하였다. 본 연구를 통해 시계열적 어휘의 변화를 분류에 반영하여 분류의 성능을 향상시킬 수 있음을 확인하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As the Internet technology advances, data on the web is increasing sharply. Many research study about incremental learning for classifying effectively in data increasing. Web document contains the time-series data such as published date. If we reflect time-series data to classification, it will be an effective classification. In this study, we analyze the time-series variation of the words. We propose an efficient classification through dividing the dataset based on the analysis of time-series information. For experiment, we corrected 1 million online news articles including time-series information. We divide the dataset and classify the dataset using SVM and Naïve Bayes. In each model, we show that classification performance is increasing Through this study, we showed that reflecting, time-series information can improve the classification performance.
    번역하기

    As the Internet technology advances, data on the web is increasing sharply. Many research study about incremental learning for classifying effectively in data increasing. Web document contains the time-series data such as published date. If we reflect...

    As the Internet technology advances, data on the web is increasing sharply. Many research study about incremental learning for classifying effectively in data increasing. Web document contains the time-series data such as published date. If we reflect time-series data to classification, it will be an effective classification. In this study, we analyze the time-series variation of the words. We propose an efficient classification through dividing the dataset based on the analysis of time-series information. For experiment, we corrected 1 million online news articles including time-series information. We divide the dataset and classify the dataset using SVM and Naïve Bayes. In each model, we show that classification performance is increasing Through this study, we showed that reflecting, time-series information can improve the classification performance.

    더보기

    목차 (Table of Contents)

    • 요약
    • Abstract
    • Ⅰ. 서론
    • Ⅱ. 관련연구 및 연구배경
    • Ⅲ. 데이터 집합 선정 모델
    • 요약
    • Abstract
    • Ⅰ. 서론
    • Ⅱ. 관련연구 및 연구배경
    • Ⅲ. 데이터 집합 선정 모델
    • Ⅳ. 성능평가
    • Ⅴ. 결론
    • 참고문헌
    더보기

    참고문헌 (Reference)

    1 정도헌, "최대 개념강도 인지기법을 이용한 데이터베이스 자동선택 방법에 관한 연구" 한국정보관리학회 27 (27): 265-281, 2010

    2 정도헌, "빅데이터 마이닝을 위한 점진적 학습 기술 개발" KISTI 2015

    3 "https://nodejs.org"

    4 "http://www.highcharts.com"

    5 "http://visjs.org"

    6 Derry Tanti Wijaya, "Understanding Semantic Change of Words Over Centuries" 2011

    7 Do-Heon Jeong, "Time gap analysis by he topic model-based temporal technique" 8 (8): 776-790, 2014

    8 C. Cortes, "Support-Vector Net –works" 20 (20): 273-297, 1995

    9 B. Croft, "Machine Learning and Information Retrieval" 1995

    10 E. Jessica, "Forecast : Mobile Data Traffic, Worldwide, 2011-2018" Gartner 2015

    1 정도헌, "최대 개념강도 인지기법을 이용한 데이터베이스 자동선택 방법에 관한 연구" 한국정보관리학회 27 (27): 265-281, 2010

    2 정도헌, "빅데이터 마이닝을 위한 점진적 학습 기술 개발" KISTI 2015

    3 "https://nodejs.org"

    4 "http://www.highcharts.com"

    5 "http://visjs.org"

    6 Derry Tanti Wijaya, "Understanding Semantic Change of Words Over Centuries" 2011

    7 Do-Heon Jeong, "Time gap analysis by he topic model-based temporal technique" 8 (8): 776-790, 2014

    8 C. Cortes, "Support-Vector Net –works" 20 (20): 273-297, 1995

    9 B. Croft, "Machine Learning and Information Retrieval" 1995

    10 E. Jessica, "Forecast : Mobile Data Traffic, Worldwide, 2011-2018" Gartner 2015

    11 H. Taira, "Feature selection in SVM text categorization" 1999

    12 F. Colas, "Comparison of SVM and some older classification algorithms in text classification tasks" 2006

    13 D. Jeong, "Classification Method by Integrating Feature Property Matrices for Large Scale Data" SMA 2012

    14 Pascal Soucy, "Beyond TF-IDF Weighting for Text Categorization in the Vector Space Model" 5 : 1130-1135, 2005

    15 G. Forman, "BNS Feature Scaling: An Improved Representation over TF·IDF for SVM Text Classification" ACM 2008

    16 J. Gim, "Anayzing Email Patterns with Timelines on Researcher Data" 2014

    17 Irina Rish, "An empirical study of the naive Bayes classifier" IBM 2001

    18 H. Chih, "An empirical study of feature selection for text categorization based on term weightage" 599-602, 2004

    19 Saket S. R. Mengle, "Ambiguity Measure Feature-Selection Algorithm" 60 (60): 1037-1050, 2009

    20 B. E. Boser, "A training algorithm for optimal margin classifiers" 1992

    21 Yiming Yang, "A comparative study on feature selection in text categorization" 97 : 412-420, 1997

    22 A. McCallum, "A Comparison of Event Models for Naive Bayes Text Classification" 1998

    더보기

    동일학술지(권/호) 다른 논문

    동일학술지 더보기

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2027 평가 재인증평가 신청대상 (재인증)
    2021-01-01 등재 등재학술지 유지 (재인증) KCI등재
    2018-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2015-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2011-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2008-01-01 등재 등재학술지 선정 (등재후보2차) KCI등재
    2007-05-04 학회명변경 영문명 : The Korea Contents Society -> The Korea Contents Association KCI등재후보
    2007-01-01 등재 등재후보 1차 PASS (등재후보1차) KCI등재후보
    2006-01-01 등재 등재후보학술지 유지 (등재후보1차) KCI등재후보
    2004-01-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 1.21 1.21 1.26
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    1.29 1.25 1.573 0.33
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼