인터넷 기술이 발전함에 따라 온라인상의 데이터는 급격하게 증가하고 있고, 증가하는 데이터에 대해 점진적인 기계학습 기법을 통해 효율적으로 학습하기 위한 연구가 진행되고 있다. 온...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A102679709
2017
Korean
SVM ; 나이브베이즈 ; 시계열분석 ; 기계학습 ; 분류 ; Naive Bayes ; time-Series Analysis ; Machine Learning ; Classification
310
KCI등재
학술저널
39-49(11쪽)
0
0
상세조회0
다운로드인터넷 기술이 발전함에 따라 온라인상의 데이터는 급격하게 증가하고 있고, 증가하는 데이터에 대해 점진적인 기계학습 기법을 통해 효율적으로 학습하기 위한 연구가 진행되고 있다. 온...
인터넷 기술이 발전함에 따라 온라인상의 데이터는 급격하게 증가하고 있고, 증가하는 데이터에 대해 점진적인 기계학습 기법을 통해 효율적으로 학습하기 위한 연구가 진행되고 있다. 온라인상의 문서는 대부분 게시일, 출판일과 같은 시계열적 정보를 포함하고 있고, 이를 분류에 반영한다면 효율적인 분류가 가능할 것이다. 본 연구에서는 웹 문서상에서 나타나는 어휘의 시계열적 변화를 분석하였고, 분석한 시계열 정보를 기반으로 데이터 집합을 분할하여 효율적인 분류 학습 기법을 제안한다. 실험 및 검증을 위해 온라인상의 뉴스 기사 100만 건을 시계열 정보를 포함하여 수집하였다. 수집된 데이터를 바탕으로 데이터 집합을 분할하여 Naïve Bayes 및 SVM 분류기를 사용하여 실험을 진행하였고, 각 모델에서 전체 데이터 집합 학습 대비 최대 2.02% 포인트, 2.32% 포인트의 성능 향상을 확인하였다. 본 연구를 통해 시계열적 어휘의 변화를 분류에 반영하여 분류의 성능을 향상시킬 수 있음을 확인하였다.
다국어 초록 (Multilingual Abstract)
As the Internet technology advances, data on the web is increasing sharply. Many research study about incremental learning for classifying effectively in data increasing. Web document contains the time-series data such as published date. If we reflect...
As the Internet technology advances, data on the web is increasing sharply. Many research study about incremental learning for classifying effectively in data increasing. Web document contains the time-series data such as published date. If we reflect time-series data to classification, it will be an effective classification. In this study, we analyze the time-series variation of the words. We propose an efficient classification through dividing the dataset based on the analysis of time-series information. For experiment, we corrected 1 million online news articles including time-series information. We divide the dataset and classify the dataset using SVM and Naïve Bayes. In each model, we show that classification performance is increasing Through this study, we showed that reflecting, time-series information can improve the classification performance.
목차 (Table of Contents)
참고문헌 (Reference)
1 정도헌, "최대 개념강도 인지기법을 이용한 데이터베이스 자동선택 방법에 관한 연구" 한국정보관리학회 27 (27): 265-281, 2010
2 정도헌, "빅데이터 마이닝을 위한 점진적 학습 기술 개발" KISTI 2015
3 "https://nodejs.org"
4 "http://www.highcharts.com"
5 "http://visjs.org"
6 Derry Tanti Wijaya, "Understanding Semantic Change of Words Over Centuries" 2011
7 Do-Heon Jeong, "Time gap analysis by he topic model-based temporal technique" 8 (8): 776-790, 2014
8 C. Cortes, "Support-Vector Net –works" 20 (20): 273-297, 1995
9 B. Croft, "Machine Learning and Information Retrieval" 1995
10 E. Jessica, "Forecast : Mobile Data Traffic, Worldwide, 2011-2018" Gartner 2015
1 정도헌, "최대 개념강도 인지기법을 이용한 데이터베이스 자동선택 방법에 관한 연구" 한국정보관리학회 27 (27): 265-281, 2010
2 정도헌, "빅데이터 마이닝을 위한 점진적 학습 기술 개발" KISTI 2015
3 "https://nodejs.org"
4 "http://www.highcharts.com"
5 "http://visjs.org"
6 Derry Tanti Wijaya, "Understanding Semantic Change of Words Over Centuries" 2011
7 Do-Heon Jeong, "Time gap analysis by he topic model-based temporal technique" 8 (8): 776-790, 2014
8 C. Cortes, "Support-Vector Net –works" 20 (20): 273-297, 1995
9 B. Croft, "Machine Learning and Information Retrieval" 1995
10 E. Jessica, "Forecast : Mobile Data Traffic, Worldwide, 2011-2018" Gartner 2015
11 H. Taira, "Feature selection in SVM text categorization" 1999
12 F. Colas, "Comparison of SVM and some older classification algorithms in text classification tasks" 2006
13 D. Jeong, "Classification Method by Integrating Feature Property Matrices for Large Scale Data" SMA 2012
14 Pascal Soucy, "Beyond TF-IDF Weighting for Text Categorization in the Vector Space Model" 5 : 1130-1135, 2005
15 G. Forman, "BNS Feature Scaling: An Improved Representation over TF·IDF for SVM Text Classification" ACM 2008
16 J. Gim, "Anayzing Email Patterns with Timelines on Researcher Data" 2014
17 Irina Rish, "An empirical study of the naive Bayes classifier" IBM 2001
18 H. Chih, "An empirical study of feature selection for text categorization based on term weightage" 599-602, 2004
19 Saket S. R. Mengle, "Ambiguity Measure Feature-Selection Algorithm" 60 (60): 1037-1050, 2009
20 B. E. Boser, "A training algorithm for optimal margin classifiers" 1992
21 Yiming Yang, "A comparative study on feature selection in text categorization" 97 : 412-420, 1997
22 A. McCallum, "A Comparison of Event Models for Naive Bayes Text Classification" 1998
학술지 이력
| 연월일 | 이력구분 | 이력상세 | 등재구분 |
|---|---|---|---|
| 2027 | 평가 | 재인증평가 신청대상 (재인증) | |
| 2021-01-01 | 등재 | 등재학술지 유지 (재인증) | ![]() |
| 2018-01-01 | 등재 | 등재학술지 유지 (등재유지) | ![]() |
| 2015-01-01 | 등재 | 등재학술지 유지 (등재유지) | ![]() |
| 2011-01-01 | 등재 | 등재학술지 유지 (등재유지) | ![]() |
| 2008-01-01 | 등재 | 등재학술지 선정 (등재후보2차) | ![]() |
| 2007-05-04 | 학회명변경 | 영문명 : The Korea Contents Society -> The Korea Contents Association | ![]() |
| 2007-01-01 | 등재 | 등재후보 1차 PASS (등재후보1차) | ![]() |
| 2006-01-01 | 등재 | 등재후보학술지 유지 (등재후보1차) | ![]() |
| 2004-01-01 | 등재 | 등재후보학술지 선정 (신규평가) | ![]() |
학술지 인용정보
| 기준연도 | WOS-KCI 통합IF(2년) | KCIF(2년) | KCIF(3년) |
|---|---|---|---|
| 2016 | 1.21 | 1.21 | 1.26 |
| KCIF(4년) | KCIF(5년) | 중심성지수(3년) | 즉시성지수 |
| 1.29 | 1.25 | 1.573 | 0.33 |