RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    문헌간 유사도를 이용한 자동분류에서 미분류 문헌의 활용에 관한 연구 = Utilizing Unlabeled Documents in Automatic Classification with Inter-document Similarities

    한글로보기

    https://www.riss.kr/link?id=A104245553

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This paper studies the problem of classifying documents with labeled and unlabeled learning data, especially with regards to using document similarity features. The problem of using unlabeled data is practically important because in many information systems obtaining training labels is expensive, while large quantities of unlabeled documents are readily available. There are two steps in general semi-supervised learning algorithm. First, it trains a classifier using the available labeled documents, and classifies the unlabeled documents. Then, it trains a new classifier using all the training documents which were labeled either manually or automatically. We suggested two types of semi-supervised learning algorithm with regards to using document similarity features. The one is one step semi-supervised learning which is using unlabeled documents only to generate document similarity features. And the other is two step semi-supervised learning which is using unlabeled documents as learning examples as well as similarity features. Experimental results, obtained using support vector machines and naive Bayes classifier, show that we can get improved performance with small labeled and large unlabeled documents then the performance of supervised learning which uses labeled-only data. When considering the efficiency of a classifier system, the one step semi-supervised learning algorithm which is suggested in this study could be a good solution for improving classification performance with unlabeled documents.
    번역하기

    This paper studies the problem of classifying documents with labeled and unlabeled learning data, especially with regards to using document similarity features. The problem of using unlabeled data is practically important because in many information s...

    This paper studies the problem of classifying documents with labeled and unlabeled learning data, especially with regards to using document similarity features. The problem of using unlabeled data is practically important because in many information systems obtaining training labels is expensive, while large quantities of unlabeled documents are readily available. There are two steps in general semi-supervised learning algorithm. First, it trains a classifier using the available labeled documents, and classifies the unlabeled documents. Then, it trains a new classifier using all the training documents which were labeled either manually or automatically. We suggested two types of semi-supervised learning algorithm with regards to using document similarity features. The one is one step semi-supervised learning which is using unlabeled documents only to generate document similarity features. And the other is two step semi-supervised learning which is using unlabeled documents as learning examples as well as similarity features. Experimental results, obtained using support vector machines and naive Bayes classifier, show that we can get improved performance with small labeled and large unlabeled documents then the performance of supervised learning which uses labeled-only data. When considering the efficiency of a classifier system, the one step semi-supervised learning algorithm which is suggested in this study could be a good solution for improving classification performance with unlabeled documents.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    문헌간 유사도를 자질로 사용하는 분류기에서 미분류 문헌을 학습에 활용하여 분류 성능을 높이는 방안을 모색해보았다. 자동분류를 위해서 다량의 학습문헌을 수작업으로 확보하는 것은 많은 비용이 들기 때문에 미분류 문헌의 활용은 실용적인 면에서 중요하다. 미분류 문헌을 활용하는 준지도학습 알고리즘은 대부분 수작업으로 분류된 문헌을 학습데이터로 삼아서 미분류 문헌을 분류하는 첫 번째 단계와, 수작업으로 분류된 문헌과 자동으로 분류된 문헌을 모두 학습 데이터로 삼아서 분류기를 학습시키는 두 번째 단계로 구성된다. 이 논문에서는 문헌간 유사도 자질을 적용하는 상황을 고려하여 두 가지 준지도학습 알고리즘을 검토하였다. 이중에서 1단계 준지도학습 방식은 미분류 문헌을 문헌유사도 자질 생성에만 활용하므로 간단하며, 2단계 준지도학습 방식은 미분류 문헌을 문헌유사도 자질 생성과 함께 학습 예제로도 활용하는 알고리즘이다. 지지벡터기계와 나이브베이즈 분류기를 이용한 실험 결과, 두 가지 준지도학습 방식 모두 미분류 문헌을 활용하지 않는 지도학습 방식보다 높은 성능을 보이는 것으로 나타났다. 특히 실행효율을 고려한다면 제안된 1단계 준지도학습 방식이 미분류 문헌을 활용하여 분류 성능을 높일 수 있는 좋은 방안이라는 결론을 얻었다
    번역하기

    문헌간 유사도를 자질로 사용하는 분류기에서 미분류 문헌을 학습에 활용하여 분류 성능을 높이는 방안을 모색해보았다. 자동분류를 위해서 다량의 학습문헌을 수작업으로 확보하는 것은 ...

    문헌간 유사도를 자질로 사용하는 분류기에서 미분류 문헌을 학습에 활용하여 분류 성능을 높이는 방안을 모색해보았다. 자동분류를 위해서 다량의 학습문헌을 수작업으로 확보하는 것은 많은 비용이 들기 때문에 미분류 문헌의 활용은 실용적인 면에서 중요하다. 미분류 문헌을 활용하는 준지도학습 알고리즘은 대부분 수작업으로 분류된 문헌을 학습데이터로 삼아서 미분류 문헌을 분류하는 첫 번째 단계와, 수작업으로 분류된 문헌과 자동으로 분류된 문헌을 모두 학습 데이터로 삼아서 분류기를 학습시키는 두 번째 단계로 구성된다. 이 논문에서는 문헌간 유사도 자질을 적용하는 상황을 고려하여 두 가지 준지도학습 알고리즘을 검토하였다. 이중에서 1단계 준지도학습 방식은 미분류 문헌을 문헌유사도 자질 생성에만 활용하므로 간단하며, 2단계 준지도학습 방식은 미분류 문헌을 문헌유사도 자질 생성과 함께 학습 예제로도 활용하는 알고리즘이다. 지지벡터기계와 나이브베이즈 분류기를 이용한 실험 결과, 두 가지 준지도학습 방식 모두 미분류 문헌을 활용하지 않는 지도학습 방식보다 높은 성능을 보이는 것으로 나타났다. 특히 실행효율을 고려한다면 제안된 1단계 준지도학습 방식이 미분류 문헌을 활용하여 분류 성능을 높일 수 있는 좋은 방안이라는 결론을 얻었다

    더보기

    참고문헌 (Reference)

    1 "한국어 테스트 컬렉션 HANTEC의 확장 및 보완" 210-215, 2000

    2 "정보검색연구" 서울: 구미무역(주) 출판부 2005

    3 이재윤, "자질 선정 기준과 가중치 할당 방식간의 관계를 고려한 문서 자동분류의 개선에 대한 연구" 한국문헌정보학회 39 (39): 123-146, 2005

    4 이재윤, "문헌간 유사도를 이용한 SVM 분류기의 문헌분류성능 향상에 관한 연구" 한국정보관리학회 22 (22): 261-287, 2005

    5 김판준, "기계학습을 통한 디스크립터 자동부여에 관한 연구" 한국정보관리학회 23 (23): 279-299, 2006

    6 "and R. C. Dubes. 1988. Algorithms for Clustering Data. Englewood Cliffs" Prentice-Hall.

    7 "Transductive inference for text classification using Support Vector Machines" 200-209, 1999

    8 "The value of unlabeled data for classification problems" 2000

    9 "Text classification from positive and unlabeled examples" 2002

    10 "Text classification from positive and unlabeled documents" 232-239, 2003

    1 "한국어 테스트 컬렉션 HANTEC의 확장 및 보완" 210-215, 2000

    2 "정보검색연구" 서울: 구미무역(주) 출판부 2005

    3 이재윤, "자질 선정 기준과 가중치 할당 방식간의 관계를 고려한 문서 자동분류의 개선에 대한 연구" 한국문헌정보학회 39 (39): 123-146, 2005

    4 이재윤, "문헌간 유사도를 이용한 SVM 분류기의 문헌분류성능 향상에 관한 연구" 한국정보관리학회 22 (22): 261-287, 2005

    5 김판준, "기계학습을 통한 디스크립터 자동부여에 관한 연구" 한국정보관리학회 23 (23): 279-299, 2006

    6 "and R. C. Dubes. 1988. Algorithms for Clustering Data. Englewood Cliffs" Prentice-Hall.

    7 "Transductive inference for text classification using Support Vector Machines" 200-209, 1999

    8 "The value of unlabeled data for classification problems" 2000

    9 "Text classification from positive and unlabeled examples" 2002

    10 "Text classification from positive and unlabeled documents" 232-239, 2003

    11 "Text classification from labeled and unlabeled documents using EM" 39 (39): 103-134, 2000

    12 "Semi-supervised clustering with user feedback" 2003

    13 "Semi-supervised clustering by seeding" 19-26, 2002

    14 "PAC learning from positive statistical queries" 112-126, 1998

    15 "Labeled and unlabeled data in text categorization" 2971-2976, 2004

    16 "Exploiting relations among concepts to acquire weakly labeled training data" 43-50, 2002

    17 "Enhancing supervised learning with unlabeled data" 327-334, 2000

    18 "Employing EM and pool-based active learning with keywords, EM and shrinkage" 359-367, 1998

    19 "Data Mining: Practical Machine Learning Tools and Techniques" San Francisco: Morgan Kaufmann 2005

    20 "Constrained k-means clustering with background knowledge" 577-584, 2001

    21 "Combining labeled and unlabeled data with co-training" 92-100, 1998

    22 "Combining labeled and unlabeled data for multiclass text categorization" 187-194, 2002

    23 "Co-trained support vector machines for large scale unstructured document classification using unlabeled data and syntactic information" 40 (40): 421-439, 2004

    24 "Building text classifiers using positive and unlabeled examples" 179-188, 2003

    25 "Analyzing the effectiveness and applicability of co-training" 86-93, 2000

    26 "Active + semi-supervised learning = robust multi-view learning" 435-442, 2002

    27 "A sequential algorithm for training text classifiers. Proceedings of the 17th Annual International ACM SIGIR Conference on Research and Development in Information Retrieval" 3-12,

    28 "A semi supervised support vector machines" 368-374, 1998

    29 "A probabilistic framework for semi-supervised clustering" 59-68, 2004

    30 "A fast algorithm for automatic classification. Journal of Library Automation" 31-48,

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2026 평가 재인증평가 신청대상 (재인증)
    2020-01-01 등재 등재학술지 유지 (재인증) KCI등재
    2017-01-01 등재 등재학술지 유지 (계속평가) KCI등재
    2013-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2010-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2008-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2006-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2004-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2001-01-01 등재 등재학술지 선정 (등재후보2차) KCI등재
    1998-07-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 1.21 1.21 1.48
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    1.29 1.2 2.027 0.28
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼