RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI우수등재

    기술과학 분야 문헌의 DDC 자동 분류에 관한 연구 = A Study on Automatic DDC Classification of Documents in Technology

    한글로보기

    https://www.riss.kr/link?id=A110183915

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study investigates the automatic classification of documents in the Dewey Decimal Classification (DDC) Technology class (600) using machine learning models, with the aim of overcoming the limitations of title-based classification approaches. To enhance classification performance, descriptive document information, such as summaries and introductions, was incorporated as additional classification features. Three machine learning models—Omikuji, FastText, and BERT—were employed, and classification performance was evaluated at both the main class and division levels. Accuracy and F1-score were used as evaluation metrics. The results demonstrate that BERT consistently outperformed FastText and Omikuji across most experimental conditions. With the exception of the division-level F1-score of the Omikuji model, all models showed improved performance when descriptive information was added. In particular, the BERT-based model achieved an accuracy of 79.52% at the division level, representing an improvement of approximately 8.62 percentage points compared to previous studies. The findings also indicate that classification performance generally improves as the volume of documents used in model training increases, underscoring the importance of data scale in addition to feature selection. These results suggest that competitive automatic classification performance can be achieved through appropriate model selection and enriched classification features, even within single-model approaches. Future research should expand the scope to all DDC classes and examine the applicability of the proposed approach to the Korean Decimal Classification (KDC), as well as explore additional features and alternative machine learning models.
    번역하기

    This study investigates the automatic classification of documents in the Dewey Decimal Classification (DDC) Technology class (600) using machine learning models, with the aim of overcoming the limitations of title-based classification approaches. To e...

    This study investigates the automatic classification of documents in the Dewey Decimal Classification (DDC) Technology class (600) using machine learning models, with the aim of overcoming the limitations of title-based classification approaches. To enhance classification performance, descriptive document information, such as summaries and introductions, was incorporated as additional classification features. Three machine learning models—Omikuji, FastText, and BERT—were employed, and classification performance was evaluated at both the main class and division levels. Accuracy and F1-score were used as evaluation metrics. The results demonstrate that BERT consistently outperformed FastText and Omikuji across most experimental conditions. With the exception of the division-level F1-score of the Omikuji model, all models showed improved performance when descriptive information was added. In particular, the BERT-based model achieved an accuracy of 79.52% at the division level, representing an improvement of approximately 8.62 percentage points compared to previous studies. The findings also indicate that classification performance generally improves as the volume of documents used in model training increases, underscoring the importance of data scale in addition to feature selection. These results suggest that competitive automatic classification performance can be achieved through appropriate model selection and enriched classification features, even within single-model approaches. Future research should expand the scope to all DDC classes and examine the applicability of the proposed approach to the Korean Decimal Classification (KDC), as well as explore additional features and alternative machine learning models.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼