논문, 기사, SNS에서 키워드는 주제 요약과 검색 효율성을 결정짓는 핵심 요소이나, 지금까지 키워드 추출 연구들은 추출된 키워드의 정확도 검증에 집중하며, 대다수가 상위 5개, 상위 10개와...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
논문, 기사, SNS에서 키워드는 주제 요약과 검색 효율성을 결정짓는 핵심 요소이나, 지금까지 키워드 추출 연구들은 추출된 키워드의 정확도 검증에 집중하며, 대다수가 상위 5개, 상위 10개와...
논문, 기사, SNS에서 키워드는 주제 요약과 검색 효율성을 결정짓는 핵심 요소이나, 지금까지 키워드 추출 연구들은 추출된 키워드의 정확도 검증에 집중하며, 대다수가 상위 5개, 상위 10개와 같이 임의의 고정된 키워드 개수 기준으로 성능을 비교한다. 이러한 방법은 정확한 성능 비교가 불확실하여 키워드 추출 방법의 정확도 검증을 비교하기 위하여 논문, 기사, SNS에서의 정보를 이용하여 키워드 개수를 정해야 한다.
본 연구는 이러한 문제의식에서 출발하여 논문 초록의 특성을 활용한 키워드 개수 예측 모델을 구축하는 것을 목표로 한다.
연구 방법론은 크게 두 가지 축으로 구성된다. TextRank 키워드 추출 알고리즘의 점수를 기반으로 키워드 개수를 예측하기 위해 키워드 개수 예측 방법을 4가지 제안하고, 초록 단어의 개수를 독립변수로, 실제 저자가 부여한 키워드 개수를 종속변수로 하는 순서형 회귀분석 모델의 연결함수 Logit, Complementary Log-log, Negative Log-log, Probit, Cauchit 으로 예측한 키워드 개수 예측 성능을 비교하였다.
다국어 초록 (Multilingual Abstract)
Keywords are essential elements in academic papers, determining the topic summary and search efficiency. However, existing research on automatic keyword extraction predominantly focuses on validating the accuracy of extracted keywords, with the majori...
Keywords are essential elements in academic papers, determining the topic summary and search efficiency. However, existing research on automatic keyword extraction predominantly focuses on validating the accuracy of extracted keywords, with the majority comparing performance based on an arbitrary, fixed number of keywords, such as the top 5 or 10. This fixed approach fails to reflect the reality that the necessary number of keywords varies depending on the document’s length and content complexity. While the standard for setting the appropriate number of keywords remains ambiguous, a fixed number must be determined to compare keyword accuracy validation.
This study originates from this problem statement and aims to develop a keyword count prediction model that utilizes the characteristics of the paper’s abstract. The research methodology is broadly structured around two axes. First, four different criteria for the keyword count are proposed to analyze the number of keywords based on the scores from the TextRank keyword extraction algorithm. Second, the study predicts the number of keywords using an Ordinal Regression Model, where the word count of the abstract is the independent variable and the actual number of author-assigned keywords is the dependent variable. Various link functions, including Logit, Complementary Log-log, Negative Log-log, Probit, and Cauchit, were applied and compared.
목차 (Table of Contents)