RISS 학술연구정보서비스

검색
다국어 입력

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

변환된 중국어를 복사하여 사용하시면 됩니다.

예시)
  • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
  • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
닫기
    인기검색어 순위 펼치기

    RISS 인기검색어

      검색결과 좁혀 보기

      선택해제
      • 좁혀본 항목 보기순서

        • 원문유무
        • 음성지원유무
        • 학위유형
        • 주제분류
          펼치기
        • 수여기관
          펼치기
        • 발행연도
          펼치기
        • 작성언어
        • 지도교수
          펼치기

      오늘 본 자료

      • 오늘 본 자료가 없습니다.
      더보기
      • BERT를 활용한 ESG 정성적 요소 분석을 통한 ESG 등급 검증 방안 : 언론보도 데이터 기반으로

        강철원 서울과학기술대학교 2024 국내석사

        RANK : 247807

        제 목 : BERT를 활용한 ESG 정성적 요소 분석을 통한 ESG 등급 검증 방안 : 언론보도 데이터 기반으로 ESG는 환경, 사회, 지배구조를 평가하는 개념으로 기업의 지속가능경영을 평가 하는 중요한 지표로 자리 잡고 있다. 기존의 ESG 평가는 주로 기업의 재무적 수치와 ESG 관련 정량적 정보에 의존했지만, 최근에는 정량적 평가 외에 빅데이터를 활용한 ESG 정성적 평가의 중요성이 강조되고 있다. 언론보도 데이터는 기업의 환경, 사회, 지배구조에 관련된 ESG 정보를 담고 있어 ESG 평가의 중요한 자료로 활용할 수 있다. 특히, 언론보도 데이터는 기업의 미시적인 정보까지 다각적으로 제공하여 더욱 정확한 ESG 평가를 가능하게 한다. 뿐만 아니라 BERT와 같은 LLM을 활용한 자연어 처리 기술의 발전으로 인해, 언론보도 데이터를 효과적으로 분석하여 기업의 ESG 등급을 예측하는 연구가 가능하다. 본 연구는 기업의 ESG 등급을 예측하기 위한 방법론을 제안하며, 이를 위해 언론보도 데이터와 자연어 처리를 위해 BERT를 활용하였다. 언론보도에는 기업의 ESG에 영향을 미치는 다양한 정보가 포함되어 있으며, 이를 기반으로 기업의 ESG 성과에 대한 정성적인 특성을 분석한다. BERT는 언어의 문맥을 이해하는 데 뛰어난 성능을 가진 LLM로써, 언론보도 데이터를 활용하여 기업의 ESG 요소 분류와 감성 분석을 진행하고, 이를 기반으로 기업의 E, S, G 정성적 요소 지수를 산출하는 데 활용한다. 이를 위해 검색 플랫폼 Google에서 BERT 모델의 학습용 언론보도 데이터 140,004건과 언론보도 플랫폼 빅카인즈에서 기업별 E, S, G 정성적 요소 지수 산출을 위한 언론보도 데이터 334,161건을 수집하고, 2021년부터 2022년 까지의 ESG 등급 데이터를 수집하여 연구에 활용하였다. ESG 요소 분류에서는 다양한 BERT 모델(Multilingual BERT, KoBERT, KoELECTRA-base-v3, KPF-BERT) 및 BERT 기반의 Ensemble을 사용하여 성능을 비교한 결과 BERT Ensemble이 Accuracy는 98.28%, F1-score는 98.21%, Precision은 98.27%, Recall은 98.15%로 가장 높은 성능을 보였고, 감성 분석에서도 BERT Ensemble이 Accuracy는 99.64%, F1-score는 99.62%, Precision은 99.63%, Recall은 99.62%로 가장 높은 성능을 나타냈다. 마지막으로 ESG 등급 예측에서는 2021년 E, S, G 등급 총 3개의 독립변수를 활용한 예측모형과 2021년 E 등급, S 등급, G 등급, E 정성적 요소 지수, S 정성적 요소 지수, G 정성적 요소 지수 총 6개의 독립변수를 활용한 예측모형의 성능 비교에서 Accuracy는 28.94%, F1-score는 38.22%, Precision은 32.94%, Recall은 31.35%로 큰 폭으로 성능 향상을 보였다. 이를 통해 E, S, G 정성적 요소 지수가 ESG 등급 예측에 기여하는 것으로 확인하였다. 이러한 결과는 기업의 재무적인 수치 외에도 언론보도 데이터를 활용하여 ESG 정성적 요소 분석을 통해 기업의 ESG 평가가 가능할 것으로 기대된다. ESG rating verification method through ESG qualitative element analysis using BERT : Based on media report data ESG is a concept that evaluates the environmental, social, and governance structure, and has become an important indicator for evaluating a company's sustainable management. Existing ESG evaluations mainly relied on company financial figures and ESG-related quantitative information, but recently, in addition to quantitative evaluations, the importance of qualitative ESG evaluations using big data has been emphasized. Media report data contains ESG information related to the company's environmental, social, and governance structure and can be used as important data for ESG evaluation. In particular, media coverage data provides diverse microscopic information about companies, enabling more accurate ESG evaluation. In addition, due to the development of natural language processing technology using LLM such as BERT, it is possible to conduct research to predict a company's ESG rating by effectively analyzing media coverage data. This study proposes a methodology to predict a company's ESG rating, and for this purpose, BERT was used for media coverage data and natural language processing. Media reports contain a variety of information that affects a company's ESG, and based on this, qualitative characteristics of a company's ESG performance are analyzed. BERT is an LLM with excellent performance in understanding the context of language. It uses media report data to classify and analyze corporate ESG factors and calculates the company's E, S, G qualitative factor index based on this. Use it to To this end, we collected 140,004 pieces of media report data for training of the BERT model from search platform Google and 334,161 pieces of media report data for calculating E, S, G qualitative factor indices for each company from media report platform Bigkinds, and collected data from 2021 to 2022. ESG rating data up to was collected and used for research. In ESG factor classification, the performance was compared using various BERT models (Multilingual BERT, KoBERT, KoELECTRA-base-v3, KPF-BERT) and BERT-based Ensemble, and the BERT Ensemble achieved an Accuracy of 98.28% and an F1-score of 98.21%, Precision showed the highest performance at 98.27% and Recall at 98.15%, and in sentiment analysis, BERT Ensemble showed the highest performance at Accuracy at 99.64%, F1-score at 99.62%, Precision at 99.63%, and Recall at 99.62%. Lastly, in ESG rating prediction, the prediction model using a total of three independent variables in E, S, and G grades in 2021 and in the 2021 performance comparison of the predictive model using a total of six independent variables: E, S, G, E qualitative factor index, S qualitative factor index, and G qualitative factor index, Accuracy showed a significant improvement in performance at 28.94%, F1-score at 38.22%, Precision at 32.94%, and Recall at 31.35%. Through this, it was confirmed that the E, S, and G qualitative factor indices contribute to predicting ESG ratings. These results are expected to make it possible to evaluate a company's ESG through qualitative ESG element analysis using media coverage data in addition to the company's financial figures.

      • 소셜 미디어 데이터의 감정 분석 : 사전 학습된 BERT를 이용한 이중 미세 조정 접근

        장민지 숙명여자대학교 대학원 2024 국내석사

        RANK : 247807

        With the advent of the 21st century, advancements in the IT industry, alongside the proliferation of smartphones and the internet, have brought significant changes to people's daily lives and consumption of culture. Particularly, platforms like YouTube have disrupted the traditional paradigm of broadcast content production, opening new realms for individual creators and small-scale content production. These shifts have led to a wide consumption of diverse contents alongside the innovation of streaming services and the emergence of OTT (Over-The-Top) platforms, enabling both corporations and individual creators to engage in active marketing and content creation. However, this modern social trend has not only benefitted OTT platforms. Notably, terms like 'poverty in abundance' and 'Netflix Syndrome' have emerged around Netflix, highlighting a phenomenon where the expansion of choices in movies and dramas induces users' decision-making time and mental stress. Additionally, the surge in small-scale content production on social network services like YouTube has increased the tendency of users to refer to relatively shorter review videos than the somewhat lengthy movies or dramas offered on OTT platforms. The popularity of small-scale content production and short review videos reflects users' emotions and preferences. Through emotion analysis, it's possible to understand these trends and establish more effective content production and marketing strategies. This study utilizes the BERT – BASE version of the BERT – Multilingual model to transcribe voices from videos containing YouTube creators' poetic interpretations into subtitles and experiments with emotion analysis through dual fine-tuning without pre-training, using comments from viewers. The emotion analysis execution mechanism is divided into two experiments: the first using comments and the second using subtitles, involving data exploration, preprocessing, model training, and performance evaluation based on accuracy. The study examines the correlation of labels using binary methods based on the emotion analysis of sampled subtitles and comments. For instance, subtitles of a 'cohabitation drama review' video predominantly showed 'jealousy' (Emotion Label, 31), while the comments reflected 'satisfaction' (Emotion Label, 54) and 'excitement' (Emotion Label, 55). This difference is attributed to factors like content and user response disparity, the diversity and subjectivity of emotions, and the dramatization and direction style. This natural variance between the emotions in drama subtitles and viewer comments illustrates how each individual's unique experiences, interpretations, and responses generate diverse emotional reactions. Particularly, performance evaluation results can be qualitatively compared with related studies. Despite the same learning environment, an increase in accuracy by 0.43% was proven, and BERT demonstrated somewhat higher performance through dual fine-tuning without special pre-training. 21세기 들어 IT 산업의 발전과 함께, 스마트폰과 인터넷 보급은 사람들의 일상과 문화 소비 방식에 커다란 변화를 가져왔다. 특히 유튜브와 같은 플랫폼의 등장은 전통적인 방송 콘텐츠 제작 패러다임을 깨고, 개인 크리에이터와 소규모 콘텐츠 제작에 참여할 수 있는 새로운 영역을 열었다. 이러한 변화는 폭넓은 스트리밍 서비스의 혁신과 OTT(Over – The –Top, 이하 생략) 플랫폼의 등장과 함께 다양한 콘텐츠가 소비 됐으며, 이 콘텐츠를 바탕으로 기업과 개인 크리에이터들은 활발한 마케팅 및 창작물을 제공했다. 그러나 이러한 현대의 사회적 흐름은 OTT 플랫폼들에게 이점만 남기진 않았다. 특히, 넷플릭스를 중심으로 ‘풍요 속의 빈곤’과 ‘넷플릭스 증후군’이라는 용어들이 등장함을 통해 영화나 드라마의 선택의 폭을 넓힘으로써 사용자의 선택 시간과 정신적인 스트레스를 유발하는 현상이 나타났다. 더불어, 유튜브와 같은 소셜 네트워크 서비스를 대상으로 소규모 콘텐츠 제작의 열풍을 일으키며, 다소 긴 영화나 드라마를 제공하는 OTT 플랫폼의 영화나 드라마들보다 비교적 짧은 리뷰 영상을 참고하는 사용자의 경향이 높아지고 있다. 이러한 소규모 콘텐츠 제작의 열풍과 짧은 리뷰 영상의 인기는 사용자의 감정과 선호도를 반영하는데, 감정 분석을 통해 이러한 경향을 파악하고 콘텐츠 제작 및 마케팅 전략을 보다 효과적으로 수립할 수 있다. 본 연구에서는 BERT – BASE 버전의 BERT – Multilingual 모델을 활용하여, 유튜브 크리에이터들의 서정적인 해석이 내포된 영상의 음성을 자막으로 텍스트화시키고, 해당 영상을 시청한 사용자들의 감정을 댓글을 활용하여 사전 학습 없이 이중 미세 조정을 통한 감정 분석을 실험한다. 감정 분석 실행 메커니즘은 댓글을 활용한 1차 실험과 자막을 활용한 2차 실험으로 나뉘어, 데이터 탐색, 전처리, 모델 학습, 정확도를 활용한 성능 평가로 이루어진다. 연구 결과는 표본으로 추출한 자막의 감정 분석과 댓글의 감정 분석 결과를 토대로 이진법을 활용해 라벨의 연관성을 살펴보았으며, 대표적인 예로 간 떨어지는 동거 리뷰 영상의 자막은 주로 질투하는(감정라벨, 31번)이 나왔으며, 댓글 감정 분석 결과는 만족하는(감정라벨, 54번)과 흥분한(감정라벨, 55번)이 주로 나타났다. 이는 자막과 같은 감정이 나타나지 않았는데, 주요 큰 요인으로썬 콘텐츠와 사용자 반응의 차이, 감정의 다양성과 주관성, 드라마 표현 방식과 연출 때문이다. 이러한 요인들을 종합해 볼 때, 드라마의 자막 데이터와 시청자 댓글 사이의 감정의 차이가 나타나는 것은 매우 자연스러운 현상이다. 이는 각 개인의 독특한 경험, 해석 및 반응이 어떻게 다양한 감정적 반응을 생성하는지 보여준다. 특히, 성능 평가 결과는 본연구와 관련 연구의 비교를 통해 정성적 성능 평가 결과를 살펴볼 수 있다. 이는 같은 학습 환경임에도 불구하고, 정확도는 0.43% 높음을 입증할 수 있었고, 특별한 사전 학습 없이도, BERT는 이중 미세 조정을 통해 다소 높은 성능을 보여줄 수 있음을 확인한다.

      • Developing zero anaphora resolution system based on deep learning technology

        김영태 Graduate School, Yonsei University 2020 국내박사

        RANK : 247807

        BERT is a general language representation model that enables systems to utilize deep bidirectional contextual information in natural language texts. Good word and phrase embeddings, when used as the underlying input representation, have been shown to boost the performance in language tasks. This is what is demonstrated by BERT. BERT exploits attention mechanism extensively based upon the sequence transduction model Transformer. BERT is one of the most advanced and complex models that makes use of the most recent state-of-the-art techniques in deep learning. It is necessary to achieve high performance in the task of zero anaphora resolution (ZAR) for complete understanding of texts in Korean, Japanese, Chinese, etc. Influenced by success of deep learning, models based on this technology began to be introduced recently in building ZAR systems. However, the objective of building a high quality ZAR system is far from being achieved even by using these models. To overcome an obstacle in improving ZAR performance, we have proposed to exploit BERT in designing a new model for ZAR. This approach has not been taken by others in developing a ZAR system yet. Specifically, we have chosen to use the fine-tuning approach in utilizing a BERT made available after pre-training. To demonstrate the advantages of our proposed approach by performance comparison, we built ZAR systems based on deep learning models without using BERT. We also implemented the ZAR models suggested by other researchers that make use of deep learning techniques. The performance comparisons of our proposed model with these other models have revealed that our proposed model is superior to those of others. We also experimented with various neural network architectures added on top of BERT to develop our ZAR system. It was observed that adding a complex architecture is more advantageous in improving the performance. This is a new finding related to the use of BERT for language tasks. It was also found that a BERT pre-trained solely with Korean corpus is superior to a multi-lingual BERT. We have sought the end-to-end learning paradigm by disallowing any use of hand-crafted features or dependency-analysis features. Experimental results show that the BERT-based models we propose can result in large performance improvement in ZAR over other deep learning models introduced by other researchers. BERT는 시스템이 자연어 텍스트에서 양방향 양방향 컨텍스트 정보를 활용 할 수 있게 하는 일반 언어 표현 모델입니다. 좋은 단어와 구, 절에 대한 임- 베딩을 기본 입력 표현으로 사용할 때 언어 작업의 성능을 향상시키는 것으 로 나타났습니다. 이것이 BERT가 보여주는 것입니다. BERT는 시퀀스 변환 모 델 Transformer를 기반으로 광범위하게 어텐션 메커니즘을 활용합니다. BERT 는 최신 딥러닝 기술을 활용하는 가장 발전된 복잡한 모델 중 하나입니다. 한국어, 일본어, 중국어 등의 문장을 완전히 이해하기 위해서는 ZAR (Zero Anaphora Resolution) 작업에서 높은 성능을 달성해야 합니다. 딥러닝의 성공에 영향을 받아 이 기술을 기반으로 한 모델들이 최근 ZAR 시스템에 도입되기 시작했습니다. ZAR 시스템. 그러나 고품질 ZAR 시스템 구축은 이러한 모델을 사용하더라도 달성 할 수 없습니다. 우리는 ZAR 성능 향상의 장애를 극복하기 위해 ZAR을 위한 새로운 모델 을 설계 할 때 BERT를 활용할 것을 제안했습니다. 이 접근법은 아직 다른 ZAR 시스템 개발에서 채택되지 않았습니다. 특히, 우리는 사전 훈련 후 제공 되는 BERT를 활용하기 위해 미세 조정 방법을 사용하기로 결정했습니다. 성 능 비교를 통한 제안 된 접근 방식의 장점을 보여주기 위해 BERT를 사용하 지 않은 딥러닝 모델을 기반으로 ZAR 시스템을 구축했습니다. 또한 딥러닝 기술을 사용하는 다른 연구자들이 제안한 ZAR 모델도 구현했습니다. 제안 된 모델과 이러한 다른 모델의 성능을 비교 한 결과 제안 된 모델이 다른 모델 보다 우수합니다. 또한 ZAR 시스템을 개발하기 위해 BERT 위에 다양한 신경망 아키텍처를 추가한 실험을 했습니다. 복잡한 아키텍처를 추가하는 것이 성능 향상에 더 유리하다는 것이 관찰되었습니다. 이것은 언어 작업에 BERT를 사용하는 것과 관련된 새로운 발견입니다. 또한 한국어 말뭉치로만 사전 훈련 된 BERT가 다 국어 BERT보다 우수하다는 것도 발견되었습니다. 우리는 수작업으로 만들어진 자질정보나 의존관계 분석 결과를 사용하지 않는 end-to-end 학습 패러다임을 추구했습니다. 실험 결과에 따르면 우리가 제안한 BERT 기반 모델이 다른 연구에서 소개한 딥러닝 모델보다 ZAR의 성 능을 크게 향상시킬 수 있음을 보여줍니다.

      • BERT 기반 부산도시철도 민원 자동분류 모델 : 부산도시철도 민원 자동 분류: BERT 언어 모델 활용

        김영찬 국립부경대학교 대학원 2025 국내석사

        RANK : 247807

        본 논문은 부산교통공사 인터넷 민원 시스템에 문의되는 민원을 분석하고 이를 담당 부서별로 자동 분류하는 시스템 구축에 대한 연구를 수행하였다. 부산도시철도의 2015~2023년도 공개 민원 데이터 3,089건의 전처리 과정을 통해 파인튜닝에 필요한 데이터셋을 분류 및 구축하였다. 분류에 대한 학습과 테스트 실행을 위해 데이터셋의 분류 코드와 레이블을 분류하여, Hugging Face에서 한국어 사전 훈련된 잘 알려진 네 가지 BERT 모델을 이용하여 민원 자동분류 모델을 구축하여 주요 속성을 파악하였다. 네 가지 모델 중 kikim/bert-kor-base은 82.85%의 가장 높은 자동분류 정확도를 보였다. 테스트 데이터셋의 혼동 행렬에서 오분류한 비율이 가장 높은 코드는 승무와 차량이었으며, 이는 관련 부서의 역할과 민원의 내용이 다소 겹쳐 모델이 혼동한 것으로 추정된다. 본 연구를 통해 여러 BERT 모델의 성능을 비교 분석을 통해 BERT 기반 한국어 민원 자동분류 모델 적용 가능성 및 효과를 검증할 수 있었다. 후속 연구에서는 훈련 과정에서 더 많은 언어 데이터와 복잡한 민원 유형을 포함하여 시스템의 일반화 능력을 강화하는 시스템 최적화가 필요하다. This study investigates the development of a system that analyzes customer complaints submitted to the Busan Metro Internet Civil Complaints System and automatically classifies them for assignment to the relevant departments. A dataset for fine-tuning was constructed and classified through a preprocessing process using 3,089 publicly available civil complaint cases from 2015 to 2023 in Busan Metro. For training and testing the classification process, complaint codes and labels in the dataset were categorized, and four well-known pre-trained Korean BERT models from Hugging Face were utilized to build an automatic complaint classification system and analyze key attributes. Among the four models, kikim/bert-kor-base achieved the highest automatic classification accuracy of 82.85%. In the confusion matrix of the test dataset, the codes with the highest misclassification rates were related to train crew and vehicles, which is presumed to result from overlapping roles of the related departments and the contents of the complaints. This study verified the applicability and effectiveness of Korean BERT-based automatic complaint classification models through comparative analysis of the performance of several BERT models. Future studies require optimizing the system to enhance generalization capabilities by including more language data and more complex complaint types in the training process.

      • BERT 모델을 활용한 그룹웨어의 사용법 질의응답 챗봇 구현

        사수진 충북대학교 일반대학원 2025 국내석사

        RANK : 247807

        This paper presents a study on the development of a question-answering chatbot for groupware usage using BERT (Bidirectional Encoder Representations from Transformers). The primary objective of this research is to reduce the workload of customer support teams, provide 24/7 real-time assistance, and ensure consistent and accurate responses to user queries. Additionally, this study emphasizes the effectiveness of specialized models such as KLUE BERT for question-answering tasks related to groupware usage. Experimental results demonstrate that the KLUE BERT-based model exhibits high accuracy and consistent performance in Korean question-answering tasks, confirming its potential to significantly enhance the efficiency of customer support. keyword : Chatbot, BERT, Multilingual BERT, KLUE BERT, Natural Language Processing * A thesis for the degree of Master in February 2025

      • BERT 기반 의료 딥러닝 솔루션을 위한 한국어 임상기록지 중심의 다측면적 평가 방법론 개발

        김경모 서울대학교 대학원 2025 국내박사

        RANK : 247807

        배경: 다양한 종류의 Bidirectional encoder representations from transformers (BERT) 모델들이 언어 추론, 문서 분류, 정보 추출, 지식 추론과 같은 의료 딥러닝 솔루션을 위해서 연구 되었다. 그러나 기존의 대부분의 연구들은 영문 문서, 의료 이외의 분야의 문서들로 평가하였다. 영문 중심의 자연어처리 연구 추세와 반대로 한국어 임상 자연어처리 분야는 모델을 평가하는 방법론에 대한 심도깊은 연구가 부족하였다. 목적: 본 연구의 목적은 한국어 임상 기록지의 문맥을 가장 잘 이해할 수 있는 BERT 모델을 평가하는 방법론을 제안하는 것이다. 방법: 이를 위해서 기존의 자연어처리 연구 이론에 입각하여 다섯 종류의 평가 방법론을 제안하였다. 모델이 사전 학습한 분야에 따라 한국어 임상 문맥에 대한 이해도가 다를 것이라는 실험 가설을 세우고 다섯 종류의 BERT 모델을 선택하고 각 평가방법론 내에서 모델들의 성능을 비교하였다. 영문 분야를 사전학습 한 BERT-base, 영문 의생명분야를 사전학습 한 BioBERT, 임상기록지를 사전 학습한 Clinical BERT, 한국어 분야를 사전 학습한 KoBERT, 그리고 다국어를 사전학습 한 Multilingual BERT (M-BERT)를 비교 대상으로 선택하였다. 모델을 평가하기에 앞서서 선택한 모델들의 한국어 임상 문서에 대한 이해력을 증진시키기 위해서 서울대학교병원 159,460명의 외래경과지를 사전학습 하였다. 이후 자연어 추론, 문서 분류, 문맥 이해, 시간선 추론, 지식 추론 분야에서 BERT 모델들을 평가하기 위한 미세조정 작업을 제안하였다. 자연어 추론 능력을 평가하기 위해서 두 텍스트를 모델에게 입력 후 같은 환자의 것인지 분류하도록 하였다. 문서 분류 능력을 평가하기 위해서 모델이 문서의 진료과를 분류할 수 있는지 평가하였다. 문맥 이해 능력을 평가하기 위해서 환자기록의 평가(assessment) 문단의 범위를 찾을 수 있는지 평가하였다. 시간선 추론 능력 평가에서는 네 개의 환자기록 중 가장 마지막 문서를 찾을 수 있는지 평가하였다. 지식 추론 능력 평가에서는 주어진 문서의 알맞은 진단명을 추론할 수 있는지 평가하였다. 결과: 각 평가방법론을 BERT 모델에 적용했을 때의 성능을 통해서 제안한 평가 방법론의 타당성을 검토하고 한국어 임상기록지에서 BERT 모델의 특성을 발견하였다. 첫째로 방법론의 타당성을 검증하기 위해서 제안한 평가 방법론 내에서 모든 모델들이 너무 높은 성능을 내었는지 검토하였다. 그 결과 모든 모델이 95점 이상을 달성한 문서 분류를 제외하고 모든 평가 방법론에서 모델은 적정 수준의 성능을 내었다. 또한 대부분의 평가방법론 내에서 모델 성능 사이에서 높은 표준 편차를 가졌으며 이는 제안한 평가 방법론들이 적절한 변별력을 가졌음을 시사했다. 이는 최신 대형 디코더 모델인 Mistral 7B에서 적용했을 때도 비슷한 추세를 보였으며 이는 제안한 평가 방법론이 디코더 모델을 평가 할 때도 유효함을 시사하였다. 둘째로는 제안한 평가방법론을 통해 한국어 임상기록지에서 BERT 모델의 특성을 분석하였다. BioBERT, BERT-base는 [CLS] 토큰을 사용한 문서분류 작업에서 가장 효과적으로 동작하였지만 문맥이해, 시간선추론, 지식추론에서 M-BERT가 가장 효과적으로 동작하였다. 결론: 본 연구는 한국어 임상기록지를 사용한 다양한 의료 딥러닝 연구에서 BERT 모델들을 비교하는 방법론을 제안하였다. 또한 디코더 모델에도 적용하여 제안한 평가 방법론을 적용하는 범위의 확장 가능성을 확인하였다. 제안한 평가 방법론들은 향후 의료 분야에서 다양한 자연어처리 모델들을 비교 평가하는데 활용될 수 있을 것이다. Background: Various Bidirectional Encoder Representations from Transformers (BERT) models have been studied for medical deep learning solutions such as language inference, document classification, information extraction, and knowledge reasoning. However, most previous studies have evaluated these models using English texts or documents from non-medical domains. Unlike the English-centric trend in natural language processing (NLP) research, there has been a lack of in-depth studies on methodologies to evaluate models in the Korean clinical NLP domain. Objective: This study aims to propose a methodology for evaluating BERT models to determine their ability to understand the context of Korean clinical records effectively. Methods: To achieve this, five evaluation methods were developed based on existing NLP research theories. A hypothesis was established that the models' understanding of Korean clinical contexts would vary depending on their pretraining domain. Five types of BERT models were selected for comparison: BERT-Base pretrained on general English text, BioBERT pretrained on English biomedical literature, Clinical BERT pretrained on clinical records, KoBERT pretrained on general Korean text, and Multilingual BERT (M-BERT) pretrained on multilingual corpora. Before evaluation, the selected models were further pretrained on outpatient records from 159,460 patients at Seoul National University Hospital to enhance their understanding of Korean clinical texts. Subsequently, fine-tuning tasks were devised to evaluate the models across five domains: natural language inference, document classification, context understanding, timeline inference, and knowledge reasoning. For natural language inference, the models were tasked with determining whether two texts belonged to the same patient. Document classification involved evaluating the models’ ability to classify the medical department associated with a document. Context understanding was assessed by identifying the range of the "Assessment" section in patient records. Timeline inference required the models to identify the most recent document from a set of four records, while knowledge reasoning involved inferring the appropriate diagnosis from a given document. Results: The performance of the BERT models in each evaluation task was analyzed to validate the proposed evaluation methodologies and to discover the characteristics of BERT models in Korean clinical records. First, the validity of the methodologies was confirmed by examining whether the models achieved excessively high performance across all tasks. While all models achieved over 95 points in document classification, the other tasks showed reasonable performance levels. Additionally, most evaluation tasks revealed high standard deviations in model performance, suggesting that the methodologies possess sufficient discriminative power. Similar trends were observed when applying the methodologies to Mistral 7B, a state-of-the-art large decoder model, indicating that the proposed evaluation framework is also applicable to decoder models. Second, the proposed methodologies were used to analyze the characteristics of BERT models in Korean clinical records. BioBERT and BERT-Base were the most effective in document classification tasks that utilized the [CLS] token. However, M-BERT outperformed other models in context understanding, timeline inference, and knowledge reasoning tasks. Conclusion: This study proposed methodologies for comparing BERT models in various medical deep learning tasks using Korean clinical records. It also demonstrated the potential for extending the application of these methodologies to decoder models. The proposed evaluation methodologies can serve as a valuable tool for comparing and evaluating various NLP models in the medical domain in future research.

      • BERT와 Llama를 활용한 국내 학술지 논문의 자동분류 성능 비교

        강광선 경희대학교 대학원 2024 국내석사

        RANK : 247807

        BERT와 Llama를 활용한 국내 학술지 논문의 자동분류 성능 비교 경희대학교 테크노경영대학원 AI기술경영학과 강 광 선 초거대 인공지능 오픈 AI사의 ChatGPT의 열풍으로 다양한 LLM 모델이 발표되었 다. 2023년 2월 발표한 메타의 Llama 모델은 연구 커뮤니케이션에 오픈 하면서 거 대 언어 모델의 생태계를 활성화하였다. Llama2는 SFT, RLHF를 반복 학습하여 ChatGPT 3.5와 유사한 성능을 구현 하면서 상업적으로도 이용한 모델이다. 문서 자동분류 분야에 많이 이용되고 있는 Bert 모델과 최신 LLM 모델인 Llama2 모델 을 비교하여 Llama2 모델이 Bert 모델에 대비 문서 자동분류에서 성능이 향상되었 는지 검증하려고 한다. 학습데이터는 AI-HUB에 ‘논문자료 요약’ 데이터셋 사용하였 다. 학습데이터는 1995년부터 2020년까지 데이터 16만건이며 대상 분류는 한국연구 재단의 연구 분야 분류기준으로 8개 분류로 정의되어 있다. 본 연구를 위한 python 프로그램을 작성하였으며 Bert, Llama2의 학습 및 자동분류 성능 평가를 실행하였 다. 본 실험의 결과는 학습데이터의 오차의 경우 Short model의 경우 Bert가 더 낮 았고 middle, long model의 경우 Llama2가 더 낮았다. 학습데이터 정확도의 경우 short, middle, long model에서 Llama2가 Bert 보다 높은 정확도를 보였다. 행렬 분 석한 결과 Bert의 경우 사회과학, 공학, 농수해양에서 높았으며 Llama2는 인문학, 자연과학, 의약학, 예술체육, 복합학에서 빈도가 높게 나왔다. 분류 평가에서 short 모델의 경우 Bert가 Llama2보다 정밀도, 재현율, F1 스코어에서 우세한 결과가 나 왔다. middle, long 모델의 경우 Llama2가 Bert 보다 정밀도, 재현율, F1 스코어에 서 우세한 결과가 나왔다. 두 모델의 유의수준 5%의 쌍체 비교 t-검정을 실시하였다. short 모델은 성능 차 이가 없었고 middle 모델의 경우 정밀도는 성능 차이가 있고 재현율, F1 스코어는 차이가 없는 것으로 나왔다. long 모델의 경우 재현율은 성능 차이가 없고 정밀도, F1 스코어가 성능 차이가 있는 것으로 나왔다. 문서 자동 분류 모델 선택시 입력 길이가 Short 텍스트일 경우 Bert 모델, Long 텍스트일 경우 Llama2 모델의 사용 을 고려할 필요가 있다. 자동분류 모델 선택시 입력 데이터 길이에 따라 지표 판단 의 기준이 되는 실증분석 결과를 제시 하였다. 향후 연구에서는 다양한 LLM 모델의 활용해 보고 제로샷(Zero-shot) 및 퓨샷 (Few-shot) 학습을 이용한 문서 자동분류를 영역으로 연구하고자 한다. 주제어 : 인공지능, 자동분류, BERT, Llama, LLM

      • BERT를 이용한 한국어 기계독해와 질문생성모델

        이동헌 강원대학교 대학원 2021 국내석사

        RANK : 247807

        Machine Reading Comprehension(MRC) is to analyzing and inferring a paragraph received as input by a machine. Ushing machine reading comphrehension to understand given questions and paragraphs and outputting appropriate answers is question and answer using machine reading comprehension. Building machine reading learning data is difficult task, and you have to manually create the correct answers and questions that can derive the correct answers that appear in the document. In order to solve this problem research on automatic question generation has been actively studied. As opposed to machine reading comprehension, question generation is a task of generating question that can derive correct answers by looking at documents and correct answers. BERT is a language model showing excellent performance in various natural language processing tasks in recent years, and it learns a language model with a transformer with bedirectionality for a large-scale corpus. The pre-trained BERT can be applied to natural language processing tasks by adding an output layer. In this paper, we use KorQuAD 1.0 and KorQuAD 2.0 which are Korean question and answer datasets for machine reading comprehension learning. We propose machine reading comprehension model that adds a SRU(Simaple Recurrent Unit) on pre-trained BERT model and features suitable each dataset and BERT-based Sequence-to-sequence model that adds copying mechanism to then model that automatically generates a question from the document to which the correct answer belongs. As a result of the experiment, when the proposed in this paper was applied to KorQuAD 1.0 data, EM 85.35%, F1 93.24% were shown in the development set, and when was applied to KorQuAD 2.0 data, EM 49.2%, F1 71.21% were shown. In addition, the performance of the BERT-based Transformer decoder model was better than that of the exising model and the BERT + GRU decoder model. 기계 독해는 기계가 입력으로 받은 문단을 분석하고 추론하는 것을 말한다. 기계 독해를 이용하여 주어진 질문과 문단을 이해하고 이에 알맞은 답을 출력하는 것을 기계 독해를 이용한 질의 응답이라 한다. 기계 독해 학습 데이터 구축은 어려운 작업으로, 문서에서 등장하는 정답과 정답을 도출할 수 있는 질문을 수작업으로 만들어야 한다. 이를 해결하기 위하여 최근 질문 자동 생성 연구가 활발히 연구되고 있다. 질문 생성은 기계 독해와 반대로, 문서와 정답을 보고 정답을 도출할 수 있는 질문을 생성하는 태스크이다. BERT는 최근 다양한 자연어 처리 태스크에서 뛰어난 성능을 보이고 있는 언어 모델로 대용량 코퍼스에 대하여 양방향성을 가진 트랜스포머(transformer)로 언어 모델을 학습한다. 사전 학습 된 BERT는 출력 층(layer)을 추가하여 자연어 처리 태스크에 적용할 수 있다. 본 논문에서는 기계 독해 학습을 위해 한국어 질의 응답 데이터 셋인 KorQuAD 1.0과 KorQuAD 2.0을 이용하며, 사전 학습 된 BERT 모델 위에 SRU(Simple Recurrent Unit)와 각 데이터에 적합한 자질을 추가한 모델과 정답이 속한 문서로부터 질문을 자동으로 생성해주는 모델에 복사 메커니즘을 추가한 BERT 기반의 Sequence-to-sequence 모델을 제안한다. 실험 결과, 본 논문에서 제안한 방법을 KorQuAD 1.0 데이터에 적용한 경우, 개발 셋에서 EM 85.35%, F1 93.24%의 성능을 보였으며, KorQuAD 2.0 데이터에 적용하였을 때는 EM 49.2%, F1 71.21%의 성능을 보였다. 또한, BERT 기반의 Transformer 디코더 모델의 성능이 기존 모델과 BERT + GRU 디코더 모델보다 좋았다.

      • BERT의 지식 전이학습 모형을 이용한 비즈니스 모델 캔버스(BMC) 자연어 분석 연구

        신병규 경희대학교 대학원 2021 국내박사

        RANK : 247807

        BERT의 지식 전이학습 모형을 이용한 비즈니스 모델 캔버스(BMC) 자연어 분석 연구 비즈니스 모델(Business Model)은 기업이 비즈니스를 어떻게 수행할 것인가에 대한 설계도이며 고객에게 제품과 서비스를 제공하여 수익을 창출하는 일련의 과정을 설명하는 것이다. 그러나 비즈니스 모델의 중요성에도 불구하고 고객이 원하는 제품을 파악하지 못해 시장 진입에 실패하는 경우가 50% 가까이 발생하고 있다. 비즈니스 모델이 부족하다는 것은 기업의 생존에 필요한 수익을 기대하기 어렵고 성장에 한계가 있다는 의미이다. 이와 같은 실패 요인은 비즈니스 모델이 중요한데도 불구하고 스타트업 기업을 비롯하여 국내 기업의 99%가 50인 미만의 중소기업이기 때문에 제대로 된 비즈니스 모델을 수립하는 데 한계가 있다. 즉, 기업의 비즈니스 모델을 정확하게 수립하고 평가할 수 있다면 중소기업이나 스타트업 기업의 비즈니스 성공 확률은 상승할 것으로 기대된다. 비즈니스 모델을 수립하고 평가하기 위해 접근할 수 있는 툴이 있으면 대다수의 중소기업에 도움이 될 것이다. 비즈니스 모델 캔버스(BMC, Business Model Canvas)는 고객 세그먼트, 가치제안, 채널, 고객관계, 수익원, 핵심자원, 핵심활동, 핵심파트너, 비용구조를 9개 블록으로 나누어 한 장의 캔버스에 요약한 모델이다. 비즈니스 모델 캔버스는 기업이 목표로 하는 고객에게 기업의 핵심역량을 통해 만든 핵심가치를 어떻게 전달하여 수익을 창출하는지를 직관적으로 파악할 수 있도록 비즈니스 모델을 작성할 수 있는 편리한 도구이다. 비즈니스 모델 캔버스는 기업이 어떻게 고객가치를 만들고, 전파하는지, 그리고 이로 인해 어떻게 수익을 창출하는지에 대한 원리를 이해하고 유용하게 사용할 수 있는 도구이다. 본 연구에서는 BERT(Bidirectional Encoder Representations from Transformers) 모델을 활용하여 지식 전이학습을 통한 비정형데이터인 비즈니스 모델 캔버스를 객관적으로 평가하는 모델을 개발하였다. BERT 모델을 이용하여 건설제조업과 IT기업의 비즈니스 모델 캔버스로부터 자동으로 비즈니스 모델을 추출한 후, 이를 비즈니스 모델 평가에 직접 활용하는 평가모델을 제안하고, 이에 대한 성능을 검증하였다. 이를 위해 506개 건설제조업과 542개 IT기업의 비즈니스 모델 캔버스 데이터를 수집하였다. 분석에는 트랜스포머(Transformer) 기반 딥러닝 NLP(Natural Language Processing) 중 대표적인 BERT 모델을 사용하였다. BERT는 사전학습(pre-training) 과정을 거쳐서 컴퓨터에 대용량의 언어를 이해시켰다. 이후에 지식 전이과정을 통해서 사용 목적에 맞게 분야별 언어를 집중적으로 학습하는 파인튜닝(fine-tuning)을 실시하였고, 건설제조업과 IT기업의 BMC 모델을 이해할 수 있게 학습시켰다. SKT-Brain에서 사전학습 시킨 한국어 BERT 모델을 기반으로 비즈니스 모델 캔버스의 구성요소에 대한 평가를 분류하기 위해 전이학습(Transfer Learning)을 통해 파인튜닝을 실행하였다. 본 연구는 비정형 데이터 비즈니스 모델 캔버스의 자연어 분석을 통해 비즈니스 모델 지수를 예측하는 것이다. 분석절차는 한국어 위키사전과 한국어 뉴스로 사전학습을 한 KoBERT에 비즈니스 모델 캔버스 비정형 데이터를 학습시켜 비즈니스 모델 지수를 예측하였다. 하이퍼 파라미터 파인 튜닝은 AdamWoptimizer를 사용하였고, 학습은 총 50 epoch, minibatch 사이즈는 32, 드롭아웃은 0.1, 학습률 2e-5, L2정규화 계수는 5e-5, weight decay를 L2 정규화로 0.01, 모델 구현과 실행은 python 3.7 pytorch 1.8.0 CUDA11.1를 사용하였다. 예측 모델 평가는 9개의 비즈니스 모델 캔버스의 구성요소를 독립적인 요소로 보고, Hold-Out 검증을 통해 train data와 test data를 7대3으로 구분하여 랜덤하게 진행하였다. 모델평가는 오분류표를 이용하여 정확도를 계산하였고, 교차 엔트로피를 이용하여 로스(loss)계산을 하였다. 비즈니스 모델 캔버스 9개 구성요소의 5라벨 평균 정확도는 건설제조업이 0.606이고 IT기업은 0.508이다. 근접 라벨의 평균 정확도는 건설제조업과 IT기업 모두 0.933이다. 본 연구의 학문적 의의는 기업 혁신이나 사업계획서를 바탕으로 하는 비정형데이터인 비즈니스 모델 캔버스 데이터를 KoBERT 모델을 사용하여, 기업의 비즈니스 모델 평가에 대한 예측 모델을 최초로 개발하여 연구하였다는 것이다. 이는 비즈니스 모델을 평가하고 이를 바탕으로 기업이 목표로 하는 경영 혁신을 이룰 수 있는 계기가 될 것으로 기대한다. 실무적 의의는 비즈니스 모델 캔버스 데이터의 비정형 텍스트 데이터를 이용하여, 자연어 처리 딥러닝 모델을 기반으로 산업 현장에 적용할 수 있는 평가 모델을 개발함으로써 실무적으로 유용하게 이용될 수 있을 것으로 예상된다. 또한, 제안된 모델의 응용으로 비즈니스 모델 자기 진단시스템을 제안하였다. 비즈니스 모델 자기 진단시스템을 이용하면 기업은 전문가의 도움 없이 비즈니스 모델 분석을 하고 개선 사항을 파악할 수 있을 것으로 기대된다. 비즈니스 모델 9개 블록 요인에 대한 점수를 도출하여 선도업체나 동종업종과 비교하여 자사의 강점과 보완 사항을 파악하고 미래 전략을 수립하는 데 기여하도록 하였다. 본 연구의 한계점 및 향후 연구는 다음과 같다. 비즈니스 모델 캔버스의 9개 요소가 개별 기업의 노하우가 담긴 내용이 많고, 회사 내부 정보에 대한 데이터이기 때문에 폭넓은 산업을 대상으로 하지 못하고 건설제조업과 IT기업만을 대상으로 데이터를 수집하여 연구한 점이다. 향후 건설제조업과 IT기업 이외의 타 산업의 비즈니스 모델 구성요소에 대한 데이터를 수집하여 예측 정확도의 차이가 비즈니스 모델 구성요소의 요인별 특성인지 산업 도메인 차이 인지를 비교하는 연구가 필요하다. 본 연구를 바탕으로 BERT기반 BMC모델을 활용하여 사업계획서, 인사평가, 제안서 분석, 중장기 전략 등 회사 경영에 전반적으로 적용한다면 맨파워가 부족한 많은 중소기업이나 스타트업 기업에게 큰 도움이 될 것으로 기대된다. 주제어: 비즈니스 모델, 비즈니스 모델 캔버스, BERT, 전이학습, 자연어 분석, 딥러닝, 텍스트 마이닝 Text Analysis of Business Model Canvas Using BERT's Knowledge Transfer Learning A business model is a blueprint for how a company will conduct business and describes a series of processes that generate revenue by providing products and services to customers. However, despite the importance of the business model, nearly 50% of the cases fail to enter the market because they do not understand the product they want. The lack of a business model means that it is difficult to expect the profits necessary for the survival of a company and there is a limit to growth. Despite the importance of the business model, there is a limit to establishing a proper business model because 99% of domestic companies, including start-ups, are small and medium-sized enterprises (SMEs) with fewer than 50 employees. In other words, if a company's business model can be accurately established and evaluated, the probability of business success of SMEs or start-ups is expected to increase. Most small businesses will benefit from having tools they can access to build and evaluate their business models. Business Model Canvas(BMC) is a model that summarizes Customer Segments, Value Propositions, Channels, Customer Relationships, Revenue Streams, Key Resources, Key Activities, Key Partners, Cost Structure and cost structure into 9 blocks on one canvas. The business model canvas is a convenient tool that allows you to intuitively understand how to generate profits by delivering core values to customers based on the company's core competencies to target customers. The business model canvas is a tool that can be useful and understand the principles of how a company creates, disseminates, and monetizes customer value. In this study, using the BERT (Bidirectional Encoder Representations from Transformers) model, a model was developed to objectively evaluate the business model canvas, which is unstructured data through knowledge transfer learning. After automatically extracting a business model from the business model canvas of the construction manufacturing industry and IT company using the BERT model, an evaluation model that directly uses it for business model evaluation was proposed and its performance was verified. For this purpose, business model canvas data of 506 construction and manufacturing industries and 542 IT companies were collected. For the analysis, a representative BERT model among Transformer-based deep learning NLP (Natural Language Processing) was used. BERT made the computer understand a large amount of language through a pre-training process. After that, fine-tuning was conducted to intensively learn languages for each field according to the purpose of use through the knowledge transfer process, and the BMC model of the construction manufacturing industry and IT companies was learned to be understood. Fine tuning was performed through transfer learning to classify the evaluation of the components of the business model canvas based on the Korean BERT model trained in advance by SKT-Brain. This study predicts the business model index through natural language analysis of the unstructured data business model canvas. As for the analysis procedure, the business model index was predicted by learning the business model canvas unstructured data in KoBERT, which was previously studied with Korean wiki dictionary and Korean news. For hyperparameter fine tuning, AdamWoptimizer was used, training was performed for a total of 50 epochs, minibatch size was 32, dropout was 0.1, learning rate was 2e-5, L2 regularization coefficient was 5e-5, weight decay was 0.01 for L2 regularization, model implementation and Execution was performed using python 3.7 pytorch 1.8.0 CUDA11.1. For the prediction model evaluation, the 9 business model canvas components were regarded as independent elements, and train data and test data were divided 7 to 3 through hold-out verification and randomly performed. For model evaluation, accuracy was calculated using a misclassification table, and loss was calculated using cross entropy. The average accuracy of the 5 labels of the 9 components of the business model canvas is 0.606 for the construction and manufacturing industry and 0.508 for the IT company. The average accuracy of proximity label is 0.933 for both construction manufacturing and IT companies. The academic significance of this study is that the business model canvas data, which is unstructured data based on corporate innovation or business plan, was used for the first time to develop and study a predictive model for corporate business model evaluation using the KoBERT model. It is expected that this will serve as an opportunity to evaluate the business model and achieve the management innovation targeted by the company based on it. The practical significance is expected to be practically useful by developing an evaluation model that can be applied to industrial sites based on a natural language processing deep learning model using the unstructured text data of the business model canvas data. In addition, a business model self-diagnosis system was proposed as an application of the proposed model. By using the business model self-diagnosis system, companies are expected to be able to analyze business models and identify improvements without the help of experts. By deriving scores for the 9 business model block factors, we compared them with leading companies or the same industry to identify their strengths and complements and contribute to establishing future strategies. The limitations of this study and future research are as follows. Because the 9 elements of the business model canvas contain a lot of the know-how of individual companies and are data about company internal information, it is not possible to target a wide range of industries, but only the construction manufacturing industry and IT companies. In the future, it is necessary to collect data on business model components of industries other than the construction manufacturing industry and IT companies to compare whether the difference in prediction accuracy is a characteristic of each factor of the business model component or a difference in the industry domain. Based on this study, if the BERT-based BMC model is applied to overall company management such as business plans, personnel evaluation, proposal analysis, and mid- to long-term strategies, it is expected that it will be of great help to many SMEs and startups lacking manpower. Key words: Business Model, Business Model Canvas, BERT, Transfer Learning, Natural Language Analysis, Deep Learning, Text Mining

      • Topic modeling based Siamese-LSTM-BERT model for semantic document similarity

        김동욱 Graduate School, Yonsei University 2022 국내석사

        RANK : 247807

        BERT는 다양한 자연어 처리 작업에서 우수한 성능을 보여주며, 풍부한 표현의 텍스트 임베딩을 생성한다. 그러나 BERT에 적용가능한 텍스트의 길이가 제한되기 때문에 길이가 짧은 문장들에 대해서만 연구가 주로 진행되었다. BERT 기반의 문서 임베딩 생성을 위한 여러 시도가 있었지만, 문서 내용을 발췌하여 문서의 일부분만으로만 임베딩을 생성하였다. 문서 발췌 방법은 문서의 정보 손실을 발생시키기 때문에 정확한 문서 임베딩을 형성하는 데는 한계가 있다. 문서의 발췌 정보를 사용하는 방법은 문서 구조에 대한 경험적인 지식이나 문서 내 중요한 부분을 미리 알아야 하는 조건이 요구된다. 그러나, 특정 도메인에 대한 문서의 경우 전문적인 도메인 지식이 부족할 경우 문서에서 중요한 내용을 알기 어려우며 문서의 유형이 다를 때마다 임베딩 할 문서의 위치를 매번 변경해야 하며, 잘못된 부분을 발췌할 경우 정확하지 않은 문서 임베딩이 생성될 수 있다. 본 연구에서는 BERT를 기반으로 길이가 긴 문서를 위한 문서 임베딩 방법을 제안한다. 연구에서 제안한 모델은 문서를 세그먼트로 나누고 이를 각각의 시퀀스 상태로 간주한다. 이 시퀀스는 LSTM을 통해 하나의 문서 임베딩으로 표현될 수 있으며, 문서의 전체적인 정보를 사용하는 데 도움을 줄 수 있다. 그리고 도메인에 적합한 문서 임베딩 생성을 위해 토픽 모델링을 사용하여 각각의 세그먼트에 대해서 토픽 분포 정보를 결합하여 도메인에 특화된 문서 임베딩을 생성하였다. 제안된 방식의 모델은 Siamese Network를 사용하여 토픽 모델링을 통해 추론된 토픽 분포 정보와 BERT를 통해 추론된 세그먼트가 결합된 문서 임베딩을 기반으로 문서 간의 유사성을 판별하는 작업을 수행한다. 또한, 기존 BERT의 적용가능한 최대 길이 문제를 개선하여 문서의 로컬 정보를 임베딩에 사용하는 대신 문서의 전역 정보를 기반으로 임베딩을 생성할 수 있도록 하고, 토픽정보를 활용해 도메인에 특화된 문서의 유사성 판별에 기존 연구방법론 보다 향상된 결과를 보여준다. BERT shows the state of art in various natural language processing tasks and generates text embedding of rich expressions. However, since the length of text applicable to BERT is limited, research has been mainly conducted only on short sentences. There have been several attempts to generate BERT-based document embedding, but embeddings were generated only with a part of the document by extracting the contents of the document. Since document extraction methods cause loss of information on documents, there is a limit to forming accurate document embedding. The method of using document excerpt information requires empirical knowledge of the document structure or conditions for knowing important parts of the document in advance. However, for documents for a particular domain, lack of professional domain knowledge makes it difficult to know important content in the document, and whenever the type of document differs, the location of the document to be embedded must be changed every time, and inaccurate document embedding may be generated. This study proposes a document embedding method for long documents based on BERT. The model proposed in the study divides the document into segments and regards it as each sequence state. This sequence can be expressed as a single document embedding through LSTM and can help use the document's overall information. In addition, to generate document embedding suitable for the domain, topic distribution information was combined for each segment using topic modeling to generate document embedding specific to the domain. The proposed model uses the Siamese Network to determine the similarity between documents based on document embedding that combines topic distribution information and segments. In addition, it improves the maximum applicable length problem of existing BERT so that embeddings can be generated based on global information of documents instead of using only part of documents for embedding and combines topic distribution information with document embedding to show better results than existing methodologies.

      연관 검색어 추천

      이 검색어로 많이 본 자료

      활용도 높은 자료

      해외이동버튼