RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    BERT 기반 의료 딥러닝 솔루션을 위한 한국어 임상기록지 중심의 다측면적 평가 방법론 개발 = Development of a Multifaceted Evaluation Methodology for BERT-Based Medical Deep Learning Solutions Focused on Korean Clinical Records

    한글로보기

    https://www.riss.kr/link?id=T17242060

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    배경: 다양한 종류의 Bidirectional encoder representations from transformers (BERT) 모델들이 언어 추론, 문서 분류, 정보 추출, 지식 추론과 같은 의료 딥러닝 솔루션을 위해서 연구 되었다. 그러나 기존의 대부분의 연구들은 영문 문서, 의료 이외의 분야의 문서들로 평가하였다. 영문 중심의 자연어처리 연구 추세와 반대로 한국어 임상 자연어처리 분야는 모델을 평가하는 방법론에 대한 심도깊은 연구가 부족하였다.

    목적: 본 연구의 목적은 한국어 임상 기록지의 문맥을 가장 잘 이해할 수 있는 BERT 모델을 평가하는 방법론을 제안하는 것이다.

    방법: 이를 위해서 기존의 자연어처리 연구 이론에 입각하여 다섯 종류의 평가 방법론을 제안하였다. 모델이 사전 학습한 분야에 따라 한국어 임상 문맥에 대한 이해도가 다를 것이라는 실험 가설을 세우고 다섯 종류의 BERT 모델을 선택하고 각 평가방법론 내에서 모델들의 성능을 비교하였다. 영문 분야를 사전학습 한 BERT-base, 영문 의생명분야를 사전학습 한 BioBERT, 임상기록지를 사전 학습한 Clinical BERT, 한국어 분야를 사전 학습한 KoBERT, 그리고 다국어를 사전학습 한 Multilingual BERT (M-BERT)를 비교 대상으로 선택하였다. 모델을 평가하기에 앞서서 선택한 모델들의 한국어 임상 문서에 대한 이해력을 증진시키기 위해서 서울대학교병원 159,460명의 외래경과지를 사전학습 하였다. 이후 자연어 추론, 문서 분류, 문맥 이해, 시간선 추론, 지식 추론 분야에서 BERT 모델들을 평가하기 위한 미세조정 작업을 제안하였다. 자연어 추론 능력을 평가하기 위해서 두 텍스트를 모델에게 입력 후 같은 환자의 것인지 분류하도록 하였다. 문서 분류 능력을 평가하기 위해서 모델이 문서의 진료과를 분류할 수 있는지 평가하였다. 문맥 이해 능력을 평가하기 위해서 환자기록의 평가(assessment) 문단의 범위를 찾을 수 있는지 평가하였다. 시간선 추론 능력 평가에서는 네 개의 환자기록 중 가장 마지막 문서를 찾을 수 있는지 평가하였다. 지식 추론 능력 평가에서는 주어진 문서의 알맞은 진단명을 추론할 수 있는지 평가하였다.

    결과: 각 평가방법론을 BERT 모델에 적용했을 때의 성능을 통해서 제안한 평가 방법론의 타당성을 검토하고 한국어 임상기록지에서 BERT 모델의 특성을 발견하였다. 첫째로 방법론의 타당성을 검증하기 위해서 제안한 평가 방법론 내에서 모든 모델들이 너무 높은 성능을 내었는지 검토하였다. 그 결과 모든 모델이 95점 이상을 달성한 문서 분류를 제외하고 모든 평가 방법론에서 모델은 적정 수준의 성능을 내었다. 또한 대부분의 평가방법론 내에서 모델 성능 사이에서 높은 표준 편차를 가졌으며 이는 제안한 평가 방법론들이 적절한 변별력을 가졌음을 시사했다. 이는 최신 대형 디코더 모델인 Mistral 7B에서 적용했을 때도 비슷한 추세를 보였으며 이는 제안한 평가 방법론이 디코더 모델을 평가 할 때도 유효함을 시사하였다. 둘째로는 제안한 평가방법론을 통해 한국어 임상기록지에서 BERT 모델의 특성을 분석하였다. BioBERT, BERT-base는 [CLS] 토큰을 사용한 문서분류 작업에서 가장 효과적으로 동작하였지만 문맥이해, 시간선추론, 지식추론에서 M-BERT가 가장 효과적으로 동작하였다.

    결론: 본 연구는 한국어 임상기록지를 사용한 다양한 의료 딥러닝 연구에서 BERT 모델들을 비교하는 방법론을 제안하였다. 또한 디코더 모델에도 적용하여 제안한 평가 방법론을 적용하는 범위의 확장 가능성을 확인하였다. 제안한 평가 방법론들은 향후 의료 분야에서 다양한 자연어처리 모델들을 비교 평가하는데 활용될 수 있을 것이다.
    번역하기

    배경: 다양한 종류의 Bidirectional encoder representations from transformers (BERT) 모델들이 언어 추론, 문서 분류, 정보 추출, 지식 추론과 같은 의료 딥러닝 솔루션을 위해서 연구 되었다. 그러나 기존의 ...

    배경: 다양한 종류의 Bidirectional encoder representations from transformers (BERT) 모델들이 언어 추론, 문서 분류, 정보 추출, 지식 추론과 같은 의료 딥러닝 솔루션을 위해서 연구 되었다. 그러나 기존의 대부분의 연구들은 영문 문서, 의료 이외의 분야의 문서들로 평가하였다. 영문 중심의 자연어처리 연구 추세와 반대로 한국어 임상 자연어처리 분야는 모델을 평가하는 방법론에 대한 심도깊은 연구가 부족하였다.

    목적: 본 연구의 목적은 한국어 임상 기록지의 문맥을 가장 잘 이해할 수 있는 BERT 모델을 평가하는 방법론을 제안하는 것이다.

    방법: 이를 위해서 기존의 자연어처리 연구 이론에 입각하여 다섯 종류의 평가 방법론을 제안하였다. 모델이 사전 학습한 분야에 따라 한국어 임상 문맥에 대한 이해도가 다를 것이라는 실험 가설을 세우고 다섯 종류의 BERT 모델을 선택하고 각 평가방법론 내에서 모델들의 성능을 비교하였다. 영문 분야를 사전학습 한 BERT-base, 영문 의생명분야를 사전학습 한 BioBERT, 임상기록지를 사전 학습한 Clinical BERT, 한국어 분야를 사전 학습한 KoBERT, 그리고 다국어를 사전학습 한 Multilingual BERT (M-BERT)를 비교 대상으로 선택하였다. 모델을 평가하기에 앞서서 선택한 모델들의 한국어 임상 문서에 대한 이해력을 증진시키기 위해서 서울대학교병원 159,460명의 외래경과지를 사전학습 하였다. 이후 자연어 추론, 문서 분류, 문맥 이해, 시간선 추론, 지식 추론 분야에서 BERT 모델들을 평가하기 위한 미세조정 작업을 제안하였다. 자연어 추론 능력을 평가하기 위해서 두 텍스트를 모델에게 입력 후 같은 환자의 것인지 분류하도록 하였다. 문서 분류 능력을 평가하기 위해서 모델이 문서의 진료과를 분류할 수 있는지 평가하였다. 문맥 이해 능력을 평가하기 위해서 환자기록의 평가(assessment) 문단의 범위를 찾을 수 있는지 평가하였다. 시간선 추론 능력 평가에서는 네 개의 환자기록 중 가장 마지막 문서를 찾을 수 있는지 평가하였다. 지식 추론 능력 평가에서는 주어진 문서의 알맞은 진단명을 추론할 수 있는지 평가하였다.

    결과: 각 평가방법론을 BERT 모델에 적용했을 때의 성능을 통해서 제안한 평가 방법론의 타당성을 검토하고 한국어 임상기록지에서 BERT 모델의 특성을 발견하였다. 첫째로 방법론의 타당성을 검증하기 위해서 제안한 평가 방법론 내에서 모든 모델들이 너무 높은 성능을 내었는지 검토하였다. 그 결과 모든 모델이 95점 이상을 달성한 문서 분류를 제외하고 모든 평가 방법론에서 모델은 적정 수준의 성능을 내었다. 또한 대부분의 평가방법론 내에서 모델 성능 사이에서 높은 표준 편차를 가졌으며 이는 제안한 평가 방법론들이 적절한 변별력을 가졌음을 시사했다. 이는 최신 대형 디코더 모델인 Mistral 7B에서 적용했을 때도 비슷한 추세를 보였으며 이는 제안한 평가 방법론이 디코더 모델을 평가 할 때도 유효함을 시사하였다. 둘째로는 제안한 평가방법론을 통해 한국어 임상기록지에서 BERT 모델의 특성을 분석하였다. BioBERT, BERT-base는 [CLS] 토큰을 사용한 문서분류 작업에서 가장 효과적으로 동작하였지만 문맥이해, 시간선추론, 지식추론에서 M-BERT가 가장 효과적으로 동작하였다.

    결론: 본 연구는 한국어 임상기록지를 사용한 다양한 의료 딥러닝 연구에서 BERT 모델들을 비교하는 방법론을 제안하였다. 또한 디코더 모델에도 적용하여 제안한 평가 방법론을 적용하는 범위의 확장 가능성을 확인하였다. 제안한 평가 방법론들은 향후 의료 분야에서 다양한 자연어처리 모델들을 비교 평가하는데 활용될 수 있을 것이다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Background: Various Bidirectional Encoder Representations from Transformers (BERT) models have been studied for medical deep learning solutions such as language inference, document classification, information extraction, and knowledge reasoning. However, most previous studies have evaluated these models using English texts or documents from non-medical domains. Unlike the English-centric trend in natural language processing (NLP) research, there has been a lack of in-depth studies on methodologies to evaluate models in the Korean clinical NLP domain.

    Objective: This study aims to propose a methodology for evaluating BERT models to determine their ability to understand the context of Korean clinical records effectively.

    Methods: To achieve this, five evaluation methods were developed based on existing NLP research theories. A hypothesis was established that the models' understanding of Korean clinical contexts would vary depending on their pretraining domain. Five types of BERT models were selected for comparison: BERT-Base pretrained on general English text, BioBERT pretrained on English biomedical literature, Clinical BERT pretrained on clinical records, KoBERT pretrained on general Korean text, and Multilingual BERT (M-BERT) pretrained on multilingual corpora. Before evaluation, the selected models were further pretrained on outpatient records from 159,460 patients at Seoul National University Hospital to enhance their understanding of Korean clinical texts. Subsequently, fine-tuning tasks were devised to evaluate the models across five domains: natural language inference, document classification, context understanding, timeline inference, and knowledge reasoning.

    For natural language inference, the models were tasked with determining whether two texts belonged to the same patient. Document classification involved evaluating the models’ ability to classify the medical department associated with a document. Context understanding was assessed by identifying the range of the "Assessment" section in patient records. Timeline inference required the models to identify the most recent document from a set of four records, while knowledge reasoning involved inferring the appropriate diagnosis from a given document.

    Results: The performance of the BERT models in each evaluation task was analyzed to validate the proposed evaluation methodologies and to discover the characteristics of BERT models in Korean clinical records. First, the validity of the methodologies was confirmed by examining whether the models achieved excessively high performance across all tasks. While all models achieved over 95 points in document classification, the other tasks showed reasonable performance levels. Additionally, most evaluation tasks revealed high standard deviations in model performance, suggesting that the methodologies possess sufficient discriminative power. Similar trends were observed when applying the methodologies to Mistral 7B, a state-of-the-art large decoder model, indicating that the proposed evaluation framework is also applicable to decoder models.

    Second, the proposed methodologies were used to analyze the characteristics of BERT models in Korean clinical records. BioBERT and BERT-Base were the most effective in document classification tasks that utilized the [CLS] token. However, M-BERT outperformed other models in context understanding, timeline inference, and knowledge reasoning tasks.

    Conclusion: This study proposed methodologies for comparing BERT models in various medical deep learning tasks using Korean clinical records. It also demonstrated the potential for extending the application of these methodologies to decoder models. The proposed evaluation methodologies can serve as a valuable tool for comparing and evaluating various NLP models in the medical domain in future research.
    번역하기

    Background: Various Bidirectional Encoder Representations from Transformers (BERT) models have been studied for medical deep learning solutions such as language inference, document classification, information extraction, and knowledge reasoning. Howev...

    Background: Various Bidirectional Encoder Representations from Transformers (BERT) models have been studied for medical deep learning solutions such as language inference, document classification, information extraction, and knowledge reasoning. However, most previous studies have evaluated these models using English texts or documents from non-medical domains. Unlike the English-centric trend in natural language processing (NLP) research, there has been a lack of in-depth studies on methodologies to evaluate models in the Korean clinical NLP domain.

    Objective: This study aims to propose a methodology for evaluating BERT models to determine their ability to understand the context of Korean clinical records effectively.

    Methods: To achieve this, five evaluation methods were developed based on existing NLP research theories. A hypothesis was established that the models' understanding of Korean clinical contexts would vary depending on their pretraining domain. Five types of BERT models were selected for comparison: BERT-Base pretrained on general English text, BioBERT pretrained on English biomedical literature, Clinical BERT pretrained on clinical records, KoBERT pretrained on general Korean text, and Multilingual BERT (M-BERT) pretrained on multilingual corpora. Before evaluation, the selected models were further pretrained on outpatient records from 159,460 patients at Seoul National University Hospital to enhance their understanding of Korean clinical texts. Subsequently, fine-tuning tasks were devised to evaluate the models across five domains: natural language inference, document classification, context understanding, timeline inference, and knowledge reasoning.

    For natural language inference, the models were tasked with determining whether two texts belonged to the same patient. Document classification involved evaluating the models’ ability to classify the medical department associated with a document. Context understanding was assessed by identifying the range of the "Assessment" section in patient records. Timeline inference required the models to identify the most recent document from a set of four records, while knowledge reasoning involved inferring the appropriate diagnosis from a given document.

    Results: The performance of the BERT models in each evaluation task was analyzed to validate the proposed evaluation methodologies and to discover the characteristics of BERT models in Korean clinical records. First, the validity of the methodologies was confirmed by examining whether the models achieved excessively high performance across all tasks. While all models achieved over 95 points in document classification, the other tasks showed reasonable performance levels. Additionally, most evaluation tasks revealed high standard deviations in model performance, suggesting that the methodologies possess sufficient discriminative power. Similar trends were observed when applying the methodologies to Mistral 7B, a state-of-the-art large decoder model, indicating that the proposed evaluation framework is also applicable to decoder models.

    Second, the proposed methodologies were used to analyze the characteristics of BERT models in Korean clinical records. BioBERT and BERT-Base were the most effective in document classification tasks that utilized the [CLS] token. However, M-BERT outperformed other models in context understanding, timeline inference, and knowledge reasoning tasks.

    Conclusion: This study proposed methodologies for comparing BERT models in various medical deep learning tasks using Korean clinical records. It also demonstrated the potential for extending the application of these methodologies to decoder models. The proposed evaluation methodologies can serve as a valuable tool for comparing and evaluating various NLP models in the medical domain in future research.

    더보기

    목차 (Table of Contents)

    • 1. 서 론 1
    • 1.1. 연구 배경 1
    • 1.1.1. 의료 자연어처리 기술의 중요성 1
    • 1.1.2. 의료 자연어처리 기술의 발전 3
    • 1.1.3. 기존 연구의 한계점 및 연구 목적 10
    • 1. 서 론 1
    • 1.1. 연구 배경 1
    • 1.1.1. 의료 자연어처리 기술의 중요성 1
    • 1.1.2. 의료 자연어처리 기술의 발전 3
    • 1.1.3. 기존 연구의 한계점 및 연구 목적 10
    • 2. 이론적 배경 13
    • 2.1. BERT 13
    • 2.1.1. 트랜스포머(transformer) 13
    • 2.1.2. BERT 모델의 구조 및 주요 특징 20
    • 2.1.3. 사전학습 24
    • 2.1.4. 미세조정 연구 28
    • 2.2. BERT를 사용한 의료 자연어처리 연구동향 32
    • 3. 연구방법 34
    • 3.1. 실험 설계 34
    • 3.1.1. 연구 가설 34
    • 3.1.2. 비교 대상 모델 설정 35
    • 3.1.3. 다측면적 평가 방법론 제안 40
    • 3.1.4. 학습 및 평가 과정 42
    • 3.2. 데이터 획득 및 전처리 44
    • 3.2.1. 데이터 획득 44
    • 3.2.2. 데이터 분할 47
    • 3.3. 다국어, 의과학적 BERT 모델 사전학습 49
    • 3.3.1. 데이터 전처리 49
    • 3.3.2. 사전학습 55
    • 3.4. 유사 문서 분류 능력 평가 (task 1, task 2) 62
    • 3.4.1. 이론적 배경–자연어 추론 62
    • 3.4.2. 유사 문서 분류 능력 (task 1): 평가 목표 63
    • 3.4.3. 유사 문서 분류 능력 (task 1): 데이터 전처리 65
    • 3.4.4. 유사 문서 분류 능력 (task 1): 입력 변수 66
    • 3.4.5. 유사 문단 분류 능력 (task 2): 평가 목표 67
    • 3.4.6. 유사 문단 분류 능력 (task 2): 데이터 전처리 69
    • 3.4.7. 유사 문단 분류 능력 (task 2): 입력 변수 69
    • 3.4.8. Model architectures 70
    • 3.5. 문서 진료과 분류 능력 평가 (task 3) 71
    • 3.5.1. 이론적 배경 – 문서 분류 71
    • 3.5.2. 문서 진료과 분류 능력 (task 3): 평가 목표 71
    • 3.5.3. 문서 진료과 분류 능력 (task 3): 데이터 전처리 72
    • 3.5.4. 문서 진료과 분류 능력 (task 3): 입력 변수 72
    • 3.5.5. Model Architecture 73
    • 3.6. 문맥 이해 능력 평가 (task 4, task 5) 74
    • 3.6.1. 이론적 배경 – 문맥 이해 74
    • 3.6.2. 섞인 문단에서 문맥 이해 (task 4): 평가 목표 75
    • 3.6.3. 섞인 문단에서 문맥 이해 (task 4): 데이터 전처리 75
    • 3.6.4. 섞인 문단에서 문맥 이해 (task 4): 입력 변수 77
    • 3.6.5. 순차적 문단에서 문맥 이해 (task 5): 평가 목표 78
    • 3.6.6. 순차적 문단에서 문맥 이해 (task 5): 데이터 전처리 78
    • 3.6.7. 순차적 문단에서 문맥 이해 (task 5): 입력 변수 80
    • 3.6.8. Model Architectures 81
    • 3.7. 문서 시간선 추론 능력 평가 (task 6) 83
    • 3.7.1. 이론적 배경 – 임베딩 유사도 83
    • 3.7.2. 문서 시간선 추론 능력 (task 6): 평가 목표 85
    • 3.7.3. 문서 시간선 추론 능력 (task 6): 데이터 전처리 86
    • 3.7.4. 문서 시간선 추론 능력 (task 6): 입력 변수 86
    • 3.7.5. Model Architecture 88
    • 3.8. 지식 추론 능력 평가 (task 7) 89
    • 3.8.1. 이론적 배경 – 지식 추론 89
    • 3.8.2. 지식 추론 능력 (task 7): 평가 목표 90
    • 3.8.3. 지식 추론 능력 (task 7): 데이터 전처리 90
    • 3.8.4. 지식 추론 능력 (task 7): 입력 변수 91
    • 3.8.5. Model Architecture 92
    • 3.9. 실험환경 93
    • 3.9.1. 사전 학습 환경 93
    • 3.9.2. 미세조정 학습 환경 93
    • 4. 연구결과 94
    • 4.1. 유사 문서 분류 능력 평가 (task 1, task 2) 94
    • 4.2. 문서 진료과 분류 능력 평가 (task 3) 97
    • 4.3. 문맥 이해 능력 평가 (task 4, task 5) 98
    • 4.4. 문서 시간선 추론 능력 평가 (task 6) 100
    • 4.5. 지식 추론 능력 평가 (task 7) 101
    • 4.6. 연구 가설과 연구 결과 102
    • 5. 고 찰 103
    • 5.1. 모델 적용을 통한 평가방법론의 타당성 검토 103
    • 5.1.1. 달성 점수 측면의 타당성 103
    • 5.1.2. 평가 모델 구조적 측면의 타당성 105
    • 5.2. 제안한 평가 방법론을 사용한 모델 특성 분석 107
    • 5.2.1. 문서 분류 작업에서 BERT 모델의 성능 분석 (task 1∼3) 107
    • 5.2.2. 한국어 의료문서로부터 구문을 추출하는 모델의 다국어 능력의 중요성 (task 4, task 5) 109
    • 5.2.3. 한국어 의료문서로부터 구문을 추출하는 모델의 과대평가를 방지하기 위한 방법 (task 4, task 5) 112
    • 5.2.4. 문서 시간선 추론 능력 평가 시 주의점 (task 6) 113
    • 5.2.5. 모델의 지식 추론과 모델의 다국어 능력의 중요성 (task 7) 114
    • 5.2.6. 추가적인 사전학습과 미세조정 성능의 상관관계 115
    • 5.3. 디코더 기반의 대형 생성형 언어모델 116
    • 5.3.1. 최신 언어 모델과의 비교 배경 116
    • 5.3.2. Mistral 7B 117
    • 5.3.3. Low-Rank Adaptation (LoRA) 117
    • 5.3.4. 임상 문서에서 생성형 언어모델의 추론 성능 119
    • 5.3.5. 제안한 평가방법론의 생성형 언어모델에서 활용 가능성 121
    • 5.4. 입력 토큰 길이의 제한에 대한 고찰 122
    • 5.5. 본 연구의 한계점 125
    • 6. 결 론 128
    • 참고문헌 130
    • Abstract 139
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼