RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    역번역 기반 적응 사전 학습을 통한 문서 분류 성능 및 강건성 향상

    한글로보기

    https://www.riss.kr/link?id=T15944570

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Language models (LMs) pretrained on a large text corpus and fine-tuned on a downstream text corpus and fine-tuned on a downstream task becomes a de facto training strategy for several natural language processing (NLP) tasks. Recently, an adaptive pretraining method retraining the pretrained language model with task-relevant data has shown significant performance improvements. However, current adaptive pretraining methods suffer from underfitting on the task distribution owing to a relatively small amount of data to re-pretrain the LM. To completely use the concept of adaptive pretraining, we propose a back-translated task-adaptive pretraining (BT-TAPT) method that increases the amount of task-specific data for LM re-pretraining by augmenting the task data using back-translation to generalize the LM to the target task domain. The experimental results show that the proposed BT-TAPT yields improved classification accuracy on both low- and high-resource data and better robustness to noise than the conventional adaptive pretraining method.
    번역하기

    Language models (LMs) pretrained on a large text corpus and fine-tuned on a downstream text corpus and fine-tuned on a downstream task becomes a de facto training strategy for several natural language processing (NLP) tasks. Recently, an adaptive pret...

    Language models (LMs) pretrained on a large text corpus and fine-tuned on a downstream text corpus and fine-tuned on a downstream task becomes a de facto training strategy for several natural language processing (NLP) tasks. Recently, an adaptive pretraining method retraining the pretrained language model with task-relevant data has shown significant performance improvements. However, current adaptive pretraining methods suffer from underfitting on the task distribution owing to a relatively small amount of data to re-pretrain the LM. To completely use the concept of adaptive pretraining, we propose a back-translated task-adaptive pretraining (BT-TAPT) method that increases the amount of task-specific data for LM re-pretraining by augmenting the task data using back-translation to generalize the LM to the target task domain. The experimental results show that the proposed BT-TAPT yields improved classification accuracy on both low- and high-resource data and better robustness to noise than the conventional adaptive pretraining method.

    더보기

    목차 (Table of Contents)

    • 1 서론 1
    • 2 선행 연구 6
    • 2.1 마스크 기반 언어 모델(Masked Language Model) 6
    • 2.2 적응 사전 학습(Adaptive Pretraining) 7
    • 2.3 텍스트 데이터 증강(Text Data Augmentation) 8
    • 1 서론 1
    • 2 선행 연구 6
    • 2.1 마스크 기반 언어 모델(Masked Language Model) 6
    • 2.2 적응 사전 학습(Adaptive Pretraining) 7
    • 2.3 텍스트 데이터 증강(Text Data Augmentation) 8
    • 3 방법론 10
    • 3.1 불충분한 과업 데이터 11
    • 3.2 역번역 기반 적응 사전 학습 11
    • 3.3 역번역 기반 적응 사전 학습 절차 13
    • 4 실험 및 결과 14
    • 4.1 데이터 셋 및 평가 지표 14
    • 4.2 실험 환경 15
    • 4.3 실험 결과 17
    • 4.3.1 문서 분류 성능 17
    • 4.3.2 증강 방법 비교 19
    • 4.3.3 증강 수량 비교 21
    • 4.3.4 역번역 기반 적응 사전 학습 방식 비교 22
    • 4.3.5 적은 데이터 셋에 대한 성능 비교 23
    • 4.4 노이즈에 대한 강건성 24
    • 4.4.1 노이즈 종류 25
    • 4.4.2 실험 결과 26
    • 5 결론 29
    • 참고문헌 30
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼