RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    NLP Model Comparison for storytelling creation of specific genre

    한글로보기

    https://www.riss.kr/link?id=T16525988

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문의 주요 목표는 스토리텔링 텍스트 생성을 위한 최적의 모델을 발견하기 위해 책 장르에서 지정된 데이터를 사용하여 훈련된 세 가지 구별되는 모델을 비교하는 것이다. Diana Gabaldon의 로맨스 소설 “Outlander”와 Kiera Cass의 “Happily ever after”는 RNN 모델을 GRU 레이어로 훈련시키고 GPT-2 모델 124M 파라미터를 미세 조정하기 위해 선택되었다. 그러나 처음부터 완전히 훈련된 GPT-2 모델의 경우 100개의 로맨스 소설이 데이터 세트로 사용되었다. RNN 모델의 훈련 과정을 더 쉽게 만들기 위해 문자 대 정수 부호화가 사용되었다. 반면에 GPT-2 모델은 처음부터 구축되었으며 라이브러리 “Tokenizers”는 데이터를 토큰으로 인코딩하는 데 사용되었으며, 그런 다음 이를 훈련하고 출력 텍스트로 디코딩할 수 있었다. 클라우드 기반 타이핑 어시스턴트인 "Grammarly"는 문서를 비교하기 위해 사용되었으며, 문법 오류, 어휘 다양성, 흔하지 않은 영어 단어의 발견을 강조하였다. 또한 생성된 텍스트를 읽는 인간의 의견을 아는 것은 매우 중요하므로 Jose Villalobos의 평가 기준을 인간 분석의 측정 매개 변수로 작용했다. 이 논문 결과는 스토리텔링 생성을 위한 최상의 모델이 124M 매개 변수의 GPT-2이다.
    번역하기

    본 논문의 주요 목표는 스토리텔링 텍스트 생성을 위한 최적의 모델을 발견하기 위해 책 장르에서 지정된 데이터를 사용하여 훈련된 세 가지 구별되는 모델을 비교하는 것이다. Diana Gabaldon...

    본 논문의 주요 목표는 스토리텔링 텍스트 생성을 위한 최적의 모델을 발견하기 위해 책 장르에서 지정된 데이터를 사용하여 훈련된 세 가지 구별되는 모델을 비교하는 것이다. Diana Gabaldon의 로맨스 소설 “Outlander”와 Kiera Cass의 “Happily ever after”는 RNN 모델을 GRU 레이어로 훈련시키고 GPT-2 모델 124M 파라미터를 미세 조정하기 위해 선택되었다. 그러나 처음부터 완전히 훈련된 GPT-2 모델의 경우 100개의 로맨스 소설이 데이터 세트로 사용되었다. RNN 모델의 훈련 과정을 더 쉽게 만들기 위해 문자 대 정수 부호화가 사용되었다. 반면에 GPT-2 모델은 처음부터 구축되었으며 라이브러리 “Tokenizers”는 데이터를 토큰으로 인코딩하는 데 사용되었으며, 그런 다음 이를 훈련하고 출력 텍스트로 디코딩할 수 있었다. 클라우드 기반 타이핑 어시스턴트인 "Grammarly"는 문서를 비교하기 위해 사용되었으며, 문법 오류, 어휘 다양성, 흔하지 않은 영어 단어의 발견을 강조하였다. 또한 생성된 텍스트를 읽는 인간의 의견을 아는 것은 매우 중요하므로 Jose Villalobos의 평가 기준을 인간 분석의 측정 매개 변수로 작용했다. 이 논문 결과는 스토리텔링 생성을 위한 최상의 모델이 124M 매개 변수의 GPT-2이다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The goal of this work is to compare three distinct models trained using data specified by a book genre in order to discover the optimal model for the generation of storytelling texts. The romance novels "Outlander" by Diana Gabaldon and "Happily Ever After" by Kiera Cass were chosen to train the RNN model with a GRU layer, as well as to finetune the GPT-2 model of 124M parameters. However, for the GPT-2 model, which was trained from scratch, an amount of one hundred romance novels were used as the dataset. To make the training process easier for the RNN model, a character to integer codification was used. The GPT-2 model, on the other hand, was built from the ground up, and the library tokenizer was used to encode the data into tokens, which could then be trained and decoded into the output text. To compare the generated texts of the models, the cloud-based typing assistant "Grammarly" was used to find the highlighting grammar errors, vocabulary diversity, the discovery of uncommon English words and, the overall score. In addition, the opinion of a human reading the generated texts is critical, therefore, Jose Villalobos' assessment criteria served as the human analysis, measuring the coherence, cohesion and adequacy of the generated text. The findings yield fascinating results, proving that the best model for storytelling creation is the GPT-2 model of 124M parameters.
    번역하기

    The goal of this work is to compare three distinct models trained using data specified by a book genre in order to discover the optimal model for the generation of storytelling texts. The romance novels "Outlander" by Diana Gabaldon and "Happily Ever ...

    The goal of this work is to compare three distinct models trained using data specified by a book genre in order to discover the optimal model for the generation of storytelling texts. The romance novels "Outlander" by Diana Gabaldon and "Happily Ever After" by Kiera Cass were chosen to train the RNN model with a GRU layer, as well as to finetune the GPT-2 model of 124M parameters. However, for the GPT-2 model, which was trained from scratch, an amount of one hundred romance novels were used as the dataset. To make the training process easier for the RNN model, a character to integer codification was used. The GPT-2 model, on the other hand, was built from the ground up, and the library tokenizer was used to encode the data into tokens, which could then be trained and decoded into the output text. To compare the generated texts of the models, the cloud-based typing assistant "Grammarly" was used to find the highlighting grammar errors, vocabulary diversity, the discovery of uncommon English words and, the overall score. In addition, the opinion of a human reading the generated texts is critical, therefore, Jose Villalobos' assessment criteria served as the human analysis, measuring the coherence, cohesion and adequacy of the generated text. The findings yield fascinating results, proving that the best model for storytelling creation is the GPT-2 model of 124M parameters.

    더보기

    목차 (Table of Contents)

    • CONTENTS
    • LIST OF TABLES ⅳ
    • LIST OF FIGURES ⅴ
    • 1. Introduction 1
    • 1.1. Motivation 1
    • CONTENTS
    • LIST OF TABLES ⅳ
    • LIST OF FIGURES ⅴ
    • 1. Introduction 1
    • 1.1. Motivation 1
    • 1.2. Current methods for text generation 2
    • 1.3. Related works 4
    • 2. IDE and Natural Language Processing (NLP) 7
    • 2.1. Computational Linguistics and Natural Language Processing
    • within Artificial Intelligence 7
    • 2.2. Hardware Environment 9
    • 2.3. Software Environment 9
    • 3. Recurrent Neuronal Networks (RNN) & GTP-2 12
    • 3.1. Understanding RNNs 12
    • 3.1.1. Long Short-Term Memory 13
    • 3.1.2. Gated Recurrent Units 15
    • 3.2. GPT-2 17
    • 3.2.1. Transformers 19
    • 4. Hypothesis and Training Methods 21
    • 4.1. Hypothesis 21
    • 4.2. Methodology 22
    • 4.3. Training Methods 23
    • 4.3.1. RNN model using a GRU layer 23
    • 4.3.2. GPT-2 trained from scratch with specific book genre texts 26
    • 4.3.3. Fine tuning with specific book genre text the GPT- 2models of 124 million parameters 29
    • 5. Generated Text Results and Discussion 30
    • 5.1. Evaluation Criteria 31
    • 5.2. RNN model using a GRU layer 34
    • 5.3. GPT-2 trained from scratch with specific book genre texts. 35
    • 5.4. Fine tuning with specific book genre text using the GPT-2 model of 124 million parameters 36
    • 5.5. Text Measurement and Discussion 38
    • 5.6. Comparison Results 59
    • 6. Conclusions 62
    • References 63
    • ABSTRACT 66
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼