본 논문의 주요 목표는 스토리텔링 텍스트 생성을 위한 최적의 모델을 발견하기 위해 책 장르에서 지정된 데이터를 사용하여 훈련된 세 가지 구별되는 모델을 비교하는 것이다. Diana Gabaldon...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T16525988
부산: 신라대학교 일반대학원: 신라대학교, 2022
학위논문(석사) -- 신라대학교 일반대학원 , 융합공학과 컴퓨터정보공학전공 , 2022. 8
2022
영어
004.735 판사항(5)
부산
특정 장르의 스토리텔링 창작을 위한 NLP 모델의비교
ⅴⅰ, 66장: 삽화, 도표; 26 cm.
지도교수: 김병기
참고문헌 수록
I804:21020-200000646968
0
상세조회0
다운로드본 논문의 주요 목표는 스토리텔링 텍스트 생성을 위한 최적의 모델을 발견하기 위해 책 장르에서 지정된 데이터를 사용하여 훈련된 세 가지 구별되는 모델을 비교하는 것이다. Diana Gabaldon...
본 논문의 주요 목표는 스토리텔링 텍스트 생성을 위한 최적의 모델을 발견하기 위해 책 장르에서 지정된 데이터를 사용하여 훈련된 세 가지 구별되는 모델을 비교하는 것이다. Diana Gabaldon의 로맨스 소설 “Outlander”와 Kiera Cass의 “Happily ever after”는 RNN 모델을 GRU 레이어로 훈련시키고 GPT-2 모델 124M 파라미터를 미세 조정하기 위해 선택되었다. 그러나 처음부터 완전히 훈련된 GPT-2 모델의 경우 100개의 로맨스 소설이 데이터 세트로 사용되었다. RNN 모델의 훈련 과정을 더 쉽게 만들기 위해 문자 대 정수 부호화가 사용되었다. 반면에 GPT-2 모델은 처음부터 구축되었으며 라이브러리 “Tokenizers”는 데이터를 토큰으로 인코딩하는 데 사용되었으며, 그런 다음 이를 훈련하고 출력 텍스트로 디코딩할 수 있었다. 클라우드 기반 타이핑 어시스턴트인 "Grammarly"는 문서를 비교하기 위해 사용되었으며, 문법 오류, 어휘 다양성, 흔하지 않은 영어 단어의 발견을 강조하였다. 또한 생성된 텍스트를 읽는 인간의 의견을 아는 것은 매우 중요하므로 Jose Villalobos의 평가 기준을 인간 분석의 측정 매개 변수로 작용했다. 이 논문 결과는 스토리텔링 생성을 위한 최상의 모델이 124M 매개 변수의 GPT-2이다.
다국어 초록 (Multilingual Abstract)
The goal of this work is to compare three distinct models trained using data specified by a book genre in order to discover the optimal model for the generation of storytelling texts. The romance novels "Outlander" by Diana Gabaldon and "Happily Ever ...
The goal of this work is to compare three distinct models trained using data specified by a book genre in order to discover the optimal model for the generation of storytelling texts. The romance novels "Outlander" by Diana Gabaldon and "Happily Ever After" by Kiera Cass were chosen to train the RNN model with a GRU layer, as well as to finetune the GPT-2 model of 124M parameters. However, for the GPT-2 model, which was trained from scratch, an amount of one hundred romance novels were used as the dataset. To make the training process easier for the RNN model, a character to integer codification was used. The GPT-2 model, on the other hand, was built from the ground up, and the library tokenizer was used to encode the data into tokens, which could then be trained and decoded into the output text. To compare the generated texts of the models, the cloud-based typing assistant "Grammarly" was used to find the highlighting grammar errors, vocabulary diversity, the discovery of uncommon English words and, the overall score. In addition, the opinion of a human reading the generated texts is critical, therefore, Jose Villalobos' assessment criteria served as the human analysis, measuring the coherence, cohesion and adequacy of the generated text. The findings yield fascinating results, proving that the best model for storytelling creation is the GPT-2 model of 124M parameters.
목차 (Table of Contents)