역-번역의 산물인 가상 병렬 데이터는 기계 번역기 학습에서 아주 중요하다. 하지만 가상 병렬 데이터와 인간이 태깅한 병렬 데이터 사이에는 여전히 큰 차이가 있다. 본 논문은 이러한 차이...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T16594902
서울 : 서강대학교 일반대학원, 2023
학위논문(박사) -- 서강대학교 일반대학원 , 컴퓨터공학과 , 2023. 2
2023
영어
서울
viii, 64 p. : ill. ; 26 cm.
지도교수: 서정연
I804:11029-000000070206
0
상세조회0
다운로드역-번역의 산물인 가상 병렬 데이터는 기계 번역기 학습에서 아주 중요하다. 하지만 가상 병렬 데이터와 인간이 태깅한 병렬 데이터 사이에는 여전히 큰 차이가 있다. 본 논문은 이러한 차이...
역-번역의 산물인 가상 병렬 데이터는 기계 번역기 학습에서 아주 중요하다. 하지만 가상 병렬 데이터와 인간이 태깅한 병렬 데이터 사이에는 여전히 큰 차이가 있다. 본 논문은 이러한 차이를 줄이기 위하여 번역 적절성과 번역 다양성에 초점을 맞추어 가상 병렬 데이터를 분석하였다. 가상 병렬 데이터의 번역 적절성은 다중 문장-임베딩 모델을 이용하여 병렬 문장 사이의 의미적 유사도에 근거하여 측정하였다. 또한 가상 병렬 데이터의 번역 다양성은 저-빈도 단어수와 미등록 단어의 비율에 근거하여 분석하였다. 본 논문의 분석에 의하면 역-번역 디코딩 단계의 beam search로 인하여 발생되는 다양성 부족이 가상 병렬 데이터의 가장 주요한 문제이다. 따라서 본 논문은 가상 병렬 데이터 생성 시 뉴클러스 샘플링 방법으로 beam-search 디코딩을 대체하여 다양성 부족 문제를 해결하였다. 실험 결과에 의하면 뉴클러스 샘플링에 기반한 디코딩이 beam-search 디코딩에 비하여 가상 병렬 데이터의 미등록 단어 비율을 13.27%에서 4.0%까지 낮추었다. 도메인 외 (out-domain) 번역에서 뉴클러스 샘플링에 기반한 방법이 beam search에 비해 0.51 BLEU (Bilingual Evaluation Understudy) 향상된 성능을 보였다. 또한 가상 병렬 데이터 필터링을 적용한 결과 0.18 BLEU의 추가 성능 향상을 관찰하였다. 제안 방법을 중간-자원 (medium-resourced) 도메인 내 (in-domain) 번역에 적용한 결과 0.96 - 1.23 BLEU 성능 향샹을 보였다. 제안 방법을 역-번역을 탑재한 최신 사전 학습된 (pre-training) 언어 모델 기반 기계 번역기에 적용한 결과 약간의 성능 향상을 보였다. 이러한 결과는 뉴클러스 샘플링에 기반한 디코딩 방법이 아주 풍부하고 다양한 가상 병렬 데이터를 생성할 수 있고 궁극적으로 기계 번역기의 성능을 향상할 수 있음을 시사한다.
다국어 초록 (Multilingual Abstract)
Synthetic data generated by back-translation is crucial in training neural machine translation (NMT) systems. While synthetic data has been shown to be effective, there is still a big gap between synthetic data and real data that is annotated by human...
Synthetic data generated by back-translation is crucial in training neural machine translation (NMT) systems. While synthetic data has been shown to be effective, there is still a big gap between synthetic data and real data that is annotated by human beings. This thesis focuses on two aspects of synthetic data: translation adequacy and diversity. We measure the translation adequacy according to the semantic similarities of sentence pairs in synthetic data calculated by a multilingual sentenceembedding model. Moreover, we analyze the translation diversity considering the distribution of the number of low-frequency words and the out-of-vocabulary rate in synthetic data. Our analysis demonstrates that the lack of diversity and richness problem inherited from beam search in the decoding phase is the primary issue of synthetic data. Therefore, we propose using nucleus sampling-based decoding strategy as an alternative to beam-search decoding in back-translation which significantly improves the diversity of synthetic data. The experimental results demonstrate that nucleus sampling-based decoding lowers the out-of-vocabulary rate of synthetic data to 4.0% compared to 13.27% for beam search. In out-domain translation tasks, synthetic data generated by the sampling method outperforms the generated via beam search by 0.51 BLEU (Bilingual Evaluation Understudy) score. Furthermore, we observe an additional gain of 0.18 BLEU by adding synthetic data filtering. Synthetic data generated by the nucleus sampling method outperforms beam search by 0.96 − 1.23 BLEU in medium-resourced in-domain translation tasks. By applying the proposed methods to the recently advanced pretraining model with back-translation, we achieve a slight performance boost. The study indicates that nucleus samplingbased decoding is essential for generating a rich and diverse synthetic parallel data which improves the translation performance of an NMT system.
참고문헌 (Reference)
1. Hierarchical neural story generation, A . Fan , M. Lewis , and Y. Dauphin, pp . 889 ? 898 . doi : 10.18653/v1/P18-1082 ., , 2018
2. Six challenges for neural machine translation, P. Koehn and R. Knowles ,, pp . 28 ? 39 . doi : 10.18653/ v1/W17-3204 ., , 2017
3. Sequence to sequence learning with neural networks, I. Sutskever , O. Vinyals , and Q. V. LeZ. Ghahramani , M. Welling , C. Cortes , N. D. Lawrence , and K. Q. Weinberger , Eds. , Curran Associates , Inc., pp . 3104 ? 3112 ., , 2014
4. Analyzing uncertainty in neural machine translation, Ott , M. Auli , D. Grangier , and M. Ranzato ,, PMLRpp . 3956 ? 3965, , 2018
5. Language models are unsupervised multitask learners, Radford , J. Wu , R. Child , D. Luan , D. Amodei , I. Sutskever , et al., vol . 1 , no . 8 , p. 9, , 2019
6. Neural machine translation of rare words with subword units ,, R. Sennrich , B. Haddow , and A. Birch ,, pp . 1715 ? 1725 . doi : 10.18653/v1/P16-1162 ., , 2016
7. Bilingual data cleaning for SMT using graph-based random walk ,, L. Cui , D. Zhang , S. Liu , M. Li , and M. Zhou, pp . 340 ? 345 ., , 2013
8. Bilingual word representations with monolingual quality in mind, T. Luong , H. Pham , and C. D. Manning, pp . 151 ? 159, , 2015
9. On integrating a language model into neural machine translation, C. Gulcehre , O. Firat , K. Xu , K. Cho , and Y. Bengio, vol . 45 , no . C , pp . 137 ? 148 , Sep., , 2017
10. Improving neural machine translation models with monolingual data, R. Sennrich , B. Haddow , and A. Birch ,, pp . 86 ? 96 . doi : 10.18653/v1/P16-1009 ., , 2016
1. Hierarchical neural story generation, A . Fan , M. Lewis , and Y. Dauphin, pp . 889 ? 898 . doi : 10.18653/v1/P18-1082 ., , 2018
2. Six challenges for neural machine translation, P. Koehn and R. Knowles ,, pp . 28 ? 39 . doi : 10.18653/ v1/W17-3204 ., , 2017
3. Sequence to sequence learning with neural networks, I. Sutskever , O. Vinyals , and Q. V. LeZ. Ghahramani , M. Welling , C. Cortes , N. D. Lawrence , and K. Q. Weinberger , Eds. , Curran Associates , Inc., pp . 3104 ? 3112 ., , 2014
4. Analyzing uncertainty in neural machine translation, Ott , M. Auli , D. Grangier , and M. Ranzato ,, PMLRpp . 3956 ? 3965, , 2018
5. Language models are unsupervised multitask learners, Radford , J. Wu , R. Child , D. Luan , D. Amodei , I. Sutskever , et al., vol . 1 , no . 8 , p. 9, , 2019
6. Neural machine translation of rare words with subword units ,, R. Sennrich , B. Haddow , and A. Birch ,, pp . 1715 ? 1725 . doi : 10.18653/v1/P16-1162 ., , 2016
7. Bilingual data cleaning for SMT using graph-based random walk ,, L. Cui , D. Zhang , S. Liu , M. Li , and M. Zhou, pp . 340 ? 345 ., , 2013
8. Bilingual word representations with monolingual quality in mind, T. Luong , H. Pham , and C. D. Manning, pp . 151 ? 159, , 2015
9. On integrating a language model into neural machine translation, C. Gulcehre , O. Firat , K. Xu , K. Cho , and Y. Bengio, vol . 45 , no . C , pp . 137 ? 148 , Sep., , 2017
10. Improving neural machine translation models with monolingual data, R. Sennrich , B. Haddow , and A. Birch ,, pp . 86 ? 96 . doi : 10.18653/v1/P16-1009 ., , 2016
11. Dropout : A simple way to prevent neural networks from overfitting, N. Srivastava , G. Hinton , A. Krizhevsky , I. Sutskever , and R. Salakhutdinov, vol . 15 , pp . 1929 ? 1958, , 2014
12. Investigations on translation model adaptation using monolingual data, P. Lambert , H. Schwenk , C. Servan , and S. Abdul-Rauf ,, pp . 284 ? 293 ., , 2011
13. Neural machine translation for low-resource languages without parallel corpora, A. Karakanta , J. Dehdari , and J. van Genabith, vol . 32 , no . 1 , pp . 167 ? 189, , 2018
14. Normalized word embedding and orthogonal transform for bilingual word translation, C. Xing , D. Wang , C. Liu , and Y. Lin, pp . 1006 ? 1011, , 2015
15. Improving low-resource neural machine translation with filtered pseudo-parallel corpus ,, A. Imankulova , T. Sato , and M. Komachi, pp . 70 ? 78, , 2017
16. SentencePiece : A simple and language independent subword tokenizer and detokenizer for neural text processing ,, T. Kudo and J. Richardson, pp . 66 ? 71 . doi : 10.18653/v1/D18-2012 ., , 2018
17. BART : Denoising sequence-to-sequence pre-training for natural language generation , translation , and comprehension, Lewis , Y. Liu , N. Goyal , et al., pp . 7871 ? 7880 . doi : 10.18653/v1/2020.acl-main.703 ., , 2020
18. SQuAD : 100,000+ questions for machine comprehension of textin Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, P. Rajpurkar , J. Zhang , K. Lopyrev , and P. Liang ,, pp . 2383 ? 2392 . doi : 10.18653/v1/D16- 1264 ., , 2016