RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Nucleus Sampling-Based Synthetic Parallel Data Generation Method for Neural Machine Translation = 신경망 기계 번역을 위한 뉴클러스 샘플링 기반의 가상 병렬 데이터 생성 방법

    한글로보기

    https://www.riss.kr/link?id=T16594902

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    역-번역의 산물인 가상 병렬 데이터는 기계 번역기 학습에서 아주 중요하다. 하지만 가상 병렬 데이터와 인간이 태깅한 병렬 데이터 사이에는 여전히 큰 차이가 있다. 본 논문은 이러한 차이를 줄이기 위하여 번역 적절성과 번역 다양성에 초점을 맞추어 가상 병렬 데이터를 분석하였다. 가상 병렬 데이터의 번역 적절성은 다중 문장-임베딩 모델을 이용하여 병렬 문장 사이의 의미적 유사도에 근거하여 측정하였다. 또한 가상 병렬 데이터의 번역 다양성은 저-빈도 단어수와 미등록 단어의 비율에 근거하여 분석하였다. 본 논문의 분석에 의하면 역-번역 디코딩 단계의 beam search로 인하여 발생되는 다양성 부족이 가상 병렬 데이터의 가장 주요한 문제이다. 따라서 본 논문은 가상 병렬 데이터 생성 시 뉴클러스 샘플링 방법으로 beam-search 디코딩을 대체하여 다양성 부족 문제를 해결하였다. 실험 결과에 의하면 뉴클러스 샘플링에 기반한 디코딩이 beam-search 디코딩에 비하여 가상 병렬 데이터의 미등록 단어 비율을 13.27%에서 4.0%까지 낮추었다. 도메인 외 (out-domain) 번역에서 뉴클러스 샘플링에 기반한 방법이 beam search에 비해 0.51 BLEU (Bilingual Evaluation Understudy) 향상된 성능을 보였다. 또한 가상 병렬 데이터 필터링을 적용한 결과 0.18 BLEU의 추가 성능 향상을 관찰하였다. 제안 방법을 중간-자원 (medium-resourced) 도메인 내 (in-domain) 번역에 적용한 결과 0.96 - 1.23 BLEU 성능 향샹을 보였다. 제안 방법을 역-번역을 탑재한 최신 사전 학습된 (pre-training) 언어 모델 기반 기계 번역기에 적용한 결과 약간의 성능 향상을 보였다. 이러한 결과는 뉴클러스 샘플링에 기반한 디코딩 방법이 아주 풍부하고 다양한 가상 병렬 데이터를 생성할 수 있고 궁극적으로 기계 번역기의 성능을 향상할 수 있음을 시사한다.
    번역하기

    역-번역의 산물인 가상 병렬 데이터는 기계 번역기 학습에서 아주 중요하다. 하지만 가상 병렬 데이터와 인간이 태깅한 병렬 데이터 사이에는 여전히 큰 차이가 있다. 본 논문은 이러한 차이...

    역-번역의 산물인 가상 병렬 데이터는 기계 번역기 학습에서 아주 중요하다. 하지만 가상 병렬 데이터와 인간이 태깅한 병렬 데이터 사이에는 여전히 큰 차이가 있다. 본 논문은 이러한 차이를 줄이기 위하여 번역 적절성과 번역 다양성에 초점을 맞추어 가상 병렬 데이터를 분석하였다. 가상 병렬 데이터의 번역 적절성은 다중 문장-임베딩 모델을 이용하여 병렬 문장 사이의 의미적 유사도에 근거하여 측정하였다. 또한 가상 병렬 데이터의 번역 다양성은 저-빈도 단어수와 미등록 단어의 비율에 근거하여 분석하였다. 본 논문의 분석에 의하면 역-번역 디코딩 단계의 beam search로 인하여 발생되는 다양성 부족이 가상 병렬 데이터의 가장 주요한 문제이다. 따라서 본 논문은 가상 병렬 데이터 생성 시 뉴클러스 샘플링 방법으로 beam-search 디코딩을 대체하여 다양성 부족 문제를 해결하였다. 실험 결과에 의하면 뉴클러스 샘플링에 기반한 디코딩이 beam-search 디코딩에 비하여 가상 병렬 데이터의 미등록 단어 비율을 13.27%에서 4.0%까지 낮추었다. 도메인 외 (out-domain) 번역에서 뉴클러스 샘플링에 기반한 방법이 beam search에 비해 0.51 BLEU (Bilingual Evaluation Understudy) 향상된 성능을 보였다. 또한 가상 병렬 데이터 필터링을 적용한 결과 0.18 BLEU의 추가 성능 향상을 관찰하였다. 제안 방법을 중간-자원 (medium-resourced) 도메인 내 (in-domain) 번역에 적용한 결과 0.96 - 1.23 BLEU 성능 향샹을 보였다. 제안 방법을 역-번역을 탑재한 최신 사전 학습된 (pre-training) 언어 모델 기반 기계 번역기에 적용한 결과 약간의 성능 향상을 보였다. 이러한 결과는 뉴클러스 샘플링에 기반한 디코딩 방법이 아주 풍부하고 다양한 가상 병렬 데이터를 생성할 수 있고 궁극적으로 기계 번역기의 성능을 향상할 수 있음을 시사한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Synthetic data generated by back-translation is crucial in training neural machine translation (NMT) systems. While synthetic data has been shown to be effective, there is still a big gap between synthetic data and real data that is annotated by human beings. This thesis focuses on two aspects of synthetic data: translation adequacy and diversity. We measure the translation adequacy according to the semantic similarities of sentence pairs in synthetic data calculated by a multilingual sentenceembedding model. Moreover, we analyze the translation diversity considering the distribution of the number of low-frequency words and the out-of-vocabulary rate in synthetic data. Our analysis demonstrates that the lack of diversity and richness problem inherited from beam search in the decoding phase is the primary issue of synthetic data. Therefore, we propose using nucleus sampling-based decoding strategy as an alternative to beam-search decoding in back-translation which significantly improves the diversity of synthetic data. The experimental results demonstrate that nucleus sampling-based decoding lowers the out-of-vocabulary rate of synthetic data to 4.0% compared to 13.27% for beam search. In out-domain translation tasks, synthetic data generated by the sampling method outperforms the generated via beam search by 0.51 BLEU (Bilingual Evaluation Understudy) score. Furthermore, we observe an additional gain of 0.18 BLEU by adding synthetic data filtering. Synthetic data generated by the nucleus sampling method outperforms beam search by 0.96 − 1.23 BLEU in medium-resourced in-domain translation tasks. By applying the proposed methods to the recently advanced pretraining model with back-translation, we achieve a slight performance boost. The study indicates that nucleus samplingbased decoding is essential for generating a rich and diverse synthetic parallel data which improves the translation performance of an NMT system.
    번역하기

    Synthetic data generated by back-translation is crucial in training neural machine translation (NMT) systems. While synthetic data has been shown to be effective, there is still a big gap between synthetic data and real data that is annotated by human...

    Synthetic data generated by back-translation is crucial in training neural machine translation (NMT) systems. While synthetic data has been shown to be effective, there is still a big gap between synthetic data and real data that is annotated by human beings. This thesis focuses on two aspects of synthetic data: translation adequacy and diversity. We measure the translation adequacy according to the semantic similarities of sentence pairs in synthetic data calculated by a multilingual sentenceembedding model. Moreover, we analyze the translation diversity considering the distribution of the number of low-frequency words and the out-of-vocabulary rate in synthetic data. Our analysis demonstrates that the lack of diversity and richness problem inherited from beam search in the decoding phase is the primary issue of synthetic data. Therefore, we propose using nucleus sampling-based decoding strategy as an alternative to beam-search decoding in back-translation which significantly improves the diversity of synthetic data. The experimental results demonstrate that nucleus sampling-based decoding lowers the out-of-vocabulary rate of synthetic data to 4.0% compared to 13.27% for beam search. In out-domain translation tasks, synthetic data generated by the sampling method outperforms the generated via beam search by 0.51 BLEU (Bilingual Evaluation Understudy) score. Furthermore, we observe an additional gain of 0.18 BLEU by adding synthetic data filtering. Synthetic data generated by the nucleus sampling method outperforms beam search by 0.96 − 1.23 BLEU in medium-resourced in-domain translation tasks. By applying the proposed methods to the recently advanced pretraining model with back-translation, we achieve a slight performance boost. The study indicates that nucleus samplingbased decoding is essential for generating a rich and diverse synthetic parallel data which improves the translation performance of an NMT system.

    더보기

    참고문헌 (Reference)

    1. Hierarchical neural story generation, A . Fan , M. Lewis , and Y. Dauphin, pp . 889 ? 898 . doi : 10.18653/v1/P18-1082 ., , 2018

    2. Six challenges for neural machine translation, P. Koehn and R. Knowles ,, pp . 28 ? 39 . doi : 10.18653/ v1/W17-3204 ., , 2017

    3. Sequence to sequence learning with neural networks, I. Sutskever , O. Vinyals , and Q. V. LeZ. Ghahramani , M. Welling , C. Cortes , N. D. Lawrence , and K. Q. Weinberger , Eds. , Curran Associates , Inc., pp . 3104 ? 3112 ., , 2014

    4. Analyzing uncertainty in neural machine translation, Ott , M. Auli , D. Grangier , and M. Ranzato ,, PMLRpp . 3956 ? 3965, , 2018

    5. Language models are unsupervised multitask learners, Radford , J. Wu , R. Child , D. Luan , D. Amodei , I. Sutskever , et al., vol . 1 , no . 8 , p. 9, , 2019

    6. Neural machine translation of rare words with subword units ,, R. Sennrich , B. Haddow , and A. Birch ,, pp . 1715 ? 1725 . doi : 10.18653/v1/P16-1162 ., , 2016

    7. Bilingual data cleaning for SMT using graph-based random walk ,, L. Cui , D. Zhang , S. Liu , M. Li , and M. Zhou, pp . 340 ? 345 ., , 2013

    8. Bilingual word representations with monolingual quality in mind, T. Luong , H. Pham , and C. D. Manning, pp . 151 ? 159, , 2015

    9. On integrating a language model into neural machine translation, C. Gulcehre , O. Firat , K. Xu , K. Cho , and Y. Bengio, vol . 45 , no . C , pp . 137 ? 148 , Sep., , 2017

    10. Improving neural machine translation models with monolingual data, R. Sennrich , B. Haddow , and A. Birch ,, pp . 86 ? 96 . doi : 10.18653/v1/P16-1009 ., , 2016

    1. Hierarchical neural story generation, A . Fan , M. Lewis , and Y. Dauphin, pp . 889 ? 898 . doi : 10.18653/v1/P18-1082 ., , 2018

    2. Six challenges for neural machine translation, P. Koehn and R. Knowles ,, pp . 28 ? 39 . doi : 10.18653/ v1/W17-3204 ., , 2017

    3. Sequence to sequence learning with neural networks, I. Sutskever , O. Vinyals , and Q. V. LeZ. Ghahramani , M. Welling , C. Cortes , N. D. Lawrence , and K. Q. Weinberger , Eds. , Curran Associates , Inc., pp . 3104 ? 3112 ., , 2014

    4. Analyzing uncertainty in neural machine translation, Ott , M. Auli , D. Grangier , and M. Ranzato ,, PMLRpp . 3956 ? 3965, , 2018

    5. Language models are unsupervised multitask learners, Radford , J. Wu , R. Child , D. Luan , D. Amodei , I. Sutskever , et al., vol . 1 , no . 8 , p. 9, , 2019

    6. Neural machine translation of rare words with subword units ,, R. Sennrich , B. Haddow , and A. Birch ,, pp . 1715 ? 1725 . doi : 10.18653/v1/P16-1162 ., , 2016

    7. Bilingual data cleaning for SMT using graph-based random walk ,, L. Cui , D. Zhang , S. Liu , M. Li , and M. Zhou, pp . 340 ? 345 ., , 2013

    8. Bilingual word representations with monolingual quality in mind, T. Luong , H. Pham , and C. D. Manning, pp . 151 ? 159, , 2015

    9. On integrating a language model into neural machine translation, C. Gulcehre , O. Firat , K. Xu , K. Cho , and Y. Bengio, vol . 45 , no . C , pp . 137 ? 148 , Sep., , 2017

    10. Improving neural machine translation models with monolingual data, R. Sennrich , B. Haddow , and A. Birch ,, pp . 86 ? 96 . doi : 10.18653/v1/P16-1009 ., , 2016

    11. Dropout : A simple way to prevent neural networks from overfitting, N. Srivastava , G. Hinton , A. Krizhevsky , I. Sutskever , and R. Salakhutdinov, vol . 15 , pp . 1929 ? 1958, , 2014

    12. Investigations on translation model adaptation using monolingual data, P. Lambert , H. Schwenk , C. Servan , and S. Abdul-Rauf ,, pp . 284 ? 293 ., , 2011

    13. Neural machine translation for low-resource languages without parallel corpora, A. Karakanta , J. Dehdari , and J. van Genabith, vol . 32 , no . 1 , pp . 167 ? 189, , 2018

    14. Normalized word embedding and orthogonal transform for bilingual word translation, C. Xing , D. Wang , C. Liu , and Y. Lin, pp . 1006 ? 1011, , 2015

    15. Improving low-resource neural machine translation with filtered pseudo-parallel corpus ,, A. Imankulova , T. Sato , and M. Komachi, pp . 70 ? 78, , 2017

    16. SentencePiece : A simple and language independent subword tokenizer and detokenizer for neural text processing ,, T. Kudo and J. Richardson, pp . 66 ? 71 . doi : 10.18653/v1/D18-2012 ., , 2018

    17. BART : Denoising sequence-to-sequence pre-training for natural language generation , translation , and comprehension, Lewis , Y. Liu , N. Goyal , et al., pp . 7871 ? 7880 . doi : 10.18653/v1/2020.acl-main.703 ., , 2020

    18. SQuAD : 100,000+ questions for machine comprehension of textin Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, P. Rajpurkar , J. Zhang , K. Lopyrev , and P. Liang ,, pp . 2383 ? 2392 . doi : 10.18653/v1/D16- 1264 ., , 2016

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼