기계번역은 사람의 관여 없이 자동으로 한 언어에서 다른 언어로 텍스트를 번역하기 위해서 계산을 사용하는 과정이다. 신경망을 사용하는 현재 기계번역 시스템의 용이성과 발전성 외에도...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17011635
부산 : 부산대학교 대학원, 2024
학위논문(석사) -- 부산대학교 대학원 , 정보융합공학과-컴퓨터공학전공 , 2024. 2
2024
영어
부산
32 ; 26 cm
지도교수: 권혁철
I804:21016-000000163268
0
상세조회0
다운로드기계번역은 사람의 관여 없이 자동으로 한 언어에서 다른 언어로 텍스트를 번역하기 위해서 계산을 사용하는 과정이다. 신경망을 사용하는 현재 기계번역 시스템의 용이성과 발전성 외에도...
기계번역은 사람의 관여 없이 자동으로 한 언어에서 다른 언어로 텍스트를 번역하기 위해서 계산을 사용하는 과정이다. 신경망을 사용하는 현재 기계번역 시스템의 용이성과 발전성 외에도 신경망 기계번역 시스템은 여전히 언어별 문제로 어려움을 겪고 있다. 언어별 문제의 예 중 하나는 문법적 형태의 오류와 관련된 존댓말의 번역이다. 본 연구에서는 한국어-인도네시아어 병렬 말뭉치로 파인튜닝을 하기 위해서 NLLB-200이라는 다국어 번역 모델을 사용한다. 기준 모델의 결과에서 번역에서 존댓말과 관련된 문제가 일부 있음을 보여주었다.
그 다음에 결과를 개선하기 위해서 빔서치(beam search) 디코딩 알고리즘이 사용된다. 결과는 빔 크기의 변화에 따라서 0.74~0.98점 범위의 SacreBLEU 점수는 향상되었음이 나타났다. 그 후에는 데이터 증강 기법을 적용하여 사용 가능한 단일 언어 말뭉치를 활용한다. 여기서 back-translation 기법은 모델 학습을 위한 소스(합성) - 타겟 병렬 데이터를 생성하는 데 사용된다. Back-translation은 인도네시아어 - 한국어 모델에서 24.37점이며 한국어 - 인도네시아어 모델에서 25.86점의 SacreBLEU 점수 결과를 제공한다. 결론적으로, 본 연구에서는 빔서치 디코딩 알고리즘과 결합된 back-translation이 존댓말 번역의 대부분 오류를 교정할 수 있다.
다국어 초록 (Multilingual Abstract)
Machine translation is the process of using a computation to automatically translate text from one language to another without human involvement. Apart from the ease and progress of the current translation system which uses a neural network, the Neura...
Machine translation is the process of using a computation to automatically translate text from one language to another without human involvement. Apart from the ease and progress of the current translation system which uses a neural network, the Neural Machine Translation system still suffers from language-specific problems. One of the examples for language-specific problems is the translation of honorifics, which relates to the error in grammatical form. This study uses a multilingual model named NLLB-200 to be fine-tuned with the Korean-Indonesian parallel corpus. In the result of baseline model, it was shown that there are some problems related to honorifics in translation.
To better the results, a beam search decoding algorithm is used. The result indicates that there is an improvement in SacreBLEU scores ranging from 0.74 to 0.98 points, depending on the change in beam size. Then, data augmentation method is applied to utilize available monolingual corpus. Here, back-translation technique is used for creating source (synthetic) - target parallel data for training. The back-translation gives a SacreBLEU score result of 24.37 points in Indonesian - Korean model and 25.86 points in Korean - Indonesian model. In this study, back-translation combined with beam search decoding algorithm is able to correct some errors in honorifics translation.
목차 (Table of Contents)