RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    인공지능 다국어 번역을 위한 언어학적 전 처리와 공학적 처리 방법론 연구-역 방향 번역과 합성 코퍼스 생성을 중심으로 = A Study on Linguistic Preprocessing and Engineering Processing Methodology for Artificial Intelligence Multilingual Translation

    한글로보기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In this paper, we studied the basic humanities work that linguists can perform in the preprocessing process for natural language processing. Embedding an atypical and infinite human language into a structured and finite computer resource is a very important task that has not yet been solved, but must be solved in order to complete the natural language processing by artificial intelligence. In order to solve these tasks, humanities and linguistic knowledge must be mobilized, and engineering computational skills must be supported. In this paper, we introduced the underlying technology, focusing on the creation of the artificial intelligence synthesis corpus and the back translation, and introducing various issues for accurate meaning extraction in natural language. In the text, in particular, the problems of meaning accompaniment and categorization, morphological negation, direct visibility and implicitity, and existence were dealt with, and the introduction of Word2Vec, the concept of Subword, and the use of big data were suggested as engineering solutions. In connection with this, we proposed back translation and artificial intelligence synthesis corpus construction technologies, and explained that these technologies can play a particularly large role in natural language meaning extraction and multilingual translation. In fact, this research team conducted a Korean-English artificial intelligence translation test using the above technology, supported by NIPA high-performance computing resources. Based on a total of 4.6 million Korean-English parallel data, the data were trained by repeating 120 times for 60 days. It was found from the experimental results that the translation performance improved by about 5% when the back translation was repeated 40 times for 20 days (about 0.3 → 0.312).
    번역하기

    In this paper, we studied the basic humanities work that linguists can perform in the preprocessing process for natural language processing. Embedding an atypical and infinite human language into a structured and finite computer resource is a very imp...

    In this paper, we studied the basic humanities work that linguists can perform in the preprocessing process for natural language processing. Embedding an atypical and infinite human language into a structured and finite computer resource is a very important task that has not yet been solved, but must be solved in order to complete the natural language processing by artificial intelligence. In order to solve these tasks, humanities and linguistic knowledge must be mobilized, and engineering computational skills must be supported. In this paper, we introduced the underlying technology, focusing on the creation of the artificial intelligence synthesis corpus and the back translation, and introducing various issues for accurate meaning extraction in natural language. In the text, in particular, the problems of meaning accompaniment and categorization, morphological negation, direct visibility and implicitity, and existence were dealt with, and the introduction of Word2Vec, the concept of Subword, and the use of big data were suggested as engineering solutions. In connection with this, we proposed back translation and artificial intelligence synthesis corpus construction technologies, and explained that these technologies can play a particularly large role in natural language meaning extraction and multilingual translation. In fact, this research team conducted a Korean-English artificial intelligence translation test using the above technology, supported by NIPA high-performance computing resources. Based on a total of 4.6 million Korean-English parallel data, the data were trained by repeating 120 times for 60 days. It was found from the experimental results that the translation performance improved by about 5% when the back translation was repeated 40 times for 20 days (about 0.3 → 0.312).

    더보기

    참고문헌 (Reference)

    1 이기창, "한국어 임베딩" 에이콘 2019

    2 김재훈, "통합국어정보베이스를 위한 한국어 형태 통사 태그 설정" 한국과학기술원 1996

    3 Matteo Negri, "eSCAPE: a Large Synthetic Corpus for Automatic Post-Editing"

    4 Jun-Yan Zhu, "Unfaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks"

    5 Sergey Edunov, "Understanding Back-Translation at Scale"

    6 Sutskever, I., "Sequence to sequence learning with neural networks" 2014

    7 Richard Socher, "Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank"

    8 Marek Rei, "Jointly Learning to Label Sentences and Tokens"

    9 Rico Sennrich, "Improving neural machine translation models with monolingual data"

    10 R. Bowman, "GLUE: A multi-task benchmark and analysis platform for natural language understanding" 1-20, 2019

    1 이기창, "한국어 임베딩" 에이콘 2019

    2 김재훈, "통합국어정보베이스를 위한 한국어 형태 통사 태그 설정" 한국과학기술원 1996

    3 Matteo Negri, "eSCAPE: a Large Synthetic Corpus for Automatic Post-Editing"

    4 Jun-Yan Zhu, "Unfaired Image-to-Image Translation using Cycle-Consistent Adversarial Networks"

    5 Sergey Edunov, "Understanding Back-Translation at Scale"

    6 Sutskever, I., "Sequence to sequence learning with neural networks" 2014

    7 Richard Socher, "Recursive Deep Models for Semantic Compositionality over a Sentiment Treebank"

    8 Marek Rei, "Jointly Learning to Label Sentences and Tokens"

    9 Rico Sennrich, "Improving neural machine translation models with monolingual data"

    10 R. Bowman, "GLUE: A multi-task benchmark and analysis platform for natural language understanding" 1-20, 2019

    11 Tomas Mikolov, "Distributed Representations of Words and Phrases and their Compositionality" 3111-3119, 2013

    12 K. Papineni, "BLEU: a method for automatic evaluation of machine translation" 2002

    13 Marek Rei, "Auxiliary Objectives for Neural Error Detection Models"

    더보기

    동일학술지(권/호) 다른 논문

    동일학술지 더보기

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2022 평가 재인증평가 신청대상 (재인증)
    2019-01-01 등재 등재학술지 유지 (계속평가) KCI등재
    2016-01-01 등재 등재학술지 유지 (계속평가) KCI등재
    2013-04-22 학회명변경 영문명 : FOREIGN STUDIES CENTER -> FOREIGN STUDIES INSTITUTE KCI등재
    2012-01-01 등재 등재학술지 선정 (등재후보2차) KCI등재
    2011-01-01 등재 등재후보 1차 PASS (등재후보1차) KCI등재후보
    2009-01-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    2007-09-03 학회명변경 한글명 : 외국어문학연구소 -> 외국학연구소
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 0.28 0.28 0.25
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    0.22 0.2 0.437 0.12
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼