RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    위키데이터 지식 그래프를 활용한 대규모 언어 모델의 다국어 일관성 측정 벤치마크 생성 = Benchmark Generation for Measuring Cross-Lingual Consistency of Large Language Models with Wikidata Knowledge Graph

    한글로보기

    https://www.riss.kr/link?id=T17143281

    • 저자
    • 발행사항

      서울 : 숭실대학교 대학원, 2024

    • 학위논문사항

      학위논문(석사) -- 숭실대학교 대학원 , 지능형반도체학과(일원) , 2025. 2

    • 발행연도

      2024

    • 작성언어

      한국어

    • 발행국(도시)

      서울

    • 형태사항

      51 ; 26 cm

    • 일반주기명

      지도교수: 신동화

    • UCI식별코드

      I804:11044-200000847820

    • 소장기관
      • 숭실대학교 도서관 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델은 높은 수준의 언어 이해 능력과 생성 능력을 발휘 하는 것으로 평가된다. 다양한 언어에서의 학습 데이터 증가는 모델의 다국어 능력을 고도화하는 데 핵심적인 기여를 하였다. 그러나, 기존의 연구 및 성능 평가 방식은 제한적 언어 지원, 평가 데이터의 낮은 확장성, 그리고 여러 언어로 구성된 동일 문제에 대한 응답의 일관성을 검증 하기 어렵다는 한계를 지닌다. 본 연구에서는 대규모 언어 모델의 다국 어 답변 일관성 측면을 평가할 수 있는 새로운 벤치마크를 제안한다. 특히, 언어 모델과 지식 그래프 구조의 위키데이터(Wikidata)를 결합하여 자동화된 대규모 벤치마크를 설계 및 평가하고, 이를 통해 다국어 환경에서 동일한 사실관계에 대한 일관적 응답을 체계적으로 분석할 수 있는 새로운 방법론을 제시한다. 따라서 본 연구는 기존의 언어 모델 평가 방식의 한계점을 보완하고, 언어별 답변의 일관성을 정량적으로 평가할 수 있는 새로운 기준을 제시하여 향후 대규모 언어 모델 평가에 관한 연구 의 새로운 연구 방향성을 제시한다.
    번역하기

    대규모 언어 모델은 높은 수준의 언어 이해 능력과 생성 능력을 발휘 하는 것으로 평가된다. 다양한 언어에서의 학습 데이터 증가는 모델의 다국어 능력을 고도화하는 데 핵심적인 기여를 ...

    대규모 언어 모델은 높은 수준의 언어 이해 능력과 생성 능력을 발휘 하는 것으로 평가된다. 다양한 언어에서의 학습 데이터 증가는 모델의 다국어 능력을 고도화하는 데 핵심적인 기여를 하였다. 그러나, 기존의 연구 및 성능 평가 방식은 제한적 언어 지원, 평가 데이터의 낮은 확장성, 그리고 여러 언어로 구성된 동일 문제에 대한 응답의 일관성을 검증 하기 어렵다는 한계를 지닌다. 본 연구에서는 대규모 언어 모델의 다국 어 답변 일관성 측면을 평가할 수 있는 새로운 벤치마크를 제안한다. 특히, 언어 모델과 지식 그래프 구조의 위키데이터(Wikidata)를 결합하여 자동화된 대규모 벤치마크를 설계 및 평가하고, 이를 통해 다국어 환경에서 동일한 사실관계에 대한 일관적 응답을 체계적으로 분석할 수 있는 새로운 방법론을 제시한다. 따라서 본 연구는 기존의 언어 모델 평가 방식의 한계점을 보완하고, 언어별 답변의 일관성을 정량적으로 평가할 수 있는 새로운 기준을 제시하여 향후 대규모 언어 모델 평가에 관한 연구 의 새로운 연구 방향성을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large-scale language models are evaluated to exhibit a high level of linguistic comprehension and generative capabilities. The augmentation of training data across various languages has substantially contributed to advancing the model’s multilingual proficiency. However, existing research and performance evaluation methods have limitations such as restricted language coverage, low scalability of evaluation datasets, and challenges in verifying the consistency of responses to the same factual-relationship-based tasks across multiple languages. In this work, we propose a novel benchmark to evaluate multilingual consistency of language model’s response. Specifically, we construct and evaluate large-scale automated benchmarks by combining language models and Wikidata knowledge graph structure, thereby presenting a new methodology to systematically analyze consistent responses to the same factual-relationship based tasks in multilingual environments. Consequently, this study addresses the limitations of existing language model evaluation approaches and establishes a new criterion for quantitatively assessing the consistency of multilingual responses, providing a novel research direction for the evaluation of large-scale language models.
    번역하기

    Large-scale language models are evaluated to exhibit a high level of linguistic comprehension and generative capabilities. The augmentation of training data across various languages has substantially contributed to advancing the model’s multilingual...

    Large-scale language models are evaluated to exhibit a high level of linguistic comprehension and generative capabilities. The augmentation of training data across various languages has substantially contributed to advancing the model’s multilingual proficiency. However, existing research and performance evaluation methods have limitations such as restricted language coverage, low scalability of evaluation datasets, and challenges in verifying the consistency of responses to the same factual-relationship-based tasks across multiple languages. In this work, we propose a novel benchmark to evaluate multilingual consistency of language model’s response. Specifically, we construct and evaluate large-scale automated benchmarks by combining language models and Wikidata knowledge graph structure, thereby presenting a new methodology to systematically analyze consistent responses to the same factual-relationship based tasks in multilingual environments. Consequently, this study addresses the limitations of existing language model evaluation approaches and establishes a new criterion for quantitatively assessing the consistency of multilingual responses, providing a novel research direction for the evaluation of large-scale language models.

    더보기

    목차 (Table of Contents)

    • 제1 장 서론 1
    • 1.1 연구 배경 및 연구 목적 1
    • 1.1.1 연구 배경 1
    • 1.1.2 연구 목적 2
    • 제2 장 이론적 배경 3
    • 제1 장 서론 1
    • 1.1 연구 배경 및 연구 목적 1
    • 1.1.1 연구 배경 1
    • 1.1.2 연구 목적 2
    • 제2 장 이론적 배경 3
    • 2.1 대규모 언어 모델 개요 3
    • 2.1.1 트랜스포머 구조와 핵심 메커니즘 3
    • 2.1.2 언어 모델의 한계와 도전 과제 4
    • 2.1.3 최신 언어 모델의 동향 5
    • 2.2. 언어 모델의 성능 측정 5
    • 2.2.1 기존 벤치마크의 역사와 발전 5
    • 2.2.2 기존 벤치마크의 한계점 6
    • 제3 장 데이터 구성 및 수집 방법 8
    • 3.1 데이터 8
    • 3.1.1 데이터 수집 및 전처리 8
    • 3.1.1.1 위키 데이터 사실관계 추출 8
    • 3.1.1.2 자연어 표기 추출 및 수집 8
    • 3.1.1.3 데이터 템플릿 설계 9
    • 3.2 분석 환경 및 분석 방법 10
    • 3.2.1 시스템 구성 10
    • 3.2.2 분석 방법 10
    • 제4 장 벤치마크 설계 및 생성 11
    • 4.1 벤치마크 설계 11
    • 4.2 벤치마크 생성 11
    • 4.2.1 벤치마크 생성 모델 선정 11
    • 4.2.2 언어 모델을 활용한 문장 생성 과정 12
    • 4.2.2.1 프롬프트 작성 12
    • 4.2.2.2 배치크기 최적화 13
    • 4.2.2.3 사실관계 문장 생성 과정 자동화 13
    • 4.2.2.4 거짓관계 문장 생성 과정 자동화 14
    • 4.2.3 벤치마크 평가 방식 선정 15
    • 4.2.3.1 다국어 일관성 측정 벤치마크의 평가 방식 15
    • 제5 장 벤치마크 평가 및 분석 16
    • 5.1 벤치마크 평가 수행 16
    • 5.2 벤치마크 결과 분석 16
    • 5.2.1 기존 벤치마크와의 정답률 비교 16
    • 5.2.1.1 사실관계 문장 정답률 비교 16
    • 5.2.1.2 전체 정답률 비교 17
    • 5.2.2 언어 모델의 참/거짓 판별의 경향성 분석 17
    • 5.3 일관성 점수 산정 방식 24
    • 5.3.1 일관성 점수를 통한 평가 26
    • 5.3.2 언어별 일관성 점수 비교 28
    • 5.3.3 일관성 점수 분석의 의의 29
    • 제6 장 결론 31
    • 6.1 결론 및 시사점 31
    • 6.1.1 벤치마크 정확도 테스트 31
    • 6.1.2 언어 간 답변 일관성 테스트 (Coherent Score Test) 31
    • 6.2 향후 연구 발전 방향 33
    • 참고문헌 34
    • 부 록 37
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼