대규모 언어 모델은 높은 수준의 언어 이해 능력과 생성 능력을 발휘 하는 것으로 평가된다. 다양한 언어에서의 학습 데이터 증가는 모델의 다국어 능력을 고도화하는 데 핵심적인 기여를 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17143281
서울 : 숭실대학교 대학원, 2024
학위논문(석사) -- 숭실대학교 대학원 , 지능형반도체학과(일원) , 2025. 2
2024
한국어
서울
51 ; 26 cm
지도교수: 신동화
I804:11044-200000847820
0
상세조회0
다운로드대규모 언어 모델은 높은 수준의 언어 이해 능력과 생성 능력을 발휘 하는 것으로 평가된다. 다양한 언어에서의 학습 데이터 증가는 모델의 다국어 능력을 고도화하는 데 핵심적인 기여를 ...
대규모 언어 모델은 높은 수준의 언어 이해 능력과 생성 능력을 발휘 하는 것으로 평가된다. 다양한 언어에서의 학습 데이터 증가는 모델의 다국어 능력을 고도화하는 데 핵심적인 기여를 하였다. 그러나, 기존의 연구 및 성능 평가 방식은 제한적 언어 지원, 평가 데이터의 낮은 확장성, 그리고 여러 언어로 구성된 동일 문제에 대한 응답의 일관성을 검증 하기 어렵다는 한계를 지닌다. 본 연구에서는 대규모 언어 모델의 다국 어 답변 일관성 측면을 평가할 수 있는 새로운 벤치마크를 제안한다. 특히, 언어 모델과 지식 그래프 구조의 위키데이터(Wikidata)를 결합하여 자동화된 대규모 벤치마크를 설계 및 평가하고, 이를 통해 다국어 환경에서 동일한 사실관계에 대한 일관적 응답을 체계적으로 분석할 수 있는 새로운 방법론을 제시한다. 따라서 본 연구는 기존의 언어 모델 평가 방식의 한계점을 보완하고, 언어별 답변의 일관성을 정량적으로 평가할 수 있는 새로운 기준을 제시하여 향후 대규모 언어 모델 평가에 관한 연구 의 새로운 연구 방향성을 제시한다.
다국어 초록 (Multilingual Abstract)
Large-scale language models are evaluated to exhibit a high level of linguistic comprehension and generative capabilities. The augmentation of training data across various languages has substantially contributed to advancing the model’s multilingual...
Large-scale language models are evaluated to exhibit a high level of linguistic comprehension and generative capabilities. The augmentation of training data across various languages has substantially contributed to advancing the model’s multilingual proficiency. However, existing research and performance evaluation methods have limitations such as restricted language coverage, low scalability of evaluation datasets, and challenges in verifying the consistency of responses to the same factual-relationship-based tasks across multiple languages. In this work, we propose a novel benchmark to evaluate multilingual consistency of language model’s response. Specifically, we construct and evaluate large-scale automated benchmarks by combining language models and Wikidata knowledge graph structure, thereby presenting a new methodology to systematically analyze consistent responses to the same factual-relationship based tasks in multilingual environments. Consequently, this study addresses the limitations of existing language model evaluation approaches and establishes a new criterion for quantitatively assessing the consistency of multilingual responses, providing a novel research direction for the evaluation of large-scale language models.
목차 (Table of Contents)