지시튜닝(Instruction-tuning)은 대형 언어 모델(LLM)이 사용자 지시를 보 다 정확하게 따를 수 있도록 하여, 사용성을 개선하고 해로운 출력을 감소 시키는 데 기여한다. 그러나 이러한 과정은 모...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
지시튜닝(Instruction-tuning)은 대형 언어 모델(LLM)이 사용자 지시를 보 다 정확하게 따를 수 있도록 하여, 사용성을 개선하고 해로운 출력을 감소 시키는 데 기여한다. 그러나 이러한 과정은 모...
지시튜닝(Instruction-tuning)은 대형 언어 모델(LLM)이 사용자 지시를 보 다 정확하게 따를 수 있도록 하여, 사용성을 개선하고 해로운 출력을 감소 시키는 데 기여한다. 그러나 이러한 과정은 모델의 사용자 입력에 대한 의 존도를 높여, 잘못된 정보를 여과 없이 수용하거나 환각(hallucination)을 생 성할 가능성을 증가시킬 수 있다. 기존 연구들은 주로 대형 언어 모델이 자 체 파라메트릭(parametric) 지식과 상충하는 외부 정보에 수용적인 경향이 있음을 보여주었으나, 지시튜닝이 이러한 현상에 미치는 직접적인 영향에 대한 연구는 부족하다. 본 연구에서는 지시튜닝이 대형 언어 모델의 허위 정보(misinformation) 취약성에 미치는 영향을 분석하였다. 그 결과, 지시튜닝을 거친 모델은 사용 자가 제시한 허위 정보를 수용할 가능성이 유의미하게 높았다. 베이스 모델 과의 비교를 통해, 지시튜닝이 모델의 사용자 제공 정보에 대한 의존성을 강화하여, 허위 정보 취약성이 사용자 역할에서 두드러지게 나타남을 확인 하였다. 또한 프롬프트 구조 내 사용자 역할, 허위 정보의 길이, 시스템 프롬프트의 경고 존재 여부 등 허위 정보 취약성에 영향을 미치는 추가 요인들을 탐구 하였다. 연구 결과는 지시튜닝의 의도치 않은 부작용을 완화하고, 실제 응용 환경에서 대형 언어모델의 신뢰성을 향상시키기 위한 체계적 접근의 필요성 을 시사한다.
다국어 초록 (Multilingual Abstract)
Instruction tuning enhances the usability of large language models (LLMs) by enabling them to follow user instructions more accurately and reducing harmful outputs. However, this process also increases the model’s reliance on user input, potentially...
Instruction tuning enhances the usability of large language models (LLMs) by enabling them to follow user instructions more accurately and reducing harmful outputs. However, this process also increases the model’s reliance on user input, potentially making it more likely to accept incorrect information without filtering it out or to generate hallucinations. While previous studies have largely shown that LLMs tend to be receptive to external information that conflicts with their own parametric knowledge, there has been limited research on the direct impact of instruction tuning
on this phenomenon.
In this study, we analyze how instruction tuning affects LLMs’ vulnerability to misinformation. Our findings show that instruction-tuned models are significantly more likely to accept false information provided by users. By comparing them with base models, we confirm that instruction tuning strengthens the model’s dependence on user-supplied information, making misinformation vulnerability particularly pronounced in the user role.
We also explore additional factors that influence susceptibility to misinformation, including the user role within the prompt structure, the length of the misinformation, and the presence of system-prompt warnings. The results suggest the need for systematic approaches to mitigate the unintended side effects of instruction tuning and to improve the reliability of LLMs in real-world applications.
목차 (Table of Contents)