RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    멀티모달 LLM의 환각 응답 분석과 입력 방식별 오류 특성 비교에 관한 연구 = A Study on Hallucination Response Analysis of Multimodal LLM and Error Characteristics Across Input Modalities

    한글로보기

    https://www.riss.kr/link?id=T17405604

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advances in large language models (LLMs) have led to the emergence of multimodal LLMs capable of jointly interpreting and generating information from both visual and linguistic inputs. Representative models such as GPT-4o, Gemini, Claude, and LLaVA demonstrate high performance in complex visual question answering, explanation generation, and creative reasoning tasks, thereby expanding their applicability across domains such as education, healthcare, law, and design. Despite these advancements, multimodal LLMs frequently generate responses that appear plausible yet are factually incorrect, a phenomenon commonly referred to as hallucination. This issue continues to pose a fundamental challenge to the reliability of AI systems.
    In multimodal settings, hallucination errors cannot be sufficiently captured through simple correctness judgments. Errors often involve complex interactions among factors such as text–image inconsistency, overgeneralization of visual attributes, misinterpretation of user intent, and selective information bias. Prior studies have focused predominantly on hallucinations in text-only LLMs and have lacked systematic analyses of how hallucinations emerge and differ under multimodal conditions. In addition, the absence of a structured error taxonomy and the limited integration of quantitative and qualitative evaluation frameworks have constrained progress in building reliable multimodal AI systems.
    To address these gaps, this study conducts a structured analysis of hallucination errors in state-of-the-art multimodal LLMs under three input conditions: text-only, image-only, and text+image. The proposed framework combines automated evaluation metrics (e.g., GPTScore, CLIPScore) with expert-based qualitative assessment. Furthermore, this study introduces a multimodal hallucination error taxonomy consisting of four core error types—factual error, interpretive error, exaggeration/distortion, and logical inconsistency—supported by a codebook and evaluation protocol. Representative response analyses are used to interpret underlying error mechanisms, including processing pathways, fusion failures, and interaction constraints between input modalities and model structure.
    The results demonstrate distinct error patterns depending on input modality: text-only conditions predominantly yielded reasoning and interpretive errors, image-only conditions revealed failures in fine-grained attribute perception and contextual overgeneralization, while text+image conditions exhibited pronounced cross-modal alignment failures and information integration omissions. Notably, hallucination severity varied even under identical prompts, depending on input configuration. Model-wise comparison indicated that GPT-4o and Gemini produced more stable responses overall, whereas models such as LLaVA displayed tendencies to over-rely on visual cues or misinterpret query focus.
    This study contributes in four significant ways. First, it reframes hallucination in multimodal LLMs not as a surface-level correctness issue, but as a structural phenomenon rooted in information processing and fusion mechanisms. Second, it establishes a reusable error classification and evaluation framework for future multimodal AI research. Third, it proposes targeted improvement directions for multimodal model design, including enhanced cross-modal integration, grounding stabilization, and attribute reasoning refinement. Fourth, it demonstrates the value of a combined quantitative–qualitative analysis pipeline, enabling both large-scale evaluation and in-depth interpretive insight.
    Overall, the findings of this study provide foundational guidance for improving the reliability, transparency, and safety of multimodal LLM systems. As AI becomes increasingly embedded in everyday decision-making and collaborative contexts, understanding the origins and mechanisms of hallucination will be essential to developing models that are not only capable but trustworthy.
    번역하기

    Recent advances in large language models (LLMs) have led to the emergence of multimodal LLMs capable of jointly interpreting and generating information from both visual and linguistic inputs. Representative models such as GPT-4o, Gemini, Claude, and L...

    Recent advances in large language models (LLMs) have led to the emergence of multimodal LLMs capable of jointly interpreting and generating information from both visual and linguistic inputs. Representative models such as GPT-4o, Gemini, Claude, and LLaVA demonstrate high performance in complex visual question answering, explanation generation, and creative reasoning tasks, thereby expanding their applicability across domains such as education, healthcare, law, and design. Despite these advancements, multimodal LLMs frequently generate responses that appear plausible yet are factually incorrect, a phenomenon commonly referred to as hallucination. This issue continues to pose a fundamental challenge to the reliability of AI systems.
    In multimodal settings, hallucination errors cannot be sufficiently captured through simple correctness judgments. Errors often involve complex interactions among factors such as text–image inconsistency, overgeneralization of visual attributes, misinterpretation of user intent, and selective information bias. Prior studies have focused predominantly on hallucinations in text-only LLMs and have lacked systematic analyses of how hallucinations emerge and differ under multimodal conditions. In addition, the absence of a structured error taxonomy and the limited integration of quantitative and qualitative evaluation frameworks have constrained progress in building reliable multimodal AI systems.
    To address these gaps, this study conducts a structured analysis of hallucination errors in state-of-the-art multimodal LLMs under three input conditions: text-only, image-only, and text+image. The proposed framework combines automated evaluation metrics (e.g., GPTScore, CLIPScore) with expert-based qualitative assessment. Furthermore, this study introduces a multimodal hallucination error taxonomy consisting of four core error types—factual error, interpretive error, exaggeration/distortion, and logical inconsistency—supported by a codebook and evaluation protocol. Representative response analyses are used to interpret underlying error mechanisms, including processing pathways, fusion failures, and interaction constraints between input modalities and model structure.
    The results demonstrate distinct error patterns depending on input modality: text-only conditions predominantly yielded reasoning and interpretive errors, image-only conditions revealed failures in fine-grained attribute perception and contextual overgeneralization, while text+image conditions exhibited pronounced cross-modal alignment failures and information integration omissions. Notably, hallucination severity varied even under identical prompts, depending on input configuration. Model-wise comparison indicated that GPT-4o and Gemini produced more stable responses overall, whereas models such as LLaVA displayed tendencies to over-rely on visual cues or misinterpret query focus.
    This study contributes in four significant ways. First, it reframes hallucination in multimodal LLMs not as a surface-level correctness issue, but as a structural phenomenon rooted in information processing and fusion mechanisms. Second, it establishes a reusable error classification and evaluation framework for future multimodal AI research. Third, it proposes targeted improvement directions for multimodal model design, including enhanced cross-modal integration, grounding stabilization, and attribute reasoning refinement. Fourth, it demonstrates the value of a combined quantitative–qualitative analysis pipeline, enabling both large-scale evaluation and in-depth interpretive insight.
    Overall, the findings of this study provide foundational guidance for improving the reliability, transparency, and safety of multimodal LLM systems. As AI becomes increasingly embedded in everyday decision-making and collaborative contexts, understanding the origins and mechanisms of hallucination will be essential to developing models that are not only capable but trustworthy.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 대형언어모델(Large Language Models, 이하 LLM)의 진보는 단순한 자연어 처리 능력을 넘어 시각 정보와 언어 정보를 동시에 이해하고 생성하는 멀티모달 LLM(multimodal LLM)의 등장을 견인하고 있다[1]. GPT-4o, Gemini, Claude, LLaVA 등 대표적 멀티모달 모델들은 이미지와 텍스트를 결합한 복합 질의응답, 설명 생성, 창작 등 다양한 응용 영역에서 놀라운 성능을 보여주며 실용적 가능성을 빠르게 확장해가고 있다[2]. 이러한 기술 발전은 교육, 의료, 법률, 디자인 등 고차원의 지식 작업에도 인공지능의 활용을 가능하게 한다는 점에서 사회적 기대가 크다. 그러나 이와 동시에 멀티모달 LLM의 응답이 그럴듯하지만 사실과 어긋난 오류, 즉 환각(hallucination) 문제를 자주 발생시킨다는 점에서 신뢰성 문제는 여전히 본질적 한계로 지적되고 있다[3,4].
    특히, 멀티모달 환경에서 발생하는 환각 오류는 단순히 사실과 다르다는 판단만으로는 포착되기 어렵다. 텍스트와 이미지 간 정보 불일치, 특정 속성의 과잉 일반화, 질문 의도 오해, 정보 선택 편향 등 복합적 요인이 얽혀 있어 오류의 발생 양상과 원인을 구조적으로 분석하기 어렵기 때문이다. 기존 연구들은 대부분 텍스트 기반 LLM의 환각을 정답/오답 방식으로 평가하는 데 그쳤으며 멀티모달 문맥에서 환각이 어떻게 나타나고 어떤 메커니즘에 의해 유발되는지를 체계적으로 분석한 연구는 매우 부족한 실정이다[5]. 또한, 오류 유형 분류체계의 부재, 정량·정성 평가의 미비, 입력 방식별 특성 분석의 한계로 인해 멀티모달 LLM의 신뢰성을 평가하고 개선하기 위한 기반 연구가 절실한 상황이다[6].
    본 논문은 이러한 문제의식을 바탕으로 최신 멀티모달 LLM을 대상으로 한 환각 오류의 체계적 분석 및 유형화, 그리고 오류 발생 메커니즘의 실증적 규명을 목표로 한다. 구체적으로, 동일한 질문을 text-only, image-only, text+image 세 가지 입력 조건으로 나누어 실험을 설계하고 각 조건에서 생성된 응답을 자동화 지표(GPTScore, CLIPScore 등)와 전문가 정성 평가를 병행하여 다층적으로 분석하였다. 이를 위해 본 논문에서는 사실 오류, 해석 오류, 과장/왜곡, 논리 오류 등 네 가지 핵심 오류 유형을 중심으로 멀티모달 환각 오류 분류체계를 새롭게 제안하고 코드북 기반의 평가 프로토콜을 수립하였다. 또한, 오류 발생 빈도와 유형 분포를 계량화할 수 있는 평가 지표를 설계하고 representative 사례 분석을 통해 오류의 정보처리적 경로, 융합 실패 유형, 입력-모델 간 상호작용 문제 등을 해석적으로 도출하였다.
    연구 결과는 텍스트만 주어진 조건에서는 언어적 추론 오류와 해석 오류가, 이미지만 주어진 조건에서는 세부 속성 해석 실패와 맥락 과잉일반화 오류가, 텍스트와 이미지가 결합된 조건에서는 cross-modal alignment 실패, 정보 결합 누락, 선택 편향 등이 주요 오류로 나타났다. 특히 동일 질문임에도 입력 조건에 따라 오류 유형과 심각도가 현저히 달라지는 패턴이 다수 관찰되었으며 이는 멀티모달 LLM의 환각 오류가 단순한 응답 실패를 넘어 구조적 편향과 설계상의 제약에서 비롯됨을 시사한다. 모델 간 비교에서는 GPT-4o와 Gemini가 전반적으로 안정적인 응답을 보였으나 LLaVA 등 일부 오픈소스 모델은 이미지 정보에 과도하게 의존하거나 질문 초점을 해석하는 데 실패하는 경향이 뚜렷하게 나타났다.
    본 논문의 의의는 네 가지로 정리된다. 첫째, 멀티모달 LLM의 환각 오류를 단순한 정확성 문제가 아니라 정보처리 구조, 입력 해석 메커니즘, 융합 전략의 결과로 해석함으로써 오류 분석의 해상도를 높였다. 둘째, 오류 유형 분류체계 및 코드북 기반의 정성 평가 체계를 제안하여 향후 멀티모달 AI 평가에 지속적으로 활용 가능한 이론적·실험적 기반을 마련하였다. 셋째, 모델별·입력별 오류 양상을 비교·분석함으로써 향후 cross-modal integration 강화, grounding 안정화, 세부 속성 처리 개선 등 기술적 개선 방향을 제시하였다. 넷째, 정량·정성 분석을 통합한 평가 프레임워크를 구현함으로써 대규모 실험과 사례 중심 해석을 동시에 가능하게 하였다.
    향후 본 논문은 멀티모달 LLM의 사회적 신뢰성 제고, 설명 가능성 확보, 사용자 중심의 피드백 설계 등 다양한 후속 연구의 출발점이 될 수 있을 것이며 생성형 AI의 안정적 활용을 위한 핵심 기반 기술로 확장될 수 있을 것으로 사료된다.더불어, 인간과 AI 간 협업이 일상화되는 시대에 있어 오류의 본질을 이해하고 그 메커니즘을 해석할 수 있는 기술적·인문적 통합 접근의 중요성을 환기하는 계기가 되기를 기대한다.
    번역하기

    최근 대형언어모델(Large Language Models, 이하 LLM)의 진보는 단순한 자연어 처리 능력을 넘어 시각 정보와 언어 정보를 동시에 이해하고 생성하는 멀티모달 LLM(multimodal LLM)의 등장을 견인하고 있...

    최근 대형언어모델(Large Language Models, 이하 LLM)의 진보는 단순한 자연어 처리 능력을 넘어 시각 정보와 언어 정보를 동시에 이해하고 생성하는 멀티모달 LLM(multimodal LLM)의 등장을 견인하고 있다[1]. GPT-4o, Gemini, Claude, LLaVA 등 대표적 멀티모달 모델들은 이미지와 텍스트를 결합한 복합 질의응답, 설명 생성, 창작 등 다양한 응용 영역에서 놀라운 성능을 보여주며 실용적 가능성을 빠르게 확장해가고 있다[2]. 이러한 기술 발전은 교육, 의료, 법률, 디자인 등 고차원의 지식 작업에도 인공지능의 활용을 가능하게 한다는 점에서 사회적 기대가 크다. 그러나 이와 동시에 멀티모달 LLM의 응답이 그럴듯하지만 사실과 어긋난 오류, 즉 환각(hallucination) 문제를 자주 발생시킨다는 점에서 신뢰성 문제는 여전히 본질적 한계로 지적되고 있다[3,4].
    특히, 멀티모달 환경에서 발생하는 환각 오류는 단순히 사실과 다르다는 판단만으로는 포착되기 어렵다. 텍스트와 이미지 간 정보 불일치, 특정 속성의 과잉 일반화, 질문 의도 오해, 정보 선택 편향 등 복합적 요인이 얽혀 있어 오류의 발생 양상과 원인을 구조적으로 분석하기 어렵기 때문이다. 기존 연구들은 대부분 텍스트 기반 LLM의 환각을 정답/오답 방식으로 평가하는 데 그쳤으며 멀티모달 문맥에서 환각이 어떻게 나타나고 어떤 메커니즘에 의해 유발되는지를 체계적으로 분석한 연구는 매우 부족한 실정이다[5]. 또한, 오류 유형 분류체계의 부재, 정량·정성 평가의 미비, 입력 방식별 특성 분석의 한계로 인해 멀티모달 LLM의 신뢰성을 평가하고 개선하기 위한 기반 연구가 절실한 상황이다[6].
    본 논문은 이러한 문제의식을 바탕으로 최신 멀티모달 LLM을 대상으로 한 환각 오류의 체계적 분석 및 유형화, 그리고 오류 발생 메커니즘의 실증적 규명을 목표로 한다. 구체적으로, 동일한 질문을 text-only, image-only, text+image 세 가지 입력 조건으로 나누어 실험을 설계하고 각 조건에서 생성된 응답을 자동화 지표(GPTScore, CLIPScore 등)와 전문가 정성 평가를 병행하여 다층적으로 분석하였다. 이를 위해 본 논문에서는 사실 오류, 해석 오류, 과장/왜곡, 논리 오류 등 네 가지 핵심 오류 유형을 중심으로 멀티모달 환각 오류 분류체계를 새롭게 제안하고 코드북 기반의 평가 프로토콜을 수립하였다. 또한, 오류 발생 빈도와 유형 분포를 계량화할 수 있는 평가 지표를 설계하고 representative 사례 분석을 통해 오류의 정보처리적 경로, 융합 실패 유형, 입력-모델 간 상호작용 문제 등을 해석적으로 도출하였다.
    연구 결과는 텍스트만 주어진 조건에서는 언어적 추론 오류와 해석 오류가, 이미지만 주어진 조건에서는 세부 속성 해석 실패와 맥락 과잉일반화 오류가, 텍스트와 이미지가 결합된 조건에서는 cross-modal alignment 실패, 정보 결합 누락, 선택 편향 등이 주요 오류로 나타났다. 특히 동일 질문임에도 입력 조건에 따라 오류 유형과 심각도가 현저히 달라지는 패턴이 다수 관찰되었으며 이는 멀티모달 LLM의 환각 오류가 단순한 응답 실패를 넘어 구조적 편향과 설계상의 제약에서 비롯됨을 시사한다. 모델 간 비교에서는 GPT-4o와 Gemini가 전반적으로 안정적인 응답을 보였으나 LLaVA 등 일부 오픈소스 모델은 이미지 정보에 과도하게 의존하거나 질문 초점을 해석하는 데 실패하는 경향이 뚜렷하게 나타났다.
    본 논문의 의의는 네 가지로 정리된다. 첫째, 멀티모달 LLM의 환각 오류를 단순한 정확성 문제가 아니라 정보처리 구조, 입력 해석 메커니즘, 융합 전략의 결과로 해석함으로써 오류 분석의 해상도를 높였다. 둘째, 오류 유형 분류체계 및 코드북 기반의 정성 평가 체계를 제안하여 향후 멀티모달 AI 평가에 지속적으로 활용 가능한 이론적·실험적 기반을 마련하였다. 셋째, 모델별·입력별 오류 양상을 비교·분석함으로써 향후 cross-modal integration 강화, grounding 안정화, 세부 속성 처리 개선 등 기술적 개선 방향을 제시하였다. 넷째, 정량·정성 분석을 통합한 평가 프레임워크를 구현함으로써 대규모 실험과 사례 중심 해석을 동시에 가능하게 하였다.
    향후 본 논문은 멀티모달 LLM의 사회적 신뢰성 제고, 설명 가능성 확보, 사용자 중심의 피드백 설계 등 다양한 후속 연구의 출발점이 될 수 있을 것이며 생성형 AI의 안정적 활용을 위한 핵심 기반 기술로 확장될 수 있을 것으로 사료된다.더불어, 인간과 AI 간 협업이 일상화되는 시대에 있어 오류의 본질을 이해하고 그 메커니즘을 해석할 수 있는 기술적·인문적 통합 접근의 중요성을 환기하는 계기가 되기를 기대한다.

    더보기

    목차 (Table of Contents)

    • 국문초록 ⅰ
    • 목 차 ⅳ
    • 그림목차 ⅷ
    • 도표목차 ⅸ
    • 약 어 표 ⅹ
    • 국문초록 ⅰ
    • 목 차 ⅳ
    • 그림목차 ⅷ
    • 도표목차 ⅸ
    • 약 어 표 ⅹ
    • Ⅰ. 서 론 1
    • 1.1 연구배경 및 목적 1
    • 1.2 연구내용 및 범위 4
    • 1.3 논문의 구성 7
    • Ⅱ. 관련 연구 9
    • 2.1 환각 응답의 개념과 다층적 정의 9
    • 2.1.1 텍스트 기반 환각의 정의와 맥락 9
    • 2.1.2 멀티모달 환경에서의 환각 개념 확장 10
    • 2.1.3 다층적 환각 구조 11
    • 2.1.4 환각 개념에 대한 학계 내 다양한 입장 12
    • 2.2 환각 오류의 평가 방법과 한계 13
    • 2.2.1 정량적 평가 지표의 활용 13
    • 2.2.2 멀티모달 특화 평가 방식 14
    • 2.2.3 정성적 평가 16
    • 2.2.4 기존 평가 방식의 한계와 본 논문의 대응 전략 17
    • 2.3 멀티모달 LLM의 구조와 처리 방식 19
    • 2.3.1 멀티모달 LLM의 개념과 처리 방식 19
    • 2.3.2 입력 융합 방식 20
    • 2.3.3 대표 멀티모달 LLM 아키텍처 비교 21
    • 2.3.4 환각 발생과 모델 구조의 연관성 23
    • 2.4 오류 유형 분류에 대한 기존 논의 24
    • 2.4.1 기존 오류 유형 분류 체계의 개요 25
    • 2.4.2 멀티모달 오류에 적용 시 나타나는 문제점 25
    • 2.4.3 오류 분류 기준의 모호성과 평가 일관성 문제 27
    • 2.5 기존 연구의 한계와 본 논문의 기여 29
    • 2.5.1 기존 연구의 주요 한계 29
    • 2.5.2 본 논문의 학술적·기술적 기여 31
    • Ⅲ. 분석 프레임워크 및 실험 설계 33
    • 3.1 연구 대상 및 데이터셋 34
    • 3.1.1 분석 대상 모델 선정 이유 및 특성 34
    • 3.1.2 실험용 멀티모달 데이터셋 선정 및 구축 35
    • 3.1.3 데이터셋의 멀티모달 적합성 및 한계 37
    • 3.1.4 모델-데이터셋 매핑 및 실험 설계 시 고려 사항 37
    • 3.2 입력 유형 및 실험 조건 설계 38
    • 3.2.1 입력 유형 정의 38
    • 3.2.2 질문 유형 및 과제 설계 40
    • 3.2.3 환각 오류 유발 과제 설계 전략 41
    • 3.2.4 실험 환경 및 표준화 절차 43
    • 3.2.5 설계의 학술적·실험적 기여 43
    • 3.3 오류 유형 분류 체계 및 코드북 설계 44
    • 3.4 평가 지표 및 분석 방법 48
    • Ⅳ. 구현 및 실험 결과 분석 53
    • 4.1 실험 설계 요약 및 데이터셋 구축 53
    • 4.1.1 데이터셋 구성 및 문항 설계 54
    • 4.1.2 실문항 배치 및 난이도 분포 55
    • 4.1.3 개발 환경 및 프레임워크 56
    • 4.1.4 실험 재현성 확보 방안 및 학술적 의의 56
    • 4.2 모델별 실험 결과 분석 59
    • 4.2.1 모델별 환각 오류 발생률 59
    • 4.2.2 오류 유형별 분포 61
    • 4.2.3 입력 유형별 차이 62
    • 4.2.4 모델별 강점·약점 논의 65
    • 4.2.5 대표 사례 분석 66
    • 4.3 입력 유형별 오류 특성 분석 68
    • 4.3.1 분석 개요 68
    • 4.3.2 text-only 조건 68
    • 4.3.3 image-only 조건 70
    • 4.3.4 text+image 조건 71
    • 4.3.5 통계 검정 결과 73
    • 4.4 오류 발생 메커니즘 심층 분석 76
    • 4.4.1 분석 목적과 접근 방식 76
    • 4.4.2 Representative 사례 분석 76
    • 4.4.3 정보처리 및 융합 실패 관점에서의 해석 77
    • 4.4.4 모델별·입력별 차이의 구조적 원인 78
    • Ⅴ. 결 론 81
    • 참고문헌 84
    • 영문초록 90
    • 감사의 글(Acknowledgement) 93
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼