RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Multi-Image Object Hallucination Benchmark = 다중 이미지 기반의 객체 환각현상 평가

    한글로보기

    https://www.riss.kr/link?id=T17450671

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts.
    However, current MLLMs remain fundamentally limited by object hallucination—generating plausible yet factually inconsistent descriptions about objects.
    Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose
    how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures(visual context scale,perceptual difficulty, contextual bias). Through evaluation of 30 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images.
    MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.
    번역하기

    Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts. However, current MLLMs remain fundamentally limited by object hallucination—generating plausible yet fact...

    Multimodal Large Language Models (MLLMs) are increasingly deployed in multi-image scenarios requiring complex reasoning across visual contexts.
    However, current MLLMs remain fundamentally limited by object hallucination—generating plausible yet factually inconsistent descriptions about objects.
    Existing benchmarks, designed primarily for single-image settings or providing only high-level multi-image assessments, cannot systematically diagnose
    how visual complexity and reasoning demands trigger hallucination. To address this gap, we introduce MIOH, a fine-grained multi-image object hallucination benchmark that systematically evaluates object hallucination across four foundational tasks (existence, counting, attribute, position) through three multi-image reasoning patterns (comprehensive, comparative, selective) under three controlled adversarial pressures(visual context scale,perceptual difficulty, contextual bias). Through evaluation of 30 models, we reveal that even state-of-the-art systems like GPT-5 and Gemini-2.5-Pro exhibit distinct failure patterns across different reasoning patterns and tasks. Our evaluation reveals that hallucination stems not merely from perceptual failures but from integration-stage limitations when maintaining object representations across multiple images.
    MIOH provides a controlled framework for analyzing multi-image object hallucination and serves as a critical evaluation tool for developing more reliable multimodal AI systems.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    멀티모달 거대언어 모델(Multimodal Large Language Models)은여러 이미지를 동시에 처리하며 복잡한 시각적추론을요구하는 시나리오에 점점 더 널리 활용되고있다. 그러나 현재의MLLMs는 여전히 객체 환각(object hallucination)이라는 근본적인한계를 지니고있으며, 이는 실제 이미지와 일치하지 않지만 그럴듯해 보이는 객체 정보를 생성하는 문제를 의미한다. 기존의 벤치마크들은 주로 단일 이미지 설정에 초점을 맞추거나, 다중 이미지 환경에서도 고수준의 성능만을 평가하는 데 그쳐, 시각적 복잡성과 추론 요구가 어떻게 객체 환각을 유발하는지를 체계적으로 진단하는 데 한계가 존재한다. 이러한 공백을해소하기 위해, 본연구에서는 다중 이미지 객체 환각 벤치마크 (Miulti Image Object Hallucination Benchmark)를 제안한다. 제안하는 벤치마크는 여러 장의 이미지가 주어지는 환경에서의 객체 환각을 보다 세밀하게 분석할 수 있는 벤치마크로, 객체 존재, 개수, 속성, 위치관계의 네가지 핵심질문 유형을 대상으로 한다. 또한 여러 장의 이미지를 처리하기 위해 필요한 종합, 비교, 선택의 세 가지 추론 능력과, 환각을 유발할 수 있는 이미지의 개수 증가, 시각적 어려움, 맥락적인 편향 세 가지 요소들을 결합하여 다양한 관점에서 객체 환각을 체계적으로 평가한다. 30개의 최신 모델을 대상으로 한 광범위한 실험을통해, GPT-5와 Gemini-2.5-Pro와 같은 최신 모델들조차도 과제 유형과 추론 패턴에 따라 뚜렷하게 구분되는 실패 양상을 보임을 확인하였다. 실험 결과는 객체 환각이 단순한 시각적 인식의실패에 국한된 문제가 아니라, 다수의 이미지 전반에 걸쳐 객체 표현을 유지하고 통합하는 과정에서 발생하는 정보 융합 과정의 구조적 한계에서 기인함을 보여준다. 이 논문에서 제안하는 다중 이미지 객체 환각 벤치마크는, 다중 이미지 환경에서의 객체 환각을 정밀하게분석 할 수있는 통제된 실험 환경을 제공하며, 보다 신뢰할 수있는 멀티모달 인공지능 시스템 개발을위한 핵심적인 평가 도구로 활용될 수. 있을 것이다.
    번역하기

    멀티모달 거대언어 모델(Multimodal Large Language Models)은여러 이미지를 동시에 처리하며 복잡한 시각적추론을요구하는 시나리오에 점점 더 널리 활용되고있다. 그러나 현재의MLLMs는 여전히 객체...

    멀티모달 거대언어 모델(Multimodal Large Language Models)은여러 이미지를 동시에 처리하며 복잡한 시각적추론을요구하는 시나리오에 점점 더 널리 활용되고있다. 그러나 현재의MLLMs는 여전히 객체 환각(object hallucination)이라는 근본적인한계를 지니고있으며, 이는 실제 이미지와 일치하지 않지만 그럴듯해 보이는 객체 정보를 생성하는 문제를 의미한다. 기존의 벤치마크들은 주로 단일 이미지 설정에 초점을 맞추거나, 다중 이미지 환경에서도 고수준의 성능만을 평가하는 데 그쳐, 시각적 복잡성과 추론 요구가 어떻게 객체 환각을 유발하는지를 체계적으로 진단하는 데 한계가 존재한다. 이러한 공백을해소하기 위해, 본연구에서는 다중 이미지 객체 환각 벤치마크 (Miulti Image Object Hallucination Benchmark)를 제안한다. 제안하는 벤치마크는 여러 장의 이미지가 주어지는 환경에서의 객체 환각을 보다 세밀하게 분석할 수 있는 벤치마크로, 객체 존재, 개수, 속성, 위치관계의 네가지 핵심질문 유형을 대상으로 한다. 또한 여러 장의 이미지를 처리하기 위해 필요한 종합, 비교, 선택의 세 가지 추론 능력과, 환각을 유발할 수 있는 이미지의 개수 증가, 시각적 어려움, 맥락적인 편향 세 가지 요소들을 결합하여 다양한 관점에서 객체 환각을 체계적으로 평가한다. 30개의 최신 모델을 대상으로 한 광범위한 실험을통해, GPT-5와 Gemini-2.5-Pro와 같은 최신 모델들조차도 과제 유형과 추론 패턴에 따라 뚜렷하게 구분되는 실패 양상을 보임을 확인하였다. 실험 결과는 객체 환각이 단순한 시각적 인식의실패에 국한된 문제가 아니라, 다수의 이미지 전반에 걸쳐 객체 표현을 유지하고 통합하는 과정에서 발생하는 정보 융합 과정의 구조적 한계에서 기인함을 보여준다. 이 논문에서 제안하는 다중 이미지 객체 환각 벤치마크는, 다중 이미지 환경에서의 객체 환각을 정밀하게분석 할 수있는 통제된 실험 환경을 제공하며, 보다 신뢰할 수있는 멀티모달 인공지능 시스템 개발을위한 핵심적인 평가 도구로 활용될 수. 있을 것이다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • 1. Introduction 1
    • 2. Related Work 4
    • 3. Method 7
    • Abstract i
    • Contents ii
    • 1. Introduction 1
    • 2. Related Work 4
    • 3. Method 7
    • 4. Experiment 14
    • 5. Conclusion 22
    • Abstract (In Korean) 33
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼