RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    멀티모달 거대 언어 모델의 문맥 내 학습을 활용한 추론적 영상 분할 성능 향상 연구 = Research on Enhancing Reasoning Segmentation Performance via MLLM-based In-Context Learning

    한글로보기

    https://www.riss.kr/link?id=T17372332

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    지시 영상 분할(RIS) 기술은 명시적 지시문 처리에는 탁월하나, 맥락과 상식에 기반한 추론이 필요한 암시적 지시문을 처리하는 데에는 여전히 한계를 보인다. 기존 연구들은 이러한 '의도의 간극'을 줄이기 위해 고비용의 모델 미세 조정을 시도했으나, 데이터 구축과 학습 비용 측면에서 제약이 따랐다. 이에 본 연구에서는 분할 모델의 파라미터를 고정한 상태에서 멀티모달 거대 언어 모델(MLLM)의 문맥 내 학습(In-Context Learning) 능력을 활용한 새로운 프레임워크를 제안한다.

    제안하는 기법은 복잡한 자연어 지시문을 분할 모델이 즉각 이해할 수 있는 명시적 형태로 변환하는 것이 핵심이다. 이를 위해 CLIP 기반의 동적 예제 검색을 통해 유사한 맥락을 파악하고, MLLM을 통해 지시문을 재구성하며, 자가 교정 메커니즘을 통해 환각 현상을 방지한다. ReasonSeg 벤치마크 실험 결과, 본 프레임워크는 기존의 최신(SOTA) 모델인 LISA보다 향상된 분할 정확도를 달성하였다. 결론적으로 본 연구는 추가적인 데이터 학습 없이도 거대 언어 모델의 고도화된 추론 능력을 시각적 분할 태스크에 효과적으로 전이할 수 있음을 입증하였다.
    번역하기

    지시 영상 분할(RIS) 기술은 명시적 지시문 처리에는 탁월하나, 맥락과 상식에 기반한 추론이 필요한 암시적 지시문을 처리하는 데에는 여전히 한계를 보인다. 기존 연구들은 이러한 '의도의 ...

    지시 영상 분할(RIS) 기술은 명시적 지시문 처리에는 탁월하나, 맥락과 상식에 기반한 추론이 필요한 암시적 지시문을 처리하는 데에는 여전히 한계를 보인다. 기존 연구들은 이러한 '의도의 간극'을 줄이기 위해 고비용의 모델 미세 조정을 시도했으나, 데이터 구축과 학습 비용 측면에서 제약이 따랐다. 이에 본 연구에서는 분할 모델의 파라미터를 고정한 상태에서 멀티모달 거대 언어 모델(MLLM)의 문맥 내 학습(In-Context Learning) 능력을 활용한 새로운 프레임워크를 제안한다.

    제안하는 기법은 복잡한 자연어 지시문을 분할 모델이 즉각 이해할 수 있는 명시적 형태로 변환하는 것이 핵심이다. 이를 위해 CLIP 기반의 동적 예제 검색을 통해 유사한 맥락을 파악하고, MLLM을 통해 지시문을 재구성하며, 자가 교정 메커니즘을 통해 환각 현상을 방지한다. ReasonSeg 벤치마크 실험 결과, 본 프레임워크는 기존의 최신(SOTA) 모델인 LISA보다 향상된 분할 정확도를 달성하였다. 결론적으로 본 연구는 추가적인 데이터 학습 없이도 거대 언어 모델의 고도화된 추론 능력을 시각적 분할 태스크에 효과적으로 전이할 수 있음을 입증하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    While Referring Image Segmentation (RIS) excels at handling explicit instructions, it remains limited in processing implicit instructions that require reasoning based on context and common sense. Previous studies have attempted high-cost model fine-tuning to bridge this "intent gap," but they faced constraints regarding data construction and training costs. To address this, we propose a novel framework that leverages the In-Context Learning capabilities of Multimodal Large Language Models (MLLMs) while keeping the segmentation model's parameters frozen.

    The core of the proposed method lies in converting complex natural language instructions into explicit forms that the segmentation model can readily comprehend. To achieve this, we employ CLIP-based dynamic exemplar retrieval to capture similar contexts, reconstruct instructions using the MLLM, and mitigate hallucinations through a self-correction mechanism. Experimental results on the ReasonSeg benchmark demonstrate that our framework achieves superior segmentation accuracy compared to LISA, the current state-of-the-art (SOTA) model. In conclusion, this study demonstrates that the advanced reasoning capabilities of large language models can be effectively transferred to visual segmentation tasks without the need for additional data training.
    번역하기

    While Referring Image Segmentation (RIS) excels at handling explicit instructions, it remains limited in processing implicit instructions that require reasoning based on context and common sense. Previous studies have attempted high-cost model fine-tu...

    While Referring Image Segmentation (RIS) excels at handling explicit instructions, it remains limited in processing implicit instructions that require reasoning based on context and common sense. Previous studies have attempted high-cost model fine-tuning to bridge this "intent gap," but they faced constraints regarding data construction and training costs. To address this, we propose a novel framework that leverages the In-Context Learning capabilities of Multimodal Large Language Models (MLLMs) while keeping the segmentation model's parameters frozen.

    The core of the proposed method lies in converting complex natural language instructions into explicit forms that the segmentation model can readily comprehend. To achieve this, we employ CLIP-based dynamic exemplar retrieval to capture similar contexts, reconstruct instructions using the MLLM, and mitigate hallucinations through a self-correction mechanism. Experimental results on the ReasonSeg benchmark demonstrate that our framework achieves superior segmentation accuracy compared to LISA, the current state-of-the-art (SOTA) model. In conclusion, this study demonstrates that the advanced reasoning capabilities of large language models can be effectively transferred to visual segmentation tasks without the need for additional data training.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 1
    • 제2장 관련 연구 3
    • 2.1. 지시 영상 분할 (Referring Image Segmentation) 3
    • 2.1.1 초기 CNN-RNN 기반 접근법 3
    • 2.1.2 Transformer 및 Vision Transformer 기반 접근법 6
    • 제1장 서론 1
    • 제2장 관련 연구 3
    • 2.1. 지시 영상 분할 (Referring Image Segmentation) 3
    • 2.1.1 초기 CNN-RNN 기반 접근법 3
    • 2.1.2 Transformer 및 Vision Transformer 기반 접근법 6
    • 2.1.3 기존 모델의 명시적 지시문 의존성 한계 7
    • 2.2. 추론형 영상 분할 (Reasoning Segmentation) 8
    • 2.2.1 ReasonSeg 데이터셋 8
    • 2.2.2 최신 연구 동향 및 한계 9
    • 2.3. 멀티모달 대형 언어 모델 (MLLM) 11
    • 2.3.1. MLLM의 발전 11
    • 2.3.2. 문맥 내 학습 및 사고사슬 프롬프팅 11
    • 제3장 제안 기법 12
    • 3.1. 제안 기법 12
    • 3.2. CLIP 기반 동적 예제 검색 13
    • 3.3. 문맥 내 학습 기반 지시문 재구성 15
    • 3.4. 자가 교정 메커니즘 16
    • 3.5. 최종 영상 분할 및 적용 17
    • 제4장 실험 결과 18
    • 4.1. 실험 환경 18
    • 4.1.1. 활용 데이터 18
    • 4.1.2. 성능 평가 지표 19
    • 4.2. 실험 결과 및 분석 20
    • 4.2.1. 추론적 영상 분할 성능 비교 20
    • 4.2.2. 검색 전략 비교 및 예제 수 (N-shot)에 따른 영향 21
    • 4.2.3. 프롬프트 유형 비교 21
    • 4.2.4. 자기 교정 유무에 따른 효과 21
    • 4.3. 정성적 결과 분석 22
    • 4.4. 절제 연구 23
    • 제5장 결론 24
    • 참고문헌 25
    • 영문초록 26
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼