RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Enhancing Image-Text Alignment for Vision Tasks with Large Multi-Modal Models = 대규모 멀티모달 모델 기반 비전 작업들을 위한 이미지-텍스트 정렬 강화 연구

    한글로보기

    https://www.riss.kr/link?id=T17314695

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    인공지능의 급속한 발전으로 AI는 일상생활 속에 점점 더 깊이 통합되고 있습니다. 이러한 변화는 특정 분야에 국한되지 않고, 보다 다양한 복잡한 현실 세계의 시각 과제를 다룰 수 있는 시각 이해 및 분석 시스템에 대한 수요 증가로 이어지고 있습니다. 예를 들어, 자율주행에서의 동적 장면 해석, 일상 환경에서의 인간-객체 상호작용 분석, 기존의 과제 특화형 모델들이 어려움을 겪는 개방형(open-world) 객체 인식 과제 등이 이에 해당합니다.

    이러한 수요를 충족하기 위해, CLIP과 같은 비전-언어 모델 및 Stable Diffusion과 같은 텍스트-이미지 확산 모델 등 대규모 사전학습 멀티모달 모델이 주목받고 있습니다. 이들 모델은 언어적 지식과 시각적 지식을 통합함으로써 보다 직관적이고 맥락을 이해하며 인간 친화적인 방식으로 이미지를 해석할 수 있는 가능성을 제시합니다. 그러나 여전히 중요한 과제는, 특히 이미지-텍스트 정렬이 불완전한 경우, 이러한 모델들을 새로운 시각 작업에 대해 재학습 없이 그리고 과도한 연산 자원 소모 없이 어떻게 효과적으로 활용할 수 있는가입니다.

    본 학위논문은 이러한 한계를 극복하고자, 학습이 필요 없는(training-free) 제로샷(zero-shot) 기반 프레임워크를 통해 이미지-텍스트 정렬을 직접 강화하는 방식을 제안합니다. 이미지-텍스트 정렬은 시각 과제에서 인식, 공간적 정위(localization), 의미적 정합성(grounding)을 가능하게 하는 핵심 기반입니다. 하지만 대규모 멀티모달 모델은 높은 일반화 능력에도 불구하고, 정밀한 시각적 그라운딩, 문맥 이해, 세밀한 모달 간 정렬에 있어 구조적 한계를 보입니다. 이러한 문제는 특히 밀집된 장면, 모호한 시각 개념, 상세한 텍스트 설명이 요구되는 개방형 환경에서 더욱 뚜렷하게 나타나며, 이는 멀티모달 시스템의 신뢰성과 해석 가능성을 저해하는 요인으로 작용합니다.

    첫째, 제로샷 다중 레이블 인식(multi-label recognition)을 위해, CLIP 기반 이미지-텍스트 임베딩 정렬을 강화하는 학습 없는 접근법을 제안합니다. 본 방법은 클래스 개념 표현(class concept representation)과 클래스 유도 시각 표현(class-guided visual feature)을 활용하여, 이미지 특징을 텍스트 임베딩 공간으로 사영하는 주의(attention) 기반 정렬 메커니즘을 도입합니다. 이로써 대규모 텍스트 설명 내 동시 발생(co-occurrence) 및 의미적 패턴을 포착하여 수작업 프롬프트 없이도 더 정확하고 문맥에 부합하는 인식이 가능합니다.

    둘째, 개방형 어휘 의미 분할(open-vocabulary semantic segmentation)에서 CLIP의 위치 정렬 성능을 개선하기 위해 클래스 분포 기반 어텐션 맵(Class Distribution-induced Attention Map, CDAM)을 제안합니다. CDAM은 이미지 패치 간 클래스 분포를 바탕으로 의미적 유사도를 계산하고, 클래스 일관성을 가진 주의 정보를 공간적으로 전파함으로써 기존의 노이즈 많고 조밀하지 못한 어텐션 맵을 정제합니다. 또한 다중 스케일 이미지 패치 처리, 텍스트 프롬프트 확장, 엔트로피 기반 배경 분리 기법을 결합하여 학습 없이도 정확한 제로샷 의미 분할을 실현합니다.

    셋째, 인간 중심 캡션으로부터 고해상도 이미지를 생성하기 위해, 대형 언어 모델을 활용하여 텍스트로부터 키포인트-박스 레이아웃을 생성하고, 이를 공간적 사전 지식으로 활용하는 계층적 파이프라인을 제안합니다. 해당 레이아웃은 다단계 확산 생성 과정을 안내하며, 세밀한 시각적 정합성과 복잡한 장면 내 일관성을 확보합니다.

    이러한 연구 기여들은 학습 없이도 이미지-텍스트 정렬을 강화함으로써 대규모 멀티모달 모델의 성능, 적응력, 해석 가능성을 획기적으로 향상시킬 수 있음을 보여줍니다. 이는 해당 모델들이 현실 세계의 개방형 시각 과제를 위한 유연하고 강력한 시각 추론 시스템으로 발전할 수 있는 실용적 가능성을 제시합니다.
    번역하기

    인공지능의 급속한 발전으로 AI는 일상생활 속에 점점 더 깊이 통합되고 있습니다. 이러한 변화는 특정 분야에 국한되지 않고, 보다 다양한 복잡한 현실 세계의 시각 과제를 다룰 수 있는 시...

    인공지능의 급속한 발전으로 AI는 일상생활 속에 점점 더 깊이 통합되고 있습니다. 이러한 변화는 특정 분야에 국한되지 않고, 보다 다양한 복잡한 현실 세계의 시각 과제를 다룰 수 있는 시각 이해 및 분석 시스템에 대한 수요 증가로 이어지고 있습니다. 예를 들어, 자율주행에서의 동적 장면 해석, 일상 환경에서의 인간-객체 상호작용 분석, 기존의 과제 특화형 모델들이 어려움을 겪는 개방형(open-world) 객체 인식 과제 등이 이에 해당합니다.

    이러한 수요를 충족하기 위해, CLIP과 같은 비전-언어 모델 및 Stable Diffusion과 같은 텍스트-이미지 확산 모델 등 대규모 사전학습 멀티모달 모델이 주목받고 있습니다. 이들 모델은 언어적 지식과 시각적 지식을 통합함으로써 보다 직관적이고 맥락을 이해하며 인간 친화적인 방식으로 이미지를 해석할 수 있는 가능성을 제시합니다. 그러나 여전히 중요한 과제는, 특히 이미지-텍스트 정렬이 불완전한 경우, 이러한 모델들을 새로운 시각 작업에 대해 재학습 없이 그리고 과도한 연산 자원 소모 없이 어떻게 효과적으로 활용할 수 있는가입니다.

    본 학위논문은 이러한 한계를 극복하고자, 학습이 필요 없는(training-free) 제로샷(zero-shot) 기반 프레임워크를 통해 이미지-텍스트 정렬을 직접 강화하는 방식을 제안합니다. 이미지-텍스트 정렬은 시각 과제에서 인식, 공간적 정위(localization), 의미적 정합성(grounding)을 가능하게 하는 핵심 기반입니다. 하지만 대규모 멀티모달 모델은 높은 일반화 능력에도 불구하고, 정밀한 시각적 그라운딩, 문맥 이해, 세밀한 모달 간 정렬에 있어 구조적 한계를 보입니다. 이러한 문제는 특히 밀집된 장면, 모호한 시각 개념, 상세한 텍스트 설명이 요구되는 개방형 환경에서 더욱 뚜렷하게 나타나며, 이는 멀티모달 시스템의 신뢰성과 해석 가능성을 저해하는 요인으로 작용합니다.

    첫째, 제로샷 다중 레이블 인식(multi-label recognition)을 위해, CLIP 기반 이미지-텍스트 임베딩 정렬을 강화하는 학습 없는 접근법을 제안합니다. 본 방법은 클래스 개념 표현(class concept representation)과 클래스 유도 시각 표현(class-guided visual feature)을 활용하여, 이미지 특징을 텍스트 임베딩 공간으로 사영하는 주의(attention) 기반 정렬 메커니즘을 도입합니다. 이로써 대규모 텍스트 설명 내 동시 발생(co-occurrence) 및 의미적 패턴을 포착하여 수작업 프롬프트 없이도 더 정확하고 문맥에 부합하는 인식이 가능합니다.

    둘째, 개방형 어휘 의미 분할(open-vocabulary semantic segmentation)에서 CLIP의 위치 정렬 성능을 개선하기 위해 클래스 분포 기반 어텐션 맵(Class Distribution-induced Attention Map, CDAM)을 제안합니다. CDAM은 이미지 패치 간 클래스 분포를 바탕으로 의미적 유사도를 계산하고, 클래스 일관성을 가진 주의 정보를 공간적으로 전파함으로써 기존의 노이즈 많고 조밀하지 못한 어텐션 맵을 정제합니다. 또한 다중 스케일 이미지 패치 처리, 텍스트 프롬프트 확장, 엔트로피 기반 배경 분리 기법을 결합하여 학습 없이도 정확한 제로샷 의미 분할을 실현합니다.

    셋째, 인간 중심 캡션으로부터 고해상도 이미지를 생성하기 위해, 대형 언어 모델을 활용하여 텍스트로부터 키포인트-박스 레이아웃을 생성하고, 이를 공간적 사전 지식으로 활용하는 계층적 파이프라인을 제안합니다. 해당 레이아웃은 다단계 확산 생성 과정을 안내하며, 세밀한 시각적 정합성과 복잡한 장면 내 일관성을 확보합니다.

    이러한 연구 기여들은 학습 없이도 이미지-텍스트 정렬을 강화함으로써 대규모 멀티모달 모델의 성능, 적응력, 해석 가능성을 획기적으로 향상시킬 수 있음을 보여줍니다. 이는 해당 모델들이 현실 세계의 개방형 시각 과제를 위한 유연하고 강력한 시각 추론 시스템으로 발전할 수 있는 실용적 가능성을 제시합니다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the rapid advancement of artificial intelligence, AI are becoming increasingly embedded in everyday life. This shift has created a growing demand for vision understanding systems that extend beyond task-specific applications and can address diverse, complex real-world scenarios such as dynamic scene interpretation in autonomous driving, human-object interaction analysis in daily environments, and open-world object recognition that conventional models often struggle to address.

    To meet this demand, large multi-modal models, particularly vision-language models like CLIP and text-to-image diffusion models such as Stable Diffusion, have emerged as promising approaches. By integrating linguistic and visual knowledge, these models enable intuitive and context-aware visual understanding. However, a key challenge remains: effectively leveraging these models for novel and complex vision tasks without retraining or heavy computational cost, especially in the presence of imperfect image-text alignment.

    This dissertation aims to overcome the limitations of large multi-modal models by proposing training-free, zero-shot frameworks that directly enhance image-text alignment. This alignment serves as a critical foundation for enabling accurate recognition, spatial localization, and grounding in real-world vision applications. Although these models exhibit strong generalization across tasks, they often face fundamental limitations in precise visual grounding, contextual understanding, and fine-grained cross-modal alignment. These challenges become especially apparent in open-world settings involving dense scenes, ambiguous visual concepts, and detailed textual descriptions, which ultimately hinder the reliability and interpretability of multi-modal systems in human-centered environments.

    First, for zero-shot multi-label recognition, we propose a training-free method that enhances the alignment between image and text embeddings in CLIP by introducing class concept representations and class-guided visual features. Specifically, we employ an attention-based aggregation mechanism that projects visual features into the text embedding space, allowing for fine-grained alignment with class semantics. These components capture co-occurrence and semantic patterns of target objects from large-scale textual descriptions, enabling more accurate and context-aware recognition without reliance on handcrafted prompts.

    Second, to improve the localization capability of CLIP in open-vocabulary semantic segmentation, we introduce the Class Distribution-induced Attention Map (CDAM). CDAM refines noisy and coarse attention maps by computing semantic similarity between patches and propagating class-consistent attention across spatial regions based on the class distribution of patches. This leads to more accurate localization of target categories, all without requiring any additional training.

    Lastly, for high-resolution image generation from detailed, human-centric captions, we propose a hierarchical pipeline that integrates large language models to generate keypoint-box layouts from text. These spatial priors guide a multi-stage diffusion process, addressing grounding limitations and improving coherence in complex scene synthesis.

    Collectively, these contributions demonstrate that enhancing image-text alignment, even without retraining, can dramatically enhance the performance, adaptability, and interpretability of large multi-modal models. This positions them as powerful and general-purpose visual reasoning systems for real-world, open-domain vision tasks.
    번역하기

    With the rapid advancement of artificial intelligence, AI are becoming increasingly embedded in everyday life. This shift has created a growing demand for vision understanding systems that extend beyond task-specific applications and can address diver...

    With the rapid advancement of artificial intelligence, AI are becoming increasingly embedded in everyday life. This shift has created a growing demand for vision understanding systems that extend beyond task-specific applications and can address diverse, complex real-world scenarios such as dynamic scene interpretation in autonomous driving, human-object interaction analysis in daily environments, and open-world object recognition that conventional models often struggle to address.

    To meet this demand, large multi-modal models, particularly vision-language models like CLIP and text-to-image diffusion models such as Stable Diffusion, have emerged as promising approaches. By integrating linguistic and visual knowledge, these models enable intuitive and context-aware visual understanding. However, a key challenge remains: effectively leveraging these models for novel and complex vision tasks without retraining or heavy computational cost, especially in the presence of imperfect image-text alignment.

    This dissertation aims to overcome the limitations of large multi-modal models by proposing training-free, zero-shot frameworks that directly enhance image-text alignment. This alignment serves as a critical foundation for enabling accurate recognition, spatial localization, and grounding in real-world vision applications. Although these models exhibit strong generalization across tasks, they often face fundamental limitations in precise visual grounding, contextual understanding, and fine-grained cross-modal alignment. These challenges become especially apparent in open-world settings involving dense scenes, ambiguous visual concepts, and detailed textual descriptions, which ultimately hinder the reliability and interpretability of multi-modal systems in human-centered environments.

    First, for zero-shot multi-label recognition, we propose a training-free method that enhances the alignment between image and text embeddings in CLIP by introducing class concept representations and class-guided visual features. Specifically, we employ an attention-based aggregation mechanism that projects visual features into the text embedding space, allowing for fine-grained alignment with class semantics. These components capture co-occurrence and semantic patterns of target objects from large-scale textual descriptions, enabling more accurate and context-aware recognition without reliance on handcrafted prompts.

    Second, to improve the localization capability of CLIP in open-vocabulary semantic segmentation, we introduce the Class Distribution-induced Attention Map (CDAM). CDAM refines noisy and coarse attention maps by computing semantic similarity between patches and propagating class-consistent attention across spatial regions based on the class distribution of patches. This leads to more accurate localization of target categories, all without requiring any additional training.

    Lastly, for high-resolution image generation from detailed, human-centric captions, we propose a hierarchical pipeline that integrates large language models to generate keypoint-box layouts from text. These spatial priors guide a multi-stage diffusion process, addressing grounding limitations and improving coherence in complex scene synthesis.

    Collectively, these contributions demonstrate that enhancing image-text alignment, even without retraining, can dramatically enhance the performance, adaptability, and interpretability of large multi-modal models. This positions them as powerful and general-purpose visual reasoning systems for real-world, open-domain vision tasks.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Challenges of Multi-modal Pre-trained Models in Vision Tasks 4
    • 1.1.1 Zero-shot Multi-label Recognition 4
    • 1.1.2 Open-vocabulary Semantic Segmentation 5
    • 1.1.3 High-resolution Human-centric Scene Generation 6
    • 1 Introduction 1
    • 1.1 Challenges of Multi-modal Pre-trained Models in Vision Tasks 4
    • 1.1.1 Zero-shot Multi-label Recognition 4
    • 1.1.2 Open-vocabulary Semantic Segmentation 5
    • 1.1.3 High-resolution Human-centric Scene Generation 6
    • 1.2 Contributions and Summary 7
    • 2 Zero-shot multi-label recognition with pre-trained vision-language model 11
    • 2.1 Introduction 11
    • 2.2 Related Work 14
    • 2.2.1 Multi-label image recognition with CLIP 14
    • 2.2.2 Training-free enhancement with CLIP 14
    • 2.3 Method 15
    • 2.3.1 Class Concept Representation 15
    • 2.3.2 Context-Guided Visual Feature 18
    • 2.3.3 Multi-Label Recognition with Class Concepts 20
    • 2.4 Experiments 21
    • 2.4.1 Implementation Details 21
    • 2.4.2 Evaluation on Limited Data Setting 21
    • 2.4.3 Ablation Study and Analysis 25
    • 2.5 Discussion 27
    • 2.6 Conclusion 30
    • 3 Open-vocabulary semantic segmentation with pre-trained vision-language model 31
    • 3.1 Introduction 31
    • 3.2 Related Work 34
    • 3.2.1 Open-Vocabulary Semantic Segmentation with Vision-Language Model 34
    • 3.2.2 Background Subtraction 35
    • 3.3 Methods 35
    • 3.3.1 Limitation of Semantic Segmentation with Vision-Language Model 36
    • 3.3.2 Class Distribution-induced Attention Map 37
    • 3.3.3 Entropy-based Background Thresholding 41
    • 3.4 Experiments 43
    • 3.4.1 Experimental Setup 43
    • 3.4.2 Comparison with State-of-the-Art Methods 44
    • 3.4.3 Ablation Study and Analysis 47
    • 3.4.4 Qualitative Results 50
    • 3.5 Discussion 53
    • 3.6 Conclusion 53
    • 4 High-resolution human-centric scene generation with text-to-image pretrained diffusion model 54
    • 4.1 Introduction 54
    • 4.2 Related work 58
    • 4.2.1 Controllable Human Generation 58
    • 4.2.2 Large Scene Generation Using Diffusion Models 58
    • 4.3 BeyondScene 59
    • 4.3.1 Hierarchical Keypoint-box Layout Generation 59
    • 4.3.2 Detailed Base Image Generation 61
    • 4.3.3 Instance-Aware Hierarchical Enlargement 63
    • 4.4 Experiments 65
    • 4.4.1 Experimental Settings 65
    • 4.4.2 Result 69
    • 4.4.3 Ablation Study 74
    • 4.5 Conclusion 76
    • 5 Conclusion 80
    • 5.1 Future Work 82
    • 5.1.1 Integration with Other Foundation Models for Supervised Performance 82
    • 5.1.2 Expansion to Diverse Application Domains 82
    • 5.1.3 Extension Toward Multi-modal Large Language Models 83
    • Abstract (In Korean) 104
    • 감사의 글 106
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼