인공지능의 급속한 발전으로 AI는 일상생활 속에 점점 더 깊이 통합되고 있습니다. 이러한 변화는 특정 분야에 국한되지 않고, 보다 다양한 복잡한 현실 세계의 시각 과제를 다룰 수 있는 시...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
인공지능의 급속한 발전으로 AI는 일상생활 속에 점점 더 깊이 통합되고 있습니다. 이러한 변화는 특정 분야에 국한되지 않고, 보다 다양한 복잡한 현실 세계의 시각 과제를 다룰 수 있는 시...
인공지능의 급속한 발전으로 AI는 일상생활 속에 점점 더 깊이 통합되고 있습니다. 이러한 변화는 특정 분야에 국한되지 않고, 보다 다양한 복잡한 현실 세계의 시각 과제를 다룰 수 있는 시각 이해 및 분석 시스템에 대한 수요 증가로 이어지고 있습니다. 예를 들어, 자율주행에서의 동적 장면 해석, 일상 환경에서의 인간-객체 상호작용 분석, 기존의 과제 특화형 모델들이 어려움을 겪는 개방형(open-world) 객체 인식 과제 등이 이에 해당합니다.
이러한 수요를 충족하기 위해, CLIP과 같은 비전-언어 모델 및 Stable Diffusion과 같은 텍스트-이미지 확산 모델 등 대규모 사전학습 멀티모달 모델이 주목받고 있습니다. 이들 모델은 언어적 지식과 시각적 지식을 통합함으로써 보다 직관적이고 맥락을 이해하며 인간 친화적인 방식으로 이미지를 해석할 수 있는 가능성을 제시합니다. 그러나 여전히 중요한 과제는, 특히 이미지-텍스트 정렬이 불완전한 경우, 이러한 모델들을 새로운 시각 작업에 대해 재학습 없이 그리고 과도한 연산 자원 소모 없이 어떻게 효과적으로 활용할 수 있는가입니다.
본 학위논문은 이러한 한계를 극복하고자, 학습이 필요 없는(training-free) 제로샷(zero-shot) 기반 프레임워크를 통해 이미지-텍스트 정렬을 직접 강화하는 방식을 제안합니다. 이미지-텍스트 정렬은 시각 과제에서 인식, 공간적 정위(localization), 의미적 정합성(grounding)을 가능하게 하는 핵심 기반입니다. 하지만 대규모 멀티모달 모델은 높은 일반화 능력에도 불구하고, 정밀한 시각적 그라운딩, 문맥 이해, 세밀한 모달 간 정렬에 있어 구조적 한계를 보입니다. 이러한 문제는 특히 밀집된 장면, 모호한 시각 개념, 상세한 텍스트 설명이 요구되는 개방형 환경에서 더욱 뚜렷하게 나타나며, 이는 멀티모달 시스템의 신뢰성과 해석 가능성을 저해하는 요인으로 작용합니다.
첫째, 제로샷 다중 레이블 인식(multi-label recognition)을 위해, CLIP 기반 이미지-텍스트 임베딩 정렬을 강화하는 학습 없는 접근법을 제안합니다. 본 방법은 클래스 개념 표현(class concept representation)과 클래스 유도 시각 표현(class-guided visual feature)을 활용하여, 이미지 특징을 텍스트 임베딩 공간으로 사영하는 주의(attention) 기반 정렬 메커니즘을 도입합니다. 이로써 대규모 텍스트 설명 내 동시 발생(co-occurrence) 및 의미적 패턴을 포착하여 수작업 프롬프트 없이도 더 정확하고 문맥에 부합하는 인식이 가능합니다.
둘째, 개방형 어휘 의미 분할(open-vocabulary semantic segmentation)에서 CLIP의 위치 정렬 성능을 개선하기 위해 클래스 분포 기반 어텐션 맵(Class Distribution-induced Attention Map, CDAM)을 제안합니다. CDAM은 이미지 패치 간 클래스 분포를 바탕으로 의미적 유사도를 계산하고, 클래스 일관성을 가진 주의 정보를 공간적으로 전파함으로써 기존의 노이즈 많고 조밀하지 못한 어텐션 맵을 정제합니다. 또한 다중 스케일 이미지 패치 처리, 텍스트 프롬프트 확장, 엔트로피 기반 배경 분리 기법을 결합하여 학습 없이도 정확한 제로샷 의미 분할을 실현합니다.
셋째, 인간 중심 캡션으로부터 고해상도 이미지를 생성하기 위해, 대형 언어 모델을 활용하여 텍스트로부터 키포인트-박스 레이아웃을 생성하고, 이를 공간적 사전 지식으로 활용하는 계층적 파이프라인을 제안합니다. 해당 레이아웃은 다단계 확산 생성 과정을 안내하며, 세밀한 시각적 정합성과 복잡한 장면 내 일관성을 확보합니다.
이러한 연구 기여들은 학습 없이도 이미지-텍스트 정렬을 강화함으로써 대규모 멀티모달 모델의 성능, 적응력, 해석 가능성을 획기적으로 향상시킬 수 있음을 보여줍니다. 이는 해당 모델들이 현실 세계의 개방형 시각 과제를 위한 유연하고 강력한 시각 추론 시스템으로 발전할 수 있는 실용적 가능성을 제시합니다.
다국어 초록 (Multilingual Abstract)
With the rapid advancement of artificial intelligence, AI are becoming increasingly embedded in everyday life. This shift has created a growing demand for vision understanding systems that extend beyond task-specific applications and can address diver...
With the rapid advancement of artificial intelligence, AI are becoming increasingly embedded in everyday life. This shift has created a growing demand for vision understanding systems that extend beyond task-specific applications and can address diverse, complex real-world scenarios such as dynamic scene interpretation in autonomous driving, human-object interaction analysis in daily environments, and open-world object recognition that conventional models often struggle to address.
To meet this demand, large multi-modal models, particularly vision-language models like CLIP and text-to-image diffusion models such as Stable Diffusion, have emerged as promising approaches. By integrating linguistic and visual knowledge, these models enable intuitive and context-aware visual understanding. However, a key challenge remains: effectively leveraging these models for novel and complex vision tasks without retraining or heavy computational cost, especially in the presence of imperfect image-text alignment.
This dissertation aims to overcome the limitations of large multi-modal models by proposing training-free, zero-shot frameworks that directly enhance image-text alignment. This alignment serves as a critical foundation for enabling accurate recognition, spatial localization, and grounding in real-world vision applications. Although these models exhibit strong generalization across tasks, they often face fundamental limitations in precise visual grounding, contextual understanding, and fine-grained cross-modal alignment. These challenges become especially apparent in open-world settings involving dense scenes, ambiguous visual concepts, and detailed textual descriptions, which ultimately hinder the reliability and interpretability of multi-modal systems in human-centered environments.
First, for zero-shot multi-label recognition, we propose a training-free method that enhances the alignment between image and text embeddings in CLIP by introducing class concept representations and class-guided visual features. Specifically, we employ an attention-based aggregation mechanism that projects visual features into the text embedding space, allowing for fine-grained alignment with class semantics. These components capture co-occurrence and semantic patterns of target objects from large-scale textual descriptions, enabling more accurate and context-aware recognition without reliance on handcrafted prompts.
Second, to improve the localization capability of CLIP in open-vocabulary semantic segmentation, we introduce the Class Distribution-induced Attention Map (CDAM). CDAM refines noisy and coarse attention maps by computing semantic similarity between patches and propagating class-consistent attention across spatial regions based on the class distribution of patches. This leads to more accurate localization of target categories, all without requiring any additional training.
Lastly, for high-resolution image generation from detailed, human-centric captions, we propose a hierarchical pipeline that integrates large language models to generate keypoint-box layouts from text. These spatial priors guide a multi-stage diffusion process, addressing grounding limitations and improving coherence in complex scene synthesis.
Collectively, these contributions demonstrate that enhancing image-text alignment, even without retraining, can dramatically enhance the performance, adaptability, and interpretability of large multi-modal models. This positions them as powerful and general-purpose visual reasoning systems for real-world, open-domain vision tasks.
목차 (Table of Contents)