RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Reference-Conditioned Visual Content Generation with Text-Conditioned Diffusion Models = 텍스트 조건부 확산 모델을 이용한 참조 기반 시각 컨텐츠 생성

    한글로보기

    https://www.riss.kr/link?id=T17450851

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Text-conditioned visual content generation has enabled remarkable progress in synthesizing images and videos guided by natural language. However, describing a specific and unique concept solely through text remains inherently ambiguous—text alone often fails to capture fine-grained appearance details or personalized characteristics desired by the user. To overcome this limitation, reference-conditioned generation has emerged, leveraging visual cues from provided reference data to guide the synthesis process. This dissertation explores single-reference-conditioned visual content generation through three complementary adaptation perspectives—feature-level, input-level, and loss-level—that collectively enhance fidelity, temporal coherence, and controllability in text-conditioned generation.

    As a feature-level integration approach, we introduce Edit-A-Video, which generates edited videos by extending text-to-image diffusion models to the video domain through selective temporal fine-tuning and structure-guided inversion. Beyond these components, the framework performs feature blending that explicitly integrates motionaware information across frames. This design enables temporally smooth and motion preserving video editing from a single reference sequence, faithfully reflecting both spatial structure and appearance dynamics.

    As an input-level integration approach, we propose Diptych Prompting, which reformulates zero-shot subject-driven generation as an inpainting-based diptych composition problem. Building upon recent high-capacity text-to-image diffusion models and inpainting architectures, our framework directly embeds the reference image into the inpainting input, enabling seamless subject integration and context-aware synthesis without any task-specific retraining, thereby supporting flexible identity transfer across diverse visual domains.

    As a loss-level integration approach, we present Subject Fidelity Optimization (SFO), which enhances identity preservation through negative-guided comparative learning. This objective encourages the model to explicitly distinguish between faithful and degraded generations, leading to improved fine-grained subject consistency and visual robustness under diverse conditions.

    Collectively, these contributions advance the scalability and reliability of reference-conditioned generation frameworks, bridging the gap between descriptive language and personalized visual content. The proposed methodologies demonstrate the effective use of text-conditioned generative frameworks to achieve controllable, identity-preserving, and context-aware generation across both image and video domains. By establishing diverse perspectives on feature-, input-, and loss-level adaptation, this dissertation lays the groundwork for more generalizable and expressive single-reference-conditioned generation, fostering richer and more accessible forms of human–AI creative collaboration. Beyond their immediate technical improvements, these frameworks provide a conceptual foundation for adaptive and human-controllable generative systems, pointing toward future directions in multimodal, interactive visual content generation.
    번역하기

    Text-conditioned visual content generation has enabled remarkable progress in synthesizing images and videos guided by natural language. However, describing a specific and unique concept solely through text remains inherently ambiguous—text alone of...

    Text-conditioned visual content generation has enabled remarkable progress in synthesizing images and videos guided by natural language. However, describing a specific and unique concept solely through text remains inherently ambiguous—text alone often fails to capture fine-grained appearance details or personalized characteristics desired by the user. To overcome this limitation, reference-conditioned generation has emerged, leveraging visual cues from provided reference data to guide the synthesis process. This dissertation explores single-reference-conditioned visual content generation through three complementary adaptation perspectives—feature-level, input-level, and loss-level—that collectively enhance fidelity, temporal coherence, and controllability in text-conditioned generation.

    As a feature-level integration approach, we introduce Edit-A-Video, which generates edited videos by extending text-to-image diffusion models to the video domain through selective temporal fine-tuning and structure-guided inversion. Beyond these components, the framework performs feature blending that explicitly integrates motionaware information across frames. This design enables temporally smooth and motion preserving video editing from a single reference sequence, faithfully reflecting both spatial structure and appearance dynamics.

    As an input-level integration approach, we propose Diptych Prompting, which reformulates zero-shot subject-driven generation as an inpainting-based diptych composition problem. Building upon recent high-capacity text-to-image diffusion models and inpainting architectures, our framework directly embeds the reference image into the inpainting input, enabling seamless subject integration and context-aware synthesis without any task-specific retraining, thereby supporting flexible identity transfer across diverse visual domains.

    As a loss-level integration approach, we present Subject Fidelity Optimization (SFO), which enhances identity preservation through negative-guided comparative learning. This objective encourages the model to explicitly distinguish between faithful and degraded generations, leading to improved fine-grained subject consistency and visual robustness under diverse conditions.

    Collectively, these contributions advance the scalability and reliability of reference-conditioned generation frameworks, bridging the gap between descriptive language and personalized visual content. The proposed methodologies demonstrate the effective use of text-conditioned generative frameworks to achieve controllable, identity-preserving, and context-aware generation across both image and video domains. By establishing diverse perspectives on feature-, input-, and loss-level adaptation, this dissertation lays the groundwork for more generalizable and expressive single-reference-conditioned generation, fostering richer and more accessible forms of human–AI creative collaboration. Beyond their immediate technical improvements, these frameworks provide a conceptual foundation for adaptive and human-controllable generative systems, pointing toward future directions in multimodal, interactive visual content generation.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 텍스트 기반 시각 컨텐츠 생성 기술은 자연어로부터 이미지와 비디오를 생성하는 데 있어 눈에 띄는 발전을 이루어 왔습니다. 그러나 특정하고 개별적인 개 념을 텍스트만으로 정확히 묘사하는 데에는 본질적인 한계가 존재합니다. 텍스트는 사용자가 의도하는 미세한 외형적 특징이나 개인화된 속성을 완전히 전달하기 어렵 기 때문입니다. 이러한 한계를 보완하기 위해, 참조 이미지나 영상을 활용하여 보다 정교한 시각적 단서를 제공하는 참조 기반 생성(reference-conditioned generation) 기법이 활발히 연구되고 있습니다.

    본 학위논문에서는 텍스트 조건부 생성 모델을 기반으로 한 단일 참조 기반 (single-reference-conditioned) 시각 컨텐츠 생성을 다루며, 이를 특징 수준(featurelevel), 입력 수준(input-level), 손실 함수 수준(loss-level)의 세 가지 상보적인 적응 관점에서 체계적으로 탐구합니다. 이러한 관점은 고유한 장점을 통해 생성된 이미 지 · 비디오의 정확성(fidelity), 시간적 일관성(temporal coherence), 제어 가능성(controllability)을 획기적으로 향상시키는 것을 목표로 합니다.

    먼저, 특징 수준 통합(feature-level integration) 접근으로서 Edit-A-Video를 제안 합니다. 이 방법은 기존 텍스트 기반 이미지 확산 모델을 선택적 시간축 파인튜닝과 구조 기반 인버전 기법을 통해 비디오 영역으로 확장하였으며, 프레임 간 모션 정보 를 명시적으로 활용하는 특징 블렌딩을 수행합니다. 이를 통해 단일 레퍼런스 비디 오만으로도 공간 구조와 외형적 변화를 자연스럽게 보존하는 시간적 일관성 높은 비디오 편집이 가능합니다.

    다음으로, 입력 수준 통합(input-level integration) 접근으로서 Diptych Prompting 을 제안합니다. 본 방법은 제로샷 기반 주체 생성(zero-shot subject-driven generation)을 인페인팅 구조의 이중 패널(diptych) 구성 문제로 재해석함으로써, 고용량 텍스트투-이미지 확산 모델과 인페인팅 아키텍처를 활용해 참조 이미지를 입력 단계에서 직접 삽입할 수 있도록 합니다. 이를 통해 별도의 재학습 없이도 자연스럽고 문맥을 고려한 정체성 보존 생성이 가능하며, 다양한 시각 도메인에 대해 유연한 아이덴티 티 전이가 가능합니다.

    마지막으로, 손실 수준 통합(loss-level integration) 접근으로서 Subject Fidelity Optimization (SFO)을 제안합니다. SFO는 부정 예시를 활용한 비교 학습(negativeguided comparative learning)을 통해, 모델이 올바른 주체 재현과 왜곡된 결과를 명확 히 구분하도록 유도합니다. 이를 통해 세밀한 시각적 일관성과 강건성을 강화하고, 주체 정체성 유지 성능을 크게 향상시킵니다.

    세 가지 관점을 통한 본 연구의 기여는 레퍼런스 기반 생성의 확장성과 안정성 을 크게 높이며, 텍스트가 지니는 모호성과 실제 사용자가 원하는 개인화된 시각 컨텐츠 간의 간극을 효과적으로 메우고자 합니다. 나아가 제안된 기법들은 이미지 와 비디오 생성 전반에서 제어 가능하고, 문맥 인지적인 생성이 가능함을 실험적으 로 입증하였습니다. 더불어, 본 논문은 향후 멀티모달 상호작용, 사용자 주도 생성 등으로 확장될 수 있는 기반을 제공하며, 단일 레퍼런스 기반 생성의 범용성 · 표현 력 · 사용자 적응성을 넓히기 위한 중요한 토대를 마련합니다.
    번역하기

    최근 텍스트 기반 시각 컨텐츠 생성 기술은 자연어로부터 이미지와 비디오를 생성하는 데 있어 눈에 띄는 발전을 이루어 왔습니다. 그러나 특정하고 개별적인 개 념을 텍스트만으로 정확히 ...

    최근 텍스트 기반 시각 컨텐츠 생성 기술은 자연어로부터 이미지와 비디오를 생성하는 데 있어 눈에 띄는 발전을 이루어 왔습니다. 그러나 특정하고 개별적인 개 념을 텍스트만으로 정확히 묘사하는 데에는 본질적인 한계가 존재합니다. 텍스트는 사용자가 의도하는 미세한 외형적 특징이나 개인화된 속성을 완전히 전달하기 어렵 기 때문입니다. 이러한 한계를 보완하기 위해, 참조 이미지나 영상을 활용하여 보다 정교한 시각적 단서를 제공하는 참조 기반 생성(reference-conditioned generation) 기법이 활발히 연구되고 있습니다.

    본 학위논문에서는 텍스트 조건부 생성 모델을 기반으로 한 단일 참조 기반 (single-reference-conditioned) 시각 컨텐츠 생성을 다루며, 이를 특징 수준(featurelevel), 입력 수준(input-level), 손실 함수 수준(loss-level)의 세 가지 상보적인 적응 관점에서 체계적으로 탐구합니다. 이러한 관점은 고유한 장점을 통해 생성된 이미 지 · 비디오의 정확성(fidelity), 시간적 일관성(temporal coherence), 제어 가능성(controllability)을 획기적으로 향상시키는 것을 목표로 합니다.

    먼저, 특징 수준 통합(feature-level integration) 접근으로서 Edit-A-Video를 제안 합니다. 이 방법은 기존 텍스트 기반 이미지 확산 모델을 선택적 시간축 파인튜닝과 구조 기반 인버전 기법을 통해 비디오 영역으로 확장하였으며, 프레임 간 모션 정보 를 명시적으로 활용하는 특징 블렌딩을 수행합니다. 이를 통해 단일 레퍼런스 비디 오만으로도 공간 구조와 외형적 변화를 자연스럽게 보존하는 시간적 일관성 높은 비디오 편집이 가능합니다.

    다음으로, 입력 수준 통합(input-level integration) 접근으로서 Diptych Prompting 을 제안합니다. 본 방법은 제로샷 기반 주체 생성(zero-shot subject-driven generation)을 인페인팅 구조의 이중 패널(diptych) 구성 문제로 재해석함으로써, 고용량 텍스트투-이미지 확산 모델과 인페인팅 아키텍처를 활용해 참조 이미지를 입력 단계에서 직접 삽입할 수 있도록 합니다. 이를 통해 별도의 재학습 없이도 자연스럽고 문맥을 고려한 정체성 보존 생성이 가능하며, 다양한 시각 도메인에 대해 유연한 아이덴티 티 전이가 가능합니다.

    마지막으로, 손실 수준 통합(loss-level integration) 접근으로서 Subject Fidelity Optimization (SFO)을 제안합니다. SFO는 부정 예시를 활용한 비교 학습(negativeguided comparative learning)을 통해, 모델이 올바른 주체 재현과 왜곡된 결과를 명확 히 구분하도록 유도합니다. 이를 통해 세밀한 시각적 일관성과 강건성을 강화하고, 주체 정체성 유지 성능을 크게 향상시킵니다.

    세 가지 관점을 통한 본 연구의 기여는 레퍼런스 기반 생성의 확장성과 안정성 을 크게 높이며, 텍스트가 지니는 모호성과 실제 사용자가 원하는 개인화된 시각 컨텐츠 간의 간극을 효과적으로 메우고자 합니다. 나아가 제안된 기법들은 이미지 와 비디오 생성 전반에서 제어 가능하고, 문맥 인지적인 생성이 가능함을 실험적으 로 입증하였습니다. 더불어, 본 논문은 향후 멀티모달 상호작용, 사용자 주도 생성 등으로 확장될 수 있는 기반을 제공하며, 단일 레퍼런스 기반 생성의 범용성 · 표현 력 · 사용자 적응성을 넓히기 위한 중요한 토대를 마련합니다.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Motivation for Reference-Conditioned Visual Content Generation 1
    • 1.2 Overview of Text-Conditioned Diffusion Models 4
    • 1 Introduction 1
    • 1.1 Motivation for Reference-Conditioned Visual Content Generation 1
    • 1.2 Overview of Text-Conditioned Diffusion Models 4
    • 1.3 Reference-Conditioned Generation with Text-Conditioned Diffusion Models 6
    • 1.3.1 Visual Content Editing 7
    • 1.3.2 Subject-Driven Generation 8
    • 1.4 Scope of Dissertation 9
    • 2 Background 11
    • 2.1 Generative Models 12
    • 2.1.1 Diffusion Models 12
    • 2.1.2 Flow Matching Models 16
    • 2.2 Representative Approaches for Reference-Conditioned Generation 19
    • 2.2.1 Fine-Tuning-Based Approaches 20
    • 2.2.2 Zero-Shot-Based Approaches 26
    • 3 Feature-level Reference Integration for Video Editing 30
    • 3.1 Introduction 30
    • 3.2 Method 33
    • 3.2.1 Framework 34
    • 3.2.2 Temporal-Consistent Blending 35
    • 3.2.3 Hyperparameters for Editing 37
    • 3.3 Experimental Results 39
    • 3.3.1 Implementation Details 39
    • 3.3.2 Baseline Comparisons 40
    • 3.3.3 Ablation 41
    • 3.4 Additional Experiments 44
    • 3.4.1 Additional Samples 44
    • 3.4.2 Extensions to Stable-Diffusion XL 44
    • 3.5 Concluding Remark 51
    • 4 Input-level Reference Integration for Subject-Driven Image Generation 52
    • 4.1 Introduction 52
    • 4.2 Background 55
    • 4.2.1 MM-DiT Architecture for Text-Conditioned Diffusion Models 55
    • 4.2.2 Text-Conditioned Inpainting 56
    • 4.3 Method 56
    • 4.3.1 Diptych Generation of FLUX 56
    • 4.3.2 Diptych Prompting Framework 58
    • 4.3.3 Reference Attention Enhancement 60
    • 4.4 Experimental Results 61
    • 4.4.1 Experimental Settings 61
    • 4.4.2 Baseline Comparisons 62
    • 4.4.3 Ablation Studies 65
    • 4.4.4 Applications 68
    • 4.5 Additional Experimental Results 68
    • 4.5.1 Comparison with Fine-Tuning-Based Method 68
    • 4.5.2 Additional Results 71
    • 4.5.3 Diptych Generation 71
    • 4.5.4 Background Removal Ablation 75
    • 4.5.5 Reference Attention Enhancement Ablation 77
    • 4.5.6 Stylized Image Generation 77
    • 4.5.7 Subject-Driven Image Editing 79
    • 4.5.8 Various Pose Generation 79
    • 4.6 Limitations 81
    • 4.7 Concluding Remark 82
    • 5 Loss-level Reference Integration for Subject-Driven Image Generation 84
    • 5.1 Introduction 84
    • 5.2 Background 87
    • 5.2.1 Comparative Learning Signals 87
    • 5.3 Method 88
    • 5.3.1 Problem Setting 88
    • 5.3.2 Subject Fidelity Optimization (SFO) 89
    • 5.3.3 Condition-Degradation Negative Sampling (CDNS) 92
    • 5.4 Experimental Results 93
    • 5.4.1 Experimental Settings 93
    • 5.4.2 Main Comparisons 94
    • 5.4.3 Ablation Study 98
    • 5.5 Additional Experimental Results 101
    • 5.5.1 Subject-Driven Generation 101
    • 5.5.2 Per-Subject Fine-Tuning Method Comparisons 101
    • 5.5.3 Additional Ablation 104
    • 5.5.4 Additional Samples 111
    • 5.6 Derivation 113
    • 5.6.1 Loss Derivation 113
    • 5.6.2 Conditional Mutual Information between Images and Subject Fidelity 115
    • 5.7 Concluding Remark 116
    • 6 Concluding Remark 118
    • 6.1 Summary of Dissertation 118
    • 6.2 Discussion and Outlook 119
    • 6.2.1 Dependency on Reference Data Quality 119
    • 6.2.2 Dependency on Backbone Text-Conditioned Models 120
    • 6.2.3 Toward Multi-Reference Generation 121
    • 6.2.4 Extension to the Video Domain 121
    • 6.2.5 Ethical Considerations 122
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼