RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Real-Image-Driven Image Synthesis Using Text-to-Image Foundation Models: Improving Personalization and Advancing Editing = 텍스트-이미지 생성 모델을 활용한 실제 이미지 기반 이미지 합성: 개인화 향상 및 편집 기술 증진 연구

    한글로보기

    https://www.riss.kr/link?id=T17450839

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Text-to-image foundation models have achieved remarkable progress, evolving from U-Net–based diffusion models to multimodal diffusion transformers (MM-DiT)-based rectified flow models. These advances have enabled not only open-ended text-to-image generation but also real-image-driven applications such as personalization and real-image editing.
    Despite rapid progress, key challenges remain. In personalization, non-subject elements in reference images often become entangled with the subject embedding, reducing fidelity and controllability. In real-image editing, most prior approaches were developed for U-Net architectures and do not transfer effectively to MM-DiT.
    This dissertation addresses these challenges through two contributions.
    First, we propose Selectively Informative Description (SID), which mitigates undesired embedding entanglement by employing training descriptions where the subject is identified only by its class, while non-subject elements are provided with informative descriptions.
    Second, we introduce ReFlex, a real-image editing method tailored to rectified flow–based models. ReFlex identifies key MM-DiT features for editing, proposes mid-step inversion for structure-preserving feature extraction, and incorporates attention adaptation techniques to balance editability with source preservation.
    Together, these contributions expand the capabilities of text-to-image models toward faithful and controllable real-image-driven image synthesis, improving personalization and editing while offering insights relevant to future large-scale systems.
    번역하기

    Text-to-image foundation models have achieved remarkable progress, evolving from U-Net–based diffusion models to multimodal diffusion transformers (MM-DiT)-based rectified flow models. These advances have enabled not only open-ended text-to-image ge...

    Text-to-image foundation models have achieved remarkable progress, evolving from U-Net–based diffusion models to multimodal diffusion transformers (MM-DiT)-based rectified flow models. These advances have enabled not only open-ended text-to-image generation but also real-image-driven applications such as personalization and real-image editing.
    Despite rapid progress, key challenges remain. In personalization, non-subject elements in reference images often become entangled with the subject embedding, reducing fidelity and controllability. In real-image editing, most prior approaches were developed for U-Net architectures and do not transfer effectively to MM-DiT.
    This dissertation addresses these challenges through two contributions.
    First, we propose Selectively Informative Description (SID), which mitigates undesired embedding entanglement by employing training descriptions where the subject is identified only by its class, while non-subject elements are provided with informative descriptions.
    Second, we introduce ReFlex, a real-image editing method tailored to rectified flow–based models. ReFlex identifies key MM-DiT features for editing, proposes mid-step inversion for structure-preserving feature extraction, and incorporates attention adaptation techniques to balance editability with source preservation.
    Together, these contributions expand the capabilities of text-to-image models toward faithful and controllable real-image-driven image synthesis, improving personalization and editing while offering insights relevant to future large-scale systems.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    언어 기반 이미지 생성 모델은 최근 빠르게 발전하면서, U-Net 기반 확산 모델에서 멀티모달 디퓨전 트랜스포머(MM-DiT) 기반의 rectified flow 모델로 확장되어 왔다. 그 결과, 단순히 텍스트로 이미지를 생성하는 수준을 넘어 개인화와 실제 이미지 편집처럼 입력 이미지가 주어지는 응용까지 가능해졌다. 다만 발전 속도에 비해 해결되지 않은 문제도 분명히 남아 있다. 개인화에서는 참조 이미지에 함께 담긴 배경이나 소품 같은 비대상 요소가 대상 임베딩에 섞여 들어가, 대상의 정체성이 흐려지거나 원하는 대로 제어하기 어려운 경우가 많다. 실제 이미지 편집에서는 그동안 U-Net 구조를 기준으로 만들어진 방법들이 MM-DiT 구조에서는 그대로 잘 동작하지 않는 한계가 있다.
    본 학위논문은 이러한 문제를 두 가지 기여로 다룬다. 먼저, 선택적으로 정보적인 설명(SID)을 제안한다. 이 방법은 학습 설명에서 대상을 클래스 수준으로만 지정하고, 대신 배경이나 비대상 요소는 충분히 자세히 기술하는 방식을 사용한다. 이를 통해 대상과 비대상 정보가 한데 엮이는 현상을 줄이고, 개인화 결과의 충실도와 제어 가능성을 개선한다. 다음으로, rectified flow 기반 모델에 맞춘 실제 이미지 편집 방법인 ReFlex를 제안한다. ReFlex는 편집에 유용한 MM-DiT 내부 특징을 분석해 핵심 구성 요소를 정리하고, 구조를 잘 보존하는 특징을 얻기 위해 중간 단계까지의 역추정을 이용하는 mid-step inversion을 도입한다. 또한 편집이 잘 되면서도 원본이 과도하게 훼손되지 않도록 attention adaptation 기법을 적용해, 편집 가능성과 원본 보존 사이의 균형을 맞춘다.
    이러한 결과는 텍스트-이미지 모델의 활용 범위를 개방형 생성에서 실제 이미지 기반 개인화와 편집으로 넓히는 데 기여한다. 더 나아가 앞으로 대규모 모델을 학습하고 활용하는 흐름 속에서도, 데이터 설명을 어떻게 구성할지, 그리고 모델 구조에 맞는 편집 방법을 어떻게 설계할지에 대한 기준을 제공한다.
    번역하기

    언어 기반 이미지 생성 모델은 최근 빠르게 발전하면서, U-Net 기반 확산 모델에서 멀티모달 디퓨전 트랜스포머(MM-DiT) 기반의 rectified flow 모델로 확장되어 왔다. 그 결과, 단순히 텍스트로 이...

    언어 기반 이미지 생성 모델은 최근 빠르게 발전하면서, U-Net 기반 확산 모델에서 멀티모달 디퓨전 트랜스포머(MM-DiT) 기반의 rectified flow 모델로 확장되어 왔다. 그 결과, 단순히 텍스트로 이미지를 생성하는 수준을 넘어 개인화와 실제 이미지 편집처럼 입력 이미지가 주어지는 응용까지 가능해졌다. 다만 발전 속도에 비해 해결되지 않은 문제도 분명히 남아 있다. 개인화에서는 참조 이미지에 함께 담긴 배경이나 소품 같은 비대상 요소가 대상 임베딩에 섞여 들어가, 대상의 정체성이 흐려지거나 원하는 대로 제어하기 어려운 경우가 많다. 실제 이미지 편집에서는 그동안 U-Net 구조를 기준으로 만들어진 방법들이 MM-DiT 구조에서는 그대로 잘 동작하지 않는 한계가 있다.
    본 학위논문은 이러한 문제를 두 가지 기여로 다룬다. 먼저, 선택적으로 정보적인 설명(SID)을 제안한다. 이 방법은 학습 설명에서 대상을 클래스 수준으로만 지정하고, 대신 배경이나 비대상 요소는 충분히 자세히 기술하는 방식을 사용한다. 이를 통해 대상과 비대상 정보가 한데 엮이는 현상을 줄이고, 개인화 결과의 충실도와 제어 가능성을 개선한다. 다음으로, rectified flow 기반 모델에 맞춘 실제 이미지 편집 방법인 ReFlex를 제안한다. ReFlex는 편집에 유용한 MM-DiT 내부 특징을 분석해 핵심 구성 요소를 정리하고, 구조를 잘 보존하는 특징을 얻기 위해 중간 단계까지의 역추정을 이용하는 mid-step inversion을 도입한다. 또한 편집이 잘 되면서도 원본이 과도하게 훼손되지 않도록 attention adaptation 기법을 적용해, 편집 가능성과 원본 보존 사이의 균형을 맞춘다.
    이러한 결과는 텍스트-이미지 모델의 활용 범위를 개방형 생성에서 실제 이미지 기반 개인화와 편집으로 넓히는 데 기여한다. 더 나아가 앞으로 대규모 모델을 학습하고 활용하는 흐름 속에서도, 데이터 설명을 어떻게 구성할지, 그리고 모델 구조에 맞는 편집 방법을 어떻게 설계할지에 대한 기준을 제공한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • 1.1. Dissertation Outline 3
    • 1.2. Related Publications 4
    • Chapter 2. Background 5
    • Chapter 1. Introduction 1
    • 1.1. Dissertation Outline 3
    • 1.2. Related Publications 4
    • Chapter 2. Background 5
    • 2.0.1. Text-to-image diffusion models 5
    • 2.0.2. Text-to-image rectified-flow models 6
    • Chapter 3. Selectively Informative Description can Reduce Undesired Embedding Entanglements in Text-to-Image Personalization 8
    • 3.1. Introduction 8
    • 3.2. Contributions 11
    • 3.3. Related works 12
    • 3.4. Method 12
    • 3.4.1. SID for reducing embedding entanglement 14
    • 3.4.2. VLM for generating SID 18
    • 3.5. Experiments 18
    • 3.5.1. Qualitative comparison 18
    • 3.5.2. Comparison on highly, moderately, and low-biased scenarios 22
    • 3.5.3. Analysis of cross-attention map 23
    • 3.5.4. Analysis of three key measures 27
    • 3.5.5. Human evaluation 29
    • 3.6. Discussion 37
    • 3.6.1. Negative prompt and segmentation 37
    • 3.6.2. Enhancing subject editing 38
    • 3.6.3. SID Generation via Open-Source MLLMs and Iterative Self-Correction 39
    • 3.6.4. Conclusion of SID 44
    • 3.6.5. Limitations of SID 44
    • Chapter 4. ReFlex: Text-Guided Editing of Real Images in Rectified Flow via Mid-Step Feature Extraction and Attention Adaptation 46
    • 4.1. Introduction 46
    • 4.2. Contributions 50
    • 4.3. Related works 51
    • 4.3.1. Image editing for diffusion models 51
    • 4.3.2. Image editing for rectified flow models 52
    • 4.4. Preliminaries 52
    • 4.4.1. FLUX 52
    • 4.5. Method 53
    • 4.5.1. Observations: three key features in MM-DiT 55
    • 4.5.2. Mid-step feature extraction 57
    • 4.5.3. Two feature adaptation techniques 60
    • 4.5.4. Mask generation for latent blending 63
    • 4.6. Experiments 63
    • 4.6.1. Qualitative evaluation 64
    • 4.6.2. Quantitative Evaluation 71
    • 4.6.3. User study 75
    • 4.6.4. Ablations 78
    • 4.7. Limitations 84
    • 4.8. Discussion 84
    • Chapter 5. Discussion 87
    • Chapter 6. Conclusion 89
    • Bibliography 91
    • Appendices 101
    • A. Implementations and datasets in detail for Section 3 102
    • A.1. Descriptions in style re-contextualization 102
    • A.2. Model implementations 102
    • A.3. Dataset details 103
    • A.4. VLM details 105
    • B. Selectively describing facial expression via SID 106
    • C. Additional experiment results for Section 3 108
    • C.1. Four description cases 108
    • C.2. Instruction-following VLMs 108
    • C.3. Enhancement by SID 111
    • C.4. Enhancement for a single reference image 111
    • C.5. Negative prompt and segmentation 111
    • D. Implementation details for Section 4 125
    • D.1. Implementation details for feature analysis 125
    • D.2. Implementation details for ReFlex 125
    • D.3. Baseline implementation 129
    • E. Additional quantitative comparison for Section 4.6.2 130
    • E.1. PIE-Bench 130
    • Abstract (In Korean) 132
    • Acknowledgement 133
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼