RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation = 다중 객체 이미지 생성을 위한 텍스트 임베딩에서의 방향성 객체 분리

    한글로보기

    https://www.riss.kr/link?id=T17450129

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.
    번역하기

    Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in ...

    Recent progress in text-to-image (T2I) generative models has led to significant improvements in generating high-quality images aligned with text prompts. However, these models still struggle with prompts involving multiple objects, often resulting in object neglect or object mixing. Through extensive studies, we identify four problematic scenarios, Similar Shapes, Similar Textures, Dissimilar Background Biases, and Many Objects, where inter-object relationships frequently lead to such failures. Motivated by two key observations about CLIP embeddings, we propose DOS (Directional Object Separation), a method that modifies three types of CLIP text embeddings before passing them into text-to-image models. Experimental results show that DOS consistently improves the success rate of multi-object image generation and reduces object mixing. In human evaluations, DOS significantly outperforms four competing methods, receiving 26.24%-43.04% more votes across four benchmarks. These results highlight DOS as a practical and effective solution for improving multi-object image generation.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 텍스트-이미지 변환 (T2I) 생성 모델의 발전은 텍스트 프롬프트와 일치하는 고품질 이미지 생성에 상당한 개선을 가져왔다. 하지만 이러한 모델들은 여전히 다중 객체를 포함하는 프롬프트 처리에는 어려움을 겪으며, 이는 종종 객체 누락이나 객체 혼합 문제로 이어진다. 본 논문은 광범위한 실험을 통해, 객체 간의 관계가 이러한 실패를 자주 유발하는 네 가지 문제적 시나리오 (유사한 형태, 유사한 질감, 상이한 배경 편향, 다수의 객체)를 규명했다. CLIP 임베딩에 대한 두 가지 핵심 관찰에 기반하여, 본 논문은 텍스트-이미지 변환 생성 모델에 입력하기 전에 세 가지 유형의 CLIP 텍스트 임베딩을 수정하는 방법인 DOS (방향성 객체 분리)를 제안한다. 실험 결과는 DOS가 다중 객체 이미지 생성의 성공률을 일관되게 향상시키고 객체 혼합을 감소시킨다는 것을 보여준다. 인간 평가에서 DOS는 4개의 경쟁 방법보다 뛰어난 성능을 보였으며, 4개의 벤치마크에서 26.24%-43.04% 더 많은 표를 받았다. 이러한 결과는 DOS가 다중 객체 이미지 생성을 위한 실용적이고 효과적인 해결책임을 시사한다.
    번역하기

    최근 텍스트-이미지 변환 (T2I) 생성 모델의 발전은 텍스트 프롬프트와 일치하는 고품질 이미지 생성에 상당한 개선을 가져왔다. 하지만 이러한 모델들은 여전히 다중 객체를 포함하는 프롬...

    최근 텍스트-이미지 변환 (T2I) 생성 모델의 발전은 텍스트 프롬프트와 일치하는 고품질 이미지 생성에 상당한 개선을 가져왔다. 하지만 이러한 모델들은 여전히 다중 객체를 포함하는 프롬프트 처리에는 어려움을 겪으며, 이는 종종 객체 누락이나 객체 혼합 문제로 이어진다. 본 논문은 광범위한 실험을 통해, 객체 간의 관계가 이러한 실패를 자주 유발하는 네 가지 문제적 시나리오 (유사한 형태, 유사한 질감, 상이한 배경 편향, 다수의 객체)를 규명했다. CLIP 임베딩에 대한 두 가지 핵심 관찰에 기반하여, 본 논문은 텍스트-이미지 변환 생성 모델에 입력하기 전에 세 가지 유형의 CLIP 텍스트 임베딩을 수정하는 방법인 DOS (방향성 객체 분리)를 제안한다. 실험 결과는 DOS가 다중 객체 이미지 생성의 성공률을 일관되게 향상시키고 객체 혼합을 감소시킨다는 것을 보여준다. 인간 평가에서 DOS는 4개의 경쟁 방법보다 뛰어난 성능을 보였으며, 4개의 벤치마크에서 26.24%-43.04% 더 많은 표를 받았다. 이러한 결과는 DOS가 다중 객체 이미지 생성을 위한 실용적이고 효과적인 해결책임을 시사한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • 1.1 Dissertation Outline 2
    • 1.2 Related Publication 3
    • Chapter 2. Background 4
    • 2.1 Text-to-Image Generative Models 4
    • Chapter 1. Introduction 1
    • 1.1 Dissertation Outline 2
    • 1.2 Related Publication 3
    • Chapter 2. Background 4
    • 2.1 Text-to-Image Generative Models 4
    • 2.2 Information Mix-up in CLIP Text Embeddings 4
    • 2.3 Directional Information in CLIP Text Embeddings 5
    • Chapter 3. DOS: Directional Object Separation in Text Embeddings for Multi-Object Image Generation 6
    • 3.1 Background and Overview 6
    • 3.2 Related Work 8
    • 3.3 Preliminary Analysis 9
    • 3.4 Method 11
    • 3.5 Experiments 17
    • 3.5.1 Experimental setup 17
    • 3.5.2 Experimental results 19
    • 3.5.3 Analysis 22
    • 3.6 Limitation 27
    • Chapter 4. Discussion 28
    • 4.1 Combining DOS With Attend-and-Excite 28
    • 4.2 Implicit Preservation of Attribute Binding 30
    • 4.3 Directional Encoding Across CLIP Text Embedding Types 33
    • 4.4 Generalization to Diverse Object Categories 35
    • 4.5 Robustness of Evaluation Results 36
    • 4.6 Validation on Newer Architectures 37
    • 4.7 Training Data Biases Underlying Generation Failures 38
    • Chapter 5. Conclusion 41
    • Bibliography 42
    • Appendix 46
    • A Details on Preliminary Analysis 47
    • B Experimental Details 48
    • B.1 Implementation details 48
    • B.2 Benchmark details 49
    • B.3 Evaluation metrics 51
    • B.4 Human preference study 51
    • C Additional Qualitative Results 54
    • Abstract (In Korean) 59
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼