RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Harmonizing Multimodal Inputs in Diffusion Model Customization = 멀티모달 입력 조화를 통한 확산 모델 커스터마이징 기법

    한글로보기

    https://www.riss.kr/link?id=T17450389

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    확산 모델은 텍스트, 이미지 등 다양한 입력 조건에서 고품질 이미지와 비디오를 생성할 수 있는 강력한 도구로 떠오르고 있다. 이러한 모델들이 현실 세계의 다양한 응용 분야에서 널리 사용됨에 따라, 커스터마이징 (customization) 작업 중요성이 점차 커지고 있다. 커스터마이징은 사전 훈련된 모델이 사용자가 가지고 있는 특정 객체를 이해하고 사용할 수 있도록 확장시키는 것을 목표로 한다. 일반적으로 커스터마이징 모델은 객체의 정보를 담고 있는 시각적 입력과 생성하고자 하는 이미지의 맥락 (context) 을 지시하는 텍스트 입력을 사용한다. 하지만 현재의 확산 모델 커스터마이징 연구들은 대부분 개별 입력 모달리티 (modality) 에만 집중하거나 이를 독립적으로 다루고 있어, 여러 모달리티 간의 복잡한 상호작용과 잠재적인 불일치를 종종 간과하고 있다.

    이 논문은 이러한 멀리모달 입력들 간 불일치에서 발생하는 중요한 문제들을 다루며, 특히 일반적인 객체 커스터마이징 (subject customization) 과 모션 커스터마이징 (motion customizatio) 의 맥락에서 이러한 문제들을 바라보고 있다. 이 논문에서는 시각적 입력과 텍스트 입력 간 어떻게 불일치가 발생하는지 탐구하며, 이러한 불일치로 인한 생성된 이미지 또는 동영상의 품질 저하를 초래할 수 있는지 조사한다. 또한, 이 논문은 이러한 문제들을 구체적인 여러 커스터마이징 분야에 적용하여 탐구하고 있다.

    먼저, 단일 이미지 기반의 제로샷 (zero-shot) 커스터마이징에서 불일치 문제를 살펴보았다. 제로샷 커스터마이징에서는 확산 모델이 추가 훈련 없이 개인화하고자 하는 객제를 생성하도록 한다. 우리는 개인화된 객제의 정보를 담고 있는 시각적 임베딩과 텍스트 임베딩 간의 관계를 조사하고, 이들 간의 충돌이 어떻게 생성된 이미지에 영향을 끼치는지 알아보고자 한다. 특히 객체가 특정한 동작을 취하는 이미지를 생성할 때 불일치가 어떻게 나타나는지 살펴보았다.

    다음으로, 모션 커스터마이징에서의 불일치 문제를 탐구하였다. 모션 커스터마이징은 목표로 하는 객체를 참조 동영상의 모션으로 확장하여, 해당 모션을 새로운 대상이 취하는 동영상을 생성하는 것을 목표로 한다. 모션 커스터마이징에서는 모션을 학습한 파라미터 (parameter) 와 텍스트에 따라 비디오가 생성된다. 이 때, 파라미터와 텍스트 간의 불일치가 생성된 비디오에 어떻게 영향을 미치는지 분석하였다.

    마지막으로, 이 논문에서는 시각적 입력과 텍스트 입력 간의 조화를 강화하는 방법을 제안하여, 생성된 결과물이 모션을 잘 유지함과 동시에 텍스트 입력에 충실한 수 있도록 하였다. 또한, 실험적으로 멀티모달 입력를 조화롭게 조정하는 메소드들이 확산 모델이 사용자 의도에 더 잘 부합하고, 다양한 작업에서 더 신뢰할 수 있으며, 다양한 응용 분야에 더 확장 가능하게 할 수 있음을 보여주었다.

    이 논문에서는 객체와 모션 커스터마이징에 중점을 두고 있지만, 내포하고 있는 통찰과 제안한 방법론은 멀티모달 입력을 받는 생성 모델들에 대한 더 넓은 의미를 지닐 수 있다. 또한, 커스터마이징을 멀티모달 입력 간의 관계에서 다시 바라봄으로써, 본 연구는 확산 모델 커스터마이징 분야에 새로운 관점을 제시하였다.
    번역하기

    확산 모델은 텍스트, 이미지 등 다양한 입력 조건에서 고품질 이미지와 비디오를 생성할 수 있는 강력한 도구로 떠오르고 있다. 이러한 모델들이 현실 세계의 다양한 응용 분야에서 널리 사...

    확산 모델은 텍스트, 이미지 등 다양한 입력 조건에서 고품질 이미지와 비디오를 생성할 수 있는 강력한 도구로 떠오르고 있다. 이러한 모델들이 현실 세계의 다양한 응용 분야에서 널리 사용됨에 따라, 커스터마이징 (customization) 작업 중요성이 점차 커지고 있다. 커스터마이징은 사전 훈련된 모델이 사용자가 가지고 있는 특정 객체를 이해하고 사용할 수 있도록 확장시키는 것을 목표로 한다. 일반적으로 커스터마이징 모델은 객체의 정보를 담고 있는 시각적 입력과 생성하고자 하는 이미지의 맥락 (context) 을 지시하는 텍스트 입력을 사용한다. 하지만 현재의 확산 모델 커스터마이징 연구들은 대부분 개별 입력 모달리티 (modality) 에만 집중하거나 이를 독립적으로 다루고 있어, 여러 모달리티 간의 복잡한 상호작용과 잠재적인 불일치를 종종 간과하고 있다.

    이 논문은 이러한 멀리모달 입력들 간 불일치에서 발생하는 중요한 문제들을 다루며, 특히 일반적인 객체 커스터마이징 (subject customization) 과 모션 커스터마이징 (motion customizatio) 의 맥락에서 이러한 문제들을 바라보고 있다. 이 논문에서는 시각적 입력과 텍스트 입력 간 어떻게 불일치가 발생하는지 탐구하며, 이러한 불일치로 인한 생성된 이미지 또는 동영상의 품질 저하를 초래할 수 있는지 조사한다. 또한, 이 논문은 이러한 문제들을 구체적인 여러 커스터마이징 분야에 적용하여 탐구하고 있다.

    먼저, 단일 이미지 기반의 제로샷 (zero-shot) 커스터마이징에서 불일치 문제를 살펴보았다. 제로샷 커스터마이징에서는 확산 모델이 추가 훈련 없이 개인화하고자 하는 객제를 생성하도록 한다. 우리는 개인화된 객제의 정보를 담고 있는 시각적 임베딩과 텍스트 임베딩 간의 관계를 조사하고, 이들 간의 충돌이 어떻게 생성된 이미지에 영향을 끼치는지 알아보고자 한다. 특히 객체가 특정한 동작을 취하는 이미지를 생성할 때 불일치가 어떻게 나타나는지 살펴보았다.

    다음으로, 모션 커스터마이징에서의 불일치 문제를 탐구하였다. 모션 커스터마이징은 목표로 하는 객체를 참조 동영상의 모션으로 확장하여, 해당 모션을 새로운 대상이 취하는 동영상을 생성하는 것을 목표로 한다. 모션 커스터마이징에서는 모션을 학습한 파라미터 (parameter) 와 텍스트에 따라 비디오가 생성된다. 이 때, 파라미터와 텍스트 간의 불일치가 생성된 비디오에 어떻게 영향을 미치는지 분석하였다.

    마지막으로, 이 논문에서는 시각적 입력과 텍스트 입력 간의 조화를 강화하는 방법을 제안하여, 생성된 결과물이 모션을 잘 유지함과 동시에 텍스트 입력에 충실한 수 있도록 하였다. 또한, 실험적으로 멀티모달 입력를 조화롭게 조정하는 메소드들이 확산 모델이 사용자 의도에 더 잘 부합하고, 다양한 작업에서 더 신뢰할 수 있으며, 다양한 응용 분야에 더 확장 가능하게 할 수 있음을 보여주었다.

    이 논문에서는 객체와 모션 커스터마이징에 중점을 두고 있지만, 내포하고 있는 통찰과 제안한 방법론은 멀티모달 입력을 받는 생성 모델들에 대한 더 넓은 의미를 지닐 수 있다. 또한, 커스터마이징을 멀티모달 입력 간의 관계에서 다시 바라봄으로써, 본 연구는 확산 모델 커스터마이징 분야에 새로운 관점을 제시하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Diffusion models have emerged as powerful tools for generating high-quality images and videos from diverse input conditions, including textual prompts, reference images, and structural guides. As these models gain traction in real-world applications, customization tasks have become increasingly important. Customization tasks aim to adapt a pre-trained model to user-specific subjects. Generally, customization methods take a visual input that encodes the core information of the personalized subject, and the textual input that describes novel contexts to be composed with the subject. However, most current approaches to diffusion model customization focus narrowly on individual input modalities or treat them independently, often neglecting the complex interactions and potential misalignments that arise between multiple modalities.

    This dissertation highlights critical challenges arising from such modality misalignment, particularly in the context of the subject and motion customization. We investigate how the interplay between visual and textual inputs can lead to inconsistencies, degraded generation quality, or loss of intended control. To address this, we explore a range of customization tasks that reflect these challenges in concrete and controlled settings.

    Specifically, we first examine single-image-based zero-shot subject customization, where a diffusion model is adapted to generate a personalized subject from a single visual reference without additional training. Here, we explore the relationship between the visual embedding of a personalized subject and the textual embedding of the guiding prompt. We find that conflicts between these two modalities can lead to failures in either aligning with the text or maintaining subject consistency, particularly when generating the subject in various poses.

    We then examine motion customization, where the personalized subject is extended to specific motions from a reference video. In this setting, the goal is to transfer the learned motion to a new object. To achieve this, the model is conditioned on a motion representation learned from a reference video, as well as a textual description of the new object. We analyze why these inputs are misaligned in the existing methods and examine how misaligned inputs limit the editability of the generated video.

    Finally, we propose methods to enhance the alignment between the visual and textual inputs, ensuring that the generated outputs faithfully reflect the novel context while effectively preserving the subject. Our experiments also demonstrate that harmonizing these two inputs makes diffusion models more responsive to user intent, more reliable across tasks, and more adaptable to diverse applications.

    Although our work centers on specific tasks, zero-shot subject and motion customization, the insights and methodologies developed here have broader implications for the design of multimodal generative systems. By examining the relationship between multimodal inputs in customization tasks, this research offers a new perspective to the evolving landscape of diffusion models.
    번역하기

    Diffusion models have emerged as powerful tools for generating high-quality images and videos from diverse input conditions, including textual prompts, reference images, and structural guides. As these models gain traction in real-world applications, ...

    Diffusion models have emerged as powerful tools for generating high-quality images and videos from diverse input conditions, including textual prompts, reference images, and structural guides. As these models gain traction in real-world applications, customization tasks have become increasingly important. Customization tasks aim to adapt a pre-trained model to user-specific subjects. Generally, customization methods take a visual input that encodes the core information of the personalized subject, and the textual input that describes novel contexts to be composed with the subject. However, most current approaches to diffusion model customization focus narrowly on individual input modalities or treat them independently, often neglecting the complex interactions and potential misalignments that arise between multiple modalities.

    This dissertation highlights critical challenges arising from such modality misalignment, particularly in the context of the subject and motion customization. We investigate how the interplay between visual and textual inputs can lead to inconsistencies, degraded generation quality, or loss of intended control. To address this, we explore a range of customization tasks that reflect these challenges in concrete and controlled settings.

    Specifically, we first examine single-image-based zero-shot subject customization, where a diffusion model is adapted to generate a personalized subject from a single visual reference without additional training. Here, we explore the relationship between the visual embedding of a personalized subject and the textual embedding of the guiding prompt. We find that conflicts between these two modalities can lead to failures in either aligning with the text or maintaining subject consistency, particularly when generating the subject in various poses.

    We then examine motion customization, where the personalized subject is extended to specific motions from a reference video. In this setting, the goal is to transfer the learned motion to a new object. To achieve this, the model is conditioned on a motion representation learned from a reference video, as well as a textual description of the new object. We analyze why these inputs are misaligned in the existing methods and examine how misaligned inputs limit the editability of the generated video.

    Finally, we propose methods to enhance the alignment between the visual and textual inputs, ensuring that the generated outputs faithfully reflect the novel context while effectively preserving the subject. Our experiments also demonstrate that harmonizing these two inputs makes diffusion models more responsive to user intent, more reliable across tasks, and more adaptable to diverse applications.

    Although our work centers on specific tasks, zero-shot subject and motion customization, the insights and methodologies developed here have broader implications for the design of multimodal generative systems. By examining the relationship between multimodal inputs in customization tasks, this research offers a new perspective to the evolving landscape of diffusion models.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 1
    • 1.1 Diffusion models 1
    • 1.2 Challenges 8
    • 1.3 Contributions 11
    • 1.4 Outline 12
    • 1 Introduction 1
    • 1.1 Diffusion models 1
    • 1.2 Challenges 8
    • 1.3 Contributions 11
    • 1.4 Outline 12
    • 2 Related Works 13
    • 2.1 Text-conditioned generative models 13
    • 2.2 Subject customization models 15
    • 2.3 Video generation models 18
    • 2.4 Handling complex inputs in generative models 21
    • 3 Harmonizing Visual and Textual Embeddings 26
    • 3.1 Introduction 26
    • 3.2 Problems: pose bias and identity loss 30
    • 3.3 Methods 32
    • 3.4 Experiments 36
    • 3.5 Analysis 44
    • 3.6 Limitation 51
    • 3.7 Conclusion 51
    • 4 SAVE: Structure Agnostic Video Editing 53
    • 4.1 Introduction 53
    • 4.2 Problem: Location bias of motion-related words 58
    • 4.3 Method 60
    • 4.4 Experiments 64
    • 4.5 Analysis 72
    • 4.6 Limitation 77
    • 4.7 Conclusion 79
    • 5 Analysis and Additional Results 80
    • 5.1 Deeper look into visual and textual inputs 80
    • 5.2 Subject customization results 81
    • 5.3 Motion customization results 90
    • 6 Conclusion 95
    • 6.1 Summary 95
    • 6.2 Future works 98
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼