Text-to-Image 생성은 자연어 설명을 기반으로 해당 설명과 근사한 이미지를 합성한다. 그러나 대규모 Text-to-Image 모델은 특정 참조 세트 내 대상(Subject)의 외형을 모방하고, 다른 맥락(Context)에서...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A108892906
2023
Korean
KCI등재
학술저널
567-576(10쪽)
0
상세조회0
다운로드Text-to-Image 생성은 자연어 설명을 기반으로 해당 설명과 근사한 이미지를 합성한다. 그러나 대규모 Text-to-Image 모델은 특정 참조 세트 내 대상(Subject)의 외형을 모방하고, 다른 맥락(Context)에서...
Text-to-Image 생성은 자연어 설명을 기반으로 해당 설명과 근사한 이미지를 합성한다. 그러나 대규모 Text-to-Image 모델은 특정 참조 세트 내 대상(Subject)의 외형을 모방하고, 다른 맥락(Context)에서 새로운 표현을 합성하는 능력이 부족하다. 이러한 한계를 극복하고자 DreamBooth는 Text-to-Image Diffuison 모델을 사용자 맞춤형으로 개인화(Personalization)하기 위한 파인튜닝 방법을 제안한다. DreamBooth를 통해 개인화된 Stable Diffusion 또한, 간혹 얼굴 합성 능력이 부족한 문제가 발생한다. 이에 따라 본 논문에서는 신호 대 잡음비(SNR; Signal to Noise Ratio) 조정 알고리즘을 도입하고, Beta Schedule 방식 변경 방법을 제안한다. 나아가 선행 연구와 정확한 성능 비교를 위해, 동일한 Text Prompt를 사용하여 다양한 실험을 진행했다. 실험 결과, SNR 조정 알고리즘과 Beta Schedule 방식으로 Sigmoid Schedule을 사용했을 때 가장 우수한 성능을 보였으며, Stable Diffusion의 U-Net 구조만 파인튜닝해도 효과적인 개인화가 가능함을 입증했다.
다국어 초록 (Multilingual Abstract)
Text-to-image generation produces images based on natural language descriptions. However, large-scale text-to-image models often struggle to mimic the appearance of objects within a specific reference set and generate new representations in various co...
Text-to-image generation produces images based on natural language descriptions. However, large-scale text-to-image models often struggle to mimic the appearance of objects within a specific reference set and generate new representations in various contexts. To address these limitations, DreamBooth proposes a fine-tuning method for personalizing text-to-image diffusion models. Personalized stable diffusion through DreamBooth sometimes encounters issues with insufficient facial synthesis capabilities. Therefore, in this paper, we introduce a signal-to-noise ratio (SNR) adjustment algorithm and suggest changing the Beta Schedule method to the Sigmoid Schedule. Additionally, various experiments were conducted using the same text prompts for an accurate performance comparison with previous studies. The experimental results demonstrated the best performance when the SNR adjustment algorithm and Sigmoid Schedule were used together. Furthermore, it was proven that effective personalization is achievable by fine-tuning only the U-Net structure of Stable Diffusion.
참고문헌 (Reference)
1 C. Ting, "On the Importance of Noise Scheduling for Diffusion Models"
2 A. Radford, "Learning Transferable Visual Models From Natural Language Supervision" 8748-8763, 2021
3 A. Q. Nichol, "Improved Denoising Diffusion Probabilistic Models" 8162-8171, 2021
4 S. W. Park, "How to train your pre-trained GAN models" 53 (53): 27001-27026, 2023
5 R. Rombach, "High-resolution image synthesis with latent diffusion models" 10674-10685, 2022
6 J. Lee, "Google Drive"
7 N. Ruiz, "Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation" 22500-22510, 2023
8 J. Ho, "Denoising Diffusion Probabilistic Models" 33 : 6840-6851, 2020
9 J. Song, "Denoising Diffusion Implicit Models" 2021
10 J. Sohl-Dickstein, "Deep unsupervised learning using nonequilibrium thermodynamics" 2256-2265, 2015
1 C. Ting, "On the Importance of Noise Scheduling for Diffusion Models"
2 A. Radford, "Learning Transferable Visual Models From Natural Language Supervision" 8748-8763, 2021
3 A. Q. Nichol, "Improved Denoising Diffusion Probabilistic Models" 8162-8171, 2021
4 S. W. Park, "How to train your pre-trained GAN models" 53 (53): 27001-27026, 2023
5 R. Rombach, "High-resolution image synthesis with latent diffusion models" 10674-10685, 2022
6 J. Lee, "Google Drive"
7 N. Ruiz, "Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation" 22500-22510, 2023
8 J. Ho, "Denoising Diffusion Probabilistic Models" 33 : 6840-6851, 2020
9 J. Song, "Denoising Diffusion Implicit Models" 2021
10 J. Sohl-Dickstein, "Deep unsupervised learning using nonequilibrium thermodynamics" 2256-2265, 2015
11 S. Lin, "Common Diffusion Noise Schedules and Sample Steps are Flawed"
12 R. Gozalo-Brizuela, "ChatGPT is not all you need. A State of the Art Review of large Generative AI models"
유튜브 썸네일 시각적 표현 요소 강조에 따른 시청자 태도 분석: 인물, 상품, 타이포그래피 중심으로
ESG 디자인 분석을 통한 지속 가능한 패키지 디자인 기초 요건 연구 -2021년도 국내 식·음료 패키지를 중심으로-