Text-to-image foundation models have achieved remarkable progress, evolving from U-Net–based diffusion models to multimodal diffusion transformers (MM-DiT)-based rectified flow models. These advances have enabled not only open-ended text-to-image ge...
Text-to-image foundation models have achieved remarkable progress, evolving from U-Net–based diffusion models to multimodal diffusion transformers (MM-DiT)-based rectified flow models. These advances have enabled not only open-ended text-to-image generation but also real-image-driven applications such as personalization and real-image editing.
Despite rapid progress, key challenges remain. In personalization, non-subject elements in reference images often become entangled with the subject embedding, reducing fidelity and controllability. In real-image editing, most prior approaches were developed for U-Net architectures and do not transfer effectively to MM-DiT.
This dissertation addresses these challenges through two contributions.
First, we propose Selectively Informative Description (SID), which mitigates undesired embedding entanglement by employing training descriptions where the subject is identified only by its class, while non-subject elements are provided with informative descriptions.
Second, we introduce ReFlex, a real-image editing method tailored to rectified flow–based models. ReFlex identifies key MM-DiT features for editing, proposes mid-step inversion for structure-preserving feature extraction, and incorporates attention adaptation techniques to balance editability with source preservation.
Together, these contributions expand the capabilities of text-to-image models toward faithful and controllable real-image-driven image synthesis, improving personalization and editing while offering insights relevant to future large-scale systems.