Text-conditioned visual content generation has enabled remarkable progress in synthesizing images and videos guided by natural language. However, describing a specific and unique concept solely through text remains inherently ambiguous—text alone of...
Text-conditioned visual content generation has enabled remarkable progress in synthesizing images and videos guided by natural language. However, describing a specific and unique concept solely through text remains inherently ambiguous—text alone often fails to capture fine-grained appearance details or personalized characteristics desired by the user. To overcome this limitation, reference-conditioned generation has emerged, leveraging visual cues from provided reference data to guide the synthesis process. This dissertation explores single-reference-conditioned visual content generation through three complementary adaptation perspectives—feature-level, input-level, and loss-level—that collectively enhance fidelity, temporal coherence, and controllability in text-conditioned generation.
As a feature-level integration approach, we introduce Edit-A-Video, which generates edited videos by extending text-to-image diffusion models to the video domain through selective temporal fine-tuning and structure-guided inversion. Beyond these components, the framework performs feature blending that explicitly integrates motionaware information across frames. This design enables temporally smooth and motion preserving video editing from a single reference sequence, faithfully reflecting both spatial structure and appearance dynamics.
As an input-level integration approach, we propose Diptych Prompting, which reformulates zero-shot subject-driven generation as an inpainting-based diptych composition problem. Building upon recent high-capacity text-to-image diffusion models and inpainting architectures, our framework directly embeds the reference image into the inpainting input, enabling seamless subject integration and context-aware synthesis without any task-specific retraining, thereby supporting flexible identity transfer across diverse visual domains.
As a loss-level integration approach, we present Subject Fidelity Optimization (SFO), which enhances identity preservation through negative-guided comparative learning. This objective encourages the model to explicitly distinguish between faithful and degraded generations, leading to improved fine-grained subject consistency and visual robustness under diverse conditions.
Collectively, these contributions advance the scalability and reliability of reference-conditioned generation frameworks, bridging the gap between descriptive language and personalized visual content. The proposed methodologies demonstrate the effective use of text-conditioned generative frameworks to achieve controllable, identity-preserving, and context-aware generation across both image and video domains. By establishing diverse perspectives on feature-, input-, and loss-level adaptation, this dissertation lays the groundwork for more generalizable and expressive single-reference-conditioned generation, fostering richer and more accessible forms of human–AI creative collaboration. Beyond their immediate technical improvements, these frameworks provide a conceptual foundation for adaptive and human-controllable generative systems, pointing toward future directions in multimodal, interactive visual content generation.