Modern generative AI systems, including diffusion models and large language models (LLMs), have achieved remarkable capabilities but require extensive human supervision for task-specific alignment. This thesis presents a novel paradigm of minimal-feed...
Modern generative AI systems, including diffusion models and large language models (LLMs), have achieved remarkable capabilities but require extensive human supervision for task-specific alignment. This thesis presents a novel paradigm of minimal-feedback learning that dramatically reduces the human annotation burden while maintaining alignment quality. We make two primary contributions. First, we demonstrate that only 3 minutes of binary human feedback suffices to censor unwanted visual concepts from pre-trained diffusion models. Using reward model ensembles and guided sampling, we reduce the generation of malign images from up to 68% to below 1% across four diverse censoring tasks without model retraining. Second, we extend the minimal-feedback principle to text generation through adversarial self-play, where LLMs iteratively refine their own guard prompts by alternating between attacker and defender roles. We provide a unified theoretical framework that casts both adversarial refinement and iterative self-feedback methods (including PromptWizard and ProTeGi) as fixed-point iterations, establishing convergence guarantees for black-box prompt optimization. Our cross-modal results demonstrate that treating human time—not compute—as the scarcest resource enables practical, accessible AI alignment. This work opens new directions for efficient model steering across modalities, making responsible AI deployment feasible for resource-constrained researchers and organizations.