Text-to-image diffusion models demonstrate impressive generative capabilities but often retain inappropriate or sensitive visual concepts, thereby raising ethical and societal concerns. While existing concept erasure methods attempt to mitigate this i...
Text-to-image diffusion models demonstrate impressive generative capabilities but often retain inappropriate or sensitive visual concepts, thereby raising ethical and societal concerns. While existing concept erasure methods attempt to mitigate this issue, they frequently exhibit poor robustness to paraphrased prompts and result in degraded image quality.
To overcome these limitations, we propose a reinforcement learning-based framework for diverse concept erasure tasks. The proposed method suppresses undesired concepts at the semantic level by optimizing task specific reward functions while preserving image quality and maintaining text-image alignment. Extensive experiments show that our framework performs consistently across concept categories such as object, NSFW (nudity), artistic style, celebrity, and instance. Compared to prior approaches, our method more effectively balances concept suppression and image quality, while also exhibiting greater robustness to prompt variations. These results demonstrate the practicality, scalability, and general applicability of the proposed method, enabling safer and more reliable text-to-image generation.