Currently, image generation and synthesis have remarkably progressed with various generative models in computer vision fields. Despite the photo-realistic result of recent generative models, intrinsic discrepancies are still observed in the frequency ...
Currently, image generation and synthesis have remarkably progressed with various generative models in computer vision fields. Despite the photo-realistic result of recent generative models, intrinsic discrepancies are still observed in the frequency domain. The spectral discrepancy appeared not only in generative adversarial networks but also in diffusion models. Due to the property of the frequency space, spectral anomaly affects the quality of the image generation. Therefore, discrepancies in the frequency domain have been regarded as an intrinsic challenge for better generative models.
In this study, we propose a framework to effectively mitigate the frequency domain disparity of generated images to improve the generative performance of both generative adversarial networks and diffusion models. This is realized by spectrum translation for the refinement of image generation (STIG) with the tactic of contrastive learning. We adopt the theoretical logic of frequency components in various generative networks. The key idea, here, is to refine the spectrum of the generated image via the concept of image-to-image translation under an unpaired setting with contrastive learning in terms of digital signal processing.
We evaluate the proposed framework across eight fake image datasets from various generative adversarial networks and diffusion models. And then, various cutting-edge methods are applied for validation to demonstrate the effectiveness of STIG. Our framework outperforms other cutting-edge methods showing significant decreases in FID and log frequency distance in the frequency domain. We further emphasize that STIG improves the quality of generated images by reducing the spectral anomaly. Additionally, validation results present that frequency-based deepfake detectors are easily vulnerable in the case where spectrums of fake images are manipulated by STIG.