With the recent proliferation of generative AI technologies, the creation of deepfake synthetic content has become increasingly simplified, resulting in continued damage to media credibility. Since deepfake synthetic content contains fake information ...
With the recent proliferation of generative AI technologies, the creation of deepfake synthetic content has become increasingly simplified, resulting in continued damage to media credibility. Since deepfake synthetic content contains fake information inconsistently across modalities, multi-modal AI-based deepfake detection that fuses heterogeneous modalities has been actively studied. Because the modality configuration varies across media, multi-modal AI-based deepfake detection has been developed primarily JunHo Yoon Supervised by Professor Chang Choi Dept. of IT Convergence Engineering Graduate School of Gachon University around the modality-agnostic transformer architecture, which processes all modalities uniformly as token sequences. The transformer analyzes token sequences globally through the attention mechanism, which leads to high computational complexity. To address this, Mamba, which compresses and accumulates token sequences to analyze them sequentially, has been studied for lightweight processing. However, Mamba is effective when the token sequence is large-scale, and thus has limitations in detecting deepfake synthetic content disseminated through short-form media such as Reels and YouTube Shorts. In addition, the artifacts of deepfake synthetic content appear as local inconsistencies between adjacent tokens, such as spatial inconsistency in the vision modality, temporal inconsistency in the audio modality, and word-order inconsistency in the language modality. As a result, the transformer, which analyzes token sequences globally, has limitations. In this thesis, we propose a convolution-based modality-agnostic deepfake detection model composed of minimum-size kernels to effectively analyze local artifacts between adjacent tokens. Evaluation of the proposed convolution-based deepfake detection model confirmed that, compared to the baseline, GFLOPs decreased by 30.21% to 51.15% and parameters decreased by 38.19% to 55.86%, while accuracy increased by 0.0917 to 0.1481 and f1 score increased by 0.0919 to 0.1871.