RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Beyond Transformer: Modality-Agnostic for Efficient Multi-modal Deepfake Detection

    한글로보기

    https://www.riss.kr/link?id=T17553613

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the recent proliferation of generative AI technologies, the creation of deepfake synthetic content has become increasingly simplified, resulting in continued damage to media credibility. Since deepfake synthetic content contains fake information inconsistently across modalities, multi-modal AI-based deepfake detection that fuses heterogeneous modalities has been actively studied. Because the modality configuration varies across media, multi-modal AI-based deepfake detection has been developed primarily JunHo Yoon Supervised by Professor Chang Choi Dept. of IT Convergence Engineering Graduate School of Gachon University around the modality-agnostic transformer architecture, which processes all modalities uniformly as token sequences. The transformer analyzes token sequences globally through the attention mechanism, which leads to high computational complexity. To address this, Mamba, which compresses and accumulates token sequences to analyze them sequentially, has been studied for lightweight processing. However, Mamba is effective when the token sequence is large-scale, and thus has limitations in detecting deepfake synthetic content disseminated through short-form media such as Reels and YouTube Shorts. In addition, the artifacts of deepfake synthetic content appear as local inconsistencies between adjacent tokens, such as spatial inconsistency in the vision modality, temporal inconsistency in the audio modality, and word-order inconsistency in the language modality. As a result, the transformer, which analyzes token sequences globally, has limitations. In this thesis, we propose a convolution-based modality-agnostic deepfake detection model composed of minimum-size kernels to effectively analyze local artifacts between adjacent tokens. Evaluation of the proposed convolution-based deepfake detection model confirmed that, compared to the baseline, GFLOPs decreased by 30.21% to 51.15% and parameters decreased by 38.19% to 55.86%, while accuracy increased by 0.0917 to 0.1481 and f1 score increased by 0.0919 to 0.1871.
    번역하기

    With the recent proliferation of generative AI technologies, the creation of deepfake synthetic content has become increasingly simplified, resulting in continued damage to media credibility. Since deepfake synthetic content contains fake information ...

    With the recent proliferation of generative AI technologies, the creation of deepfake synthetic content has become increasingly simplified, resulting in continued damage to media credibility. Since deepfake synthetic content contains fake information inconsistently across modalities, multi-modal AI-based deepfake detection that fuses heterogeneous modalities has been actively studied. Because the modality configuration varies across media, multi-modal AI-based deepfake detection has been developed primarily JunHo Yoon Supervised by Professor Chang Choi Dept. of IT Convergence Engineering Graduate School of Gachon University around the modality-agnostic transformer architecture, which processes all modalities uniformly as token sequences. The transformer analyzes token sequences globally through the attention mechanism, which leads to high computational complexity. To address this, Mamba, which compresses and accumulates token sequences to analyze them sequentially, has been studied for lightweight processing. However, Mamba is effective when the token sequence is large-scale, and thus has limitations in detecting deepfake synthetic content disseminated through short-form media such as Reels and YouTube Shorts. In addition, the artifacts of deepfake synthetic content appear as local inconsistencies between adjacent tokens, such as spatial inconsistency in the vision modality, temporal inconsistency in the audio modality, and word-order inconsistency in the language modality. As a result, the transformer, which analyzes token sequences globally, has limitations. In this thesis, we propose a convolution-based modality-agnostic deepfake detection model composed of minimum-size kernels to effectively analyze local artifacts between adjacent tokens. Evaluation of the proposed convolution-based deepfake detection model confirmed that, compared to the baseline, GFLOPs decreased by 30.21% to 51.15% and parameters decreased by 38.19% to 55.86%, while accuracy increased by 0.0917 to 0.1481 and f1 score increased by 0.0919 to 0.1871.

    더보기

    목차 (Table of Contents)

    • CHAPTER 1. Introduction 1
    • 1.1. Motivation and Challenge 1
    • 1.2. Methodology and Content 5
    • 1.3. Organization 8
    • CHAPTER 2. Related Work 11
    • CHAPTER 1. Introduction 1
    • 1.1. Motivation and Challenge 1
    • 1.2. Methodology and Content 5
    • 1.3. Organization 8
    • CHAPTER 2. Related Work 11
    • 2.1. Backbone Model 11
    • 2.2. Multi-modal AI 16
    • CHAPTER 3. Modality-agnostic Deepfake Detection 21
    • 3.1. Intra-Inter Mixer for Feature and Relation Representation 27
    • 3.2. Contextual Weighting for Adaptive Representation Enhancing 29
    • CHAPTER 4. Experiment and Result 32
    • 4.1. Dataset 32
    • 4.2. Experiment Environment 35
    • 4.3. Deepfake Detection Performance Evaluation 41
    • 4.4. Experiment Analysis 65
    • CHAPTER 5. Conclusion and Discussion 72
    • Reference 77
    • 국문 초록 83
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼