RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Single-Speaker Deepfake Voice Detection Using Deep Learning : A Case Study on Uzbek TTS = 딥러닝 기반 단일 화자 딥페이크 음성 탐지: 우즈벡어 TTS 사례 연구

    한글로보기

    https://www.riss.kr/link?id=T17389362

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The rapid development of deep learning and speech synthesis technologies has significantly advanced human-computer interaction, particularly through Text-to-Speech (TTS) and voice cloning systems. While these innovations have facilitated accessibility, entertainment, and automation, they have also introduced new security vulnerabilities. One major concern is the rise of deepfake voices—synthetic or cloned speech generated by neural networks—which can be exploited for identity theft, misinformation, and voice phishing attacks. This threat has become particularly relevant in emerging digital ecosystems such as Uzbekistan, where language-specific tools for deepfake voice detection remain largely unavailable. This research presents a deep learning–based framework for detecting synthetic and cloned voices in the Uzbek language, with a focus on single-speaker scenarios under limited data conditions. A custom dataset was developed, comprising approximately 5 hours of real recorded speech from a female speaker, 5 hours of TTS-generated audio, and 5 hours of cloned voice samples created using systems such as ElevenLabs, Google TTS, and OpenAI Voice Cloner. To enhance data diversity, augmentation techniques including noise addition, reverberation, and pitch shifting were applied. The proposed model combines Wav2Vec2- based contextual feature extraction with a lightweight convolutional neural network (CNN) classifier for efficient and robust binary discrimination between real and fake speech. Experimental evaluations demonstrated an accuracy of 88.86% and a ROC-AUC of 0.94, indicating the model’s strong capacity to generalize despite the small dataset and single- speaker constraint. Comparative analysis with baseline models (MFCC+SVM and Mel- spectrogram+CNN) showed that the Wav2Vec2-CNN framework achieved superior results while maintaining computational efficiency.
    번역하기

    The rapid development of deep learning and speech synthesis technologies has significantly advanced human-computer interaction, particularly through Text-to-Speech (TTS) and voice cloning systems. While these innovations have facilitated accessibility...

    The rapid development of deep learning and speech synthesis technologies has significantly advanced human-computer interaction, particularly through Text-to-Speech (TTS) and voice cloning systems. While these innovations have facilitated accessibility, entertainment, and automation, they have also introduced new security vulnerabilities. One major concern is the rise of deepfake voices—synthetic or cloned speech generated by neural networks—which can be exploited for identity theft, misinformation, and voice phishing attacks. This threat has become particularly relevant in emerging digital ecosystems such as Uzbekistan, where language-specific tools for deepfake voice detection remain largely unavailable. This research presents a deep learning–based framework for detecting synthetic and cloned voices in the Uzbek language, with a focus on single-speaker scenarios under limited data conditions. A custom dataset was developed, comprising approximately 5 hours of real recorded speech from a female speaker, 5 hours of TTS-generated audio, and 5 hours of cloned voice samples created using systems such as ElevenLabs, Google TTS, and OpenAI Voice Cloner. To enhance data diversity, augmentation techniques including noise addition, reverberation, and pitch shifting were applied. The proposed model combines Wav2Vec2- based contextual feature extraction with a lightweight convolutional neural network (CNN) classifier for efficient and robust binary discrimination between real and fake speech. Experimental evaluations demonstrated an accuracy of 88.86% and a ROC-AUC of 0.94, indicating the model’s strong capacity to generalize despite the small dataset and single- speaker constraint. Comparative analysis with baseline models (MFCC+SVM and Mel- spectrogram+CNN) showed that the Wav2Vec2-CNN framework achieved superior results while maintaining computational efficiency.

    더보기

    목차 (Table of Contents)

    • List of Figures..............................................................................................................................ⅱ
    • List of tables................................................................................................................................ⅲ
    • Abstract........................................................................................................................................ⅳ
    • 1. Introduction 1
    • 1.1 Scope of the Study 3
    • List of Figures..............................................................................................................................ⅱ
    • List of tables................................................................................................................................ⅲ
    • Abstract........................................................................................................................................ⅳ
    • 1. Introduction 1
    • 1.1 Scope of the Study 3
    • 1.2 Importance of Deep learning 3
    • 1.3 Research objectives 4
    • 2. Background and Related Works 5
    • 2.1 Deep Learning for Speech Processing 5
    • 2.2 Deepfake Voice Synthesis 5
    • 2.3 Fake Voice Detection and Anti-Spoofing 6
    • 2.4 Low-Resource Language Challenges 9
    • 2.5 Related Works Summary 10
    • 3. Methodology 11`
    • 3.1 Overview of the Methodology 11
    • 3.2 Data Collection and Sources 11
    • 3.2.1 Real Speech Collecton 12
    • 3.2.2 Text-to-Speech (TTS) Audio Collection 12
    • 3.2.3 Voice-Cloned Audio Collection 13
    • 3.3 Data Preprocessing 13
    • 3.3.1 Noise Reduction 14
    • 3.3.2 Resampling and Format Standardization 14
    • 3.3.3 Data Augmentation 14
    • 3.4 Feature Extraction 15
    • 3.4.1 Embedding Extraction Process 16
    • 3.5 Model Architecture 16
    • 3.5.1 Wav2Vec2 Encoder 17
    • 3.5.2 Convolutional Neural Network Classifier 18
    • 3.5.3 Fully Connected Layers 18
    • 3.5.4 Architectural Advantages 19
    • 3.6 Training Configuration and Environment 20
    • 3.6.1 Training Parameters 21
    • 4. Results and Discussion 21
    • 4.1 Training and Validation Performance 22
    • 4.1.1 Accuracy Curve 22
    • 4.1.2 Loss Curve 23
    • 4.1.3 Training Progress Visualization 24
    • 4.2 Evaluation Metrics 26
    • 4.3 Confusion Matrix and ROC Curve Analysis 29
    • 4.4 Discussion 31
    • 5. Conclusion 32
    • References 34
    • 국문초록 38
    • ACKNOWLEDGEMENT 39
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼