RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    통계적 음향 특징 기반 로지스틱 회귀를 이용한 합성 음성 탐지 연구

    한글로보기

    https://www.riss.kr/link?id=T17385774

    • 저자
    • 발행사항

      구미 : 국립금오공과대학교 산업대학원, 2026

    • 학위논문사항
    • 발행연도

      2026

    • 작성언어

      한국어

    • 발행국(도시)

      경상북도

    • 형태사항

      ; 26 cm

    • 일반주기명

      지도교수: 이해연

    • UCI식별코드

      I804:47006-000000017814

    • 소장기관
      • 국립금오공과대학교 도서관 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 인공지능 음성 합성 기술(TTS, Voice Conversion)의 급격한 발전으로 실제 음성과 구분이 어려운 합성 음성이 생성되고 있으며, 이는 음성 피싱(voice phishing), 위조 음성 인증(biometric spoofing) 등 다양한 보안 위협을 초래하고 있다. 이에 따라 실제 음성과 인공지능(AI) 기반 합성 음성을 구별하기 위한 음성 위변조 탐지(Voice Spoof Detection) 기술의 필요성이 증가하고 있다.
    본 연구에서는 다양한 통계적 음향 특성을 활용하여 경량형 로지스틱 회귀(Logistic Regression) 기반 음성 위변조 탐지 모델을 설계하였다. 제안된 방법은 RMS, ZCR, Spectral Flatness, Log-Mel Spectrogram, MFCC, F0 등 총 212차원의 음향 특징을 추출하여 학습에 사용하였으며, 데이터 종속성의 영향을 최소화하기 위해 GroupKFold 기반 5-Fold 교차검증 실험을 수행하였다.
    실험 결과, Accuracy, Precision, Recall, F1-score 등의 성능 지표는 Fold 간 변동폭이 최대 0.03 이하이며, 평균 ROC-AUC는 약 0.99 수준으로 확인되었다. 이는 제안된 모델이 특정 데이터에 편향되지 않고 안정적인 일반화 성능을 확보했음을 의미한다. 또한 단순한 선형 모델 구조임에도 빠른 연산 속도를 제공하여 실시간 탐지 환경이나 자원 제약이 존재하는 모바일·엣지 디바이스 환경에 적용 가능성이 높다.
    본 연구의 의의는 복잡한 심층 신경망 구조 없이도 설명 가능성(Explain ability)을 갖춘 고성능 음성 위변조 탐지가 가능함을 실증적으로 제시한 데 있으며, 향후 다양한 합성 음성 공격 유형과 대규모 데이터 환경에서의 확장 연구로 발전할 수 있을 것으로 기대한다.

    핵심어(Keywords): 합성 음성 탐지, 음성 위변조, 통계적 음향 특징, 로지스틱 회귀, GroupKFold, ROC-AUC
    번역하기

    최근 인공지능 음성 합성 기술(TTS, Voice Conversion)의 급격한 발전으로 실제 음성과 구분이 어려운 합성 음성이 생성되고 있으며, 이는 음성 피싱(voice phishing), 위조 음성 인증(biometric spoofing) 등 ...

    최근 인공지능 음성 합성 기술(TTS, Voice Conversion)의 급격한 발전으로 실제 음성과 구분이 어려운 합성 음성이 생성되고 있으며, 이는 음성 피싱(voice phishing), 위조 음성 인증(biometric spoofing) 등 다양한 보안 위협을 초래하고 있다. 이에 따라 실제 음성과 인공지능(AI) 기반 합성 음성을 구별하기 위한 음성 위변조 탐지(Voice Spoof Detection) 기술의 필요성이 증가하고 있다.
    본 연구에서는 다양한 통계적 음향 특성을 활용하여 경량형 로지스틱 회귀(Logistic Regression) 기반 음성 위변조 탐지 모델을 설계하였다. 제안된 방법은 RMS, ZCR, Spectral Flatness, Log-Mel Spectrogram, MFCC, F0 등 총 212차원의 음향 특징을 추출하여 학습에 사용하였으며, 데이터 종속성의 영향을 최소화하기 위해 GroupKFold 기반 5-Fold 교차검증 실험을 수행하였다.
    실험 결과, Accuracy, Precision, Recall, F1-score 등의 성능 지표는 Fold 간 변동폭이 최대 0.03 이하이며, 평균 ROC-AUC는 약 0.99 수준으로 확인되었다. 이는 제안된 모델이 특정 데이터에 편향되지 않고 안정적인 일반화 성능을 확보했음을 의미한다. 또한 단순한 선형 모델 구조임에도 빠른 연산 속도를 제공하여 실시간 탐지 환경이나 자원 제약이 존재하는 모바일·엣지 디바이스 환경에 적용 가능성이 높다.
    본 연구의 의의는 복잡한 심층 신경망 구조 없이도 설명 가능성(Explain ability)을 갖춘 고성능 음성 위변조 탐지가 가능함을 실증적으로 제시한 데 있으며, 향후 다양한 합성 음성 공격 유형과 대규모 데이터 환경에서의 확장 연구로 발전할 수 있을 것으로 기대한다.

    핵심어(Keywords): 합성 음성 탐지, 음성 위변조, 통계적 음향 특징, 로지스틱 회귀, GroupKFold, ROC-AUC

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the rapid advancement of artificial intelligence-based speech synthesis technologies, such as Text-to-Speech (TTS) and Voice Conversion (VC), synthetic speech has become increasingly similar to natural human speech, making it difficult to distinguish between real and artificially generated audio. This phenomenon has led to serious security threats, including voice phishing and biometric spoofing, thereby increasing the need for reliable voice spoof detection technologies capable of differentiating synthetic speech from genuine speech.
    This paper presents a lightweight voice spoofing detection method that leverages statistical acoustic feature information. To minimize computational complexity, the proposed approach utilizes a Logistic Regression algorithm based on 212 feature parameters reflecting the physical characteristics of speech. During the model validation phase, a GroupKFold-based 5-fold cross-validation was implemented to ensure data independence and mitigate speaker-dependent bias.
    Experimental results revealed minimal performance variance across all evaluation metrics, alongside a near-optimal ROC-AUC, confirming the model's exceptional generalization capabilities. Due to its streamlined linear structure and significantly low computational overhead, the proposed model offers the distinct advantage of enabling real-time operation on resource-constrained mobile environments.
    The key contribution of this research lies in demonstrating that explainable and high-performance voice spoof detection is achievable without relying on deep neural network architectures. Future work will focus on expanding the model to handle a wider range of synthetic speech generation methods and larger-scale real-world datasets.

    Keywords : Voice Spoof Detection, Synthetic Speech Detection, Statistical Acoustic Features, Logistic Regression, GroupKFold, ROC-AUC
    번역하기

    With the rapid advancement of artificial intelligence-based speech synthesis technologies, such as Text-to-Speech (TTS) and Voice Conversion (VC), synthetic speech has become increasingly similar to natural human speech, making it difficult to disting...

    With the rapid advancement of artificial intelligence-based speech synthesis technologies, such as Text-to-Speech (TTS) and Voice Conversion (VC), synthetic speech has become increasingly similar to natural human speech, making it difficult to distinguish between real and artificially generated audio. This phenomenon has led to serious security threats, including voice phishing and biometric spoofing, thereby increasing the need for reliable voice spoof detection technologies capable of differentiating synthetic speech from genuine speech.
    This paper presents a lightweight voice spoofing detection method that leverages statistical acoustic feature information. To minimize computational complexity, the proposed approach utilizes a Logistic Regression algorithm based on 212 feature parameters reflecting the physical characteristics of speech. During the model validation phase, a GroupKFold-based 5-fold cross-validation was implemented to ensure data independence and mitigate speaker-dependent bias.
    Experimental results revealed minimal performance variance across all evaluation metrics, alongside a near-optimal ROC-AUC, confirming the model's exceptional generalization capabilities. Due to its streamlined linear structure and significantly low computational overhead, the proposed model offers the distinct advantage of enabling real-time operation on resource-constrained mobile environments.
    The key contribution of this research lies in demonstrating that explainable and high-performance voice spoof detection is achievable without relying on deep neural network architectures. Future work will focus on expanding the model to handle a wider range of synthetic speech generation methods and larger-scale real-world datasets.

    Keywords : Voice Spoof Detection, Synthetic Speech Detection, Statistical Acoustic Features, Logistic Regression, GroupKFold, ROC-AUC

    더보기

    목차 (Table of Contents)

    • 제 1 장 서 론 1
    • 1.1 연구배경 및 문제제기 1
    • 1.2 연구 필용성 1
    • 1.3 연구 목적과 기여 2
    • 제 1 장 서 론 1
    • 1.1 연구배경 및 문제제기 1
    • 1.2 연구 필용성 1
    • 1.3 연구 목적과 기여 2
    • 제 2 장 관련 연구 및 이론적 배경 4
    • 2.1 합성 음성의 신호적 특성과 통계적 차이 4
    • 2.2 MFCC·CQCC 기반 전통적 합성 음성 탐지 연구 5
    • 2.3 딥러닝 기반 합성 음성 탐지 연구 7
    • 2.4 본 연구의 차별점 7
    • 제 3 장 제안 방법 9
    • 3.1 설계 개요 9
    • 3.2 음성 신호 전처리 10
    • 3.3 음향 특징 추출 10
    • 3.3.1 RMS 11
    • 3.3.2 ZCR 11
    • 3.3.3 Spectral Centroid/Bandwidth/Rolloff 12
    • 3.3.4 Spectral Flatness 13
    • 3.3.5 Log-Mel Spectrogram 13
    • 3.3.6 MFCC 14
    • 3.3.7 F0 및 F0 변동성 14
    • 3.3.8 최종 특성 벡터 구성 15
    • 3.4 특징 정규화 15
    • 3.5 분류기 설계 : 로지스틱 회귀 16
    • 제 4 장 실험 및 결과 18
    • 4.1 실험 환경 18
    • 4.2 데이터셋 구성 18
    • 4.3 특징 추출 및 모델 학습 19
    • 4.4 성능 평가 지표 20
    • 4.5 실험 결과 20
    • 4.6 ROC 곡선 분석 21
    • 4.7 혼동행렬 분석 22
    • 4.8 특징 기여도 분석 24
    • 4.9 기존 연구와의 비교 25
    • 4.10 결과 요약 26
    • 제 5 장 결론 및 향후 연구 27
    • 5.1 연구 요약 27
    • 5.2 연구의 의의 27
    • 5.3 향후 연구 방향 28
    • 5.4 결론 29
    • [참고 문헌] 30
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼