RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Advancing Deep Speaker Verification through Representation and Decision Learning = 표현 학습과 판단 최적화를 통한 딥러닝 기반 화자검증 기술 고도화

    한글로보기

    https://www.riss.kr/link?id=T17314392

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Speaker verification requires robust and discriminative speaker representations. Recent advances in deep learning have significantly improved speaker verification systems, particularly when embedding extraction and scoring are jointly optimized. In this paper, I propose novel deep learning-based methods to enhance the robustness and reliability of speaker verification. First, I introduce an information preservation pooling technique that maximizes mutual information between frame-level and utterance-level features to prevent information loss during pooling. Second, I present a generalized score comparison-based learning objective that directly optimizes speaker verification performance by enforcing score separation between positive and negative pairs. Finally, I propose an evidential scoring network that quantifies prediction uncertainty in verification decisions, improving system reliability under challenging conditions. Extensive experiments demonstrate that the proposed methods consistently improve speaker verification performance and robustness across diverse datasets and model architectures.
    번역하기

    Speaker verification requires robust and discriminative speaker representations. Recent advances in deep learning have significantly improved speaker verification systems, particularly when embedding extraction and scoring are jointly optimized. In th...

    Speaker verification requires robust and discriminative speaker representations. Recent advances in deep learning have significantly improved speaker verification systems, particularly when embedding extraction and scoring are jointly optimized. In this paper, I propose novel deep learning-based methods to enhance the robustness and reliability of speaker verification. First, I introduce an information preservation pooling technique that maximizes mutual information between frame-level and utterance-level features to prevent information loss during pooling. Second, I present a generalized score comparison-based learning objective that directly optimizes speaker verification performance by enforcing score separation between positive and negative pairs. Finally, I propose an evidential scoring network that quantifies prediction uncertainty in verification decisions, improving system reliability under challenging conditions. Extensive experiments demonstrate that the proposed methods consistently improve speaker verification performance and robustness across diverse datasets and model architectures.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    현대의 화자인증 시스템은 보안, 사용자 인증, 음성 기반 인터페이스 등 다양한 분야에서 활용되며, 이에 따라 신뢰도 높은 화자 표현 기술의 중요성이 커지고 있다. 최근 딥러닝 기반의 종단 간 학습 방식은 임베딩 추출과 스코어링을 통합하여 전반적인 화자인증 성능을 크게 향상시키고 있다. 이와 같은 배경에서, 본 논문은 딥러닝 기반의 화자인증 시스템의 강인성, 정확도, 결정 신뢰도를 향상시키기 위한 세 가지 새로운 기법을 제안한다.
    우선, 프레임 수준과 문장 수준 간의 상호 정보를 최대화하는 정보 보존 풀링 기법을 통해, 일반적인 풀링 과정에서 발생할 수 있는 정보 손실 문제를 완화하고 보다 정밀한 화자 표현을 가능하게 한다.
    둘째, 검증 목적에 최적화된 일반화된 점수 비교 학습 목적(GSC loss)을 도입하여, 동일/비동일 화자 음성 간 점수 분리를 효과적으로 학습한다.
    셋째, 예측 불확실성을 정량화할 수 있는 증거 기반 스코어링 네트워크(ESN)를 제안하여 OOD(분포 바깥) 환경에서도 신뢰도 높은 판단을 가능하게 한다.
    제안된 기법들은 ECAPA-TDNN, MFA-Conformer 등 다양한 모델과 결합하여, VoxCeleb 및 CN-Celeb 데이터셋에서 수행된 다양한 실험을 통해 일관된 성능 향상과 시스템의 신뢰도 개선을 보였다. 본 연구는 기존 화자 임베딩 시스템에 유연하게 통합될 수 있으며, 향후 불확실성 인식 및 검증 중심 학습 방식 연구의 기반을 제공한다.
    번역하기

    현대의 화자인증 시스템은 보안, 사용자 인증, 음성 기반 인터페이스 등 다양한 분야에서 활용되며, 이에 따라 신뢰도 높은 화자 표현 기술의 중요성이 커지고 있다. 최근 딥러닝 기반의 종...

    현대의 화자인증 시스템은 보안, 사용자 인증, 음성 기반 인터페이스 등 다양한 분야에서 활용되며, 이에 따라 신뢰도 높은 화자 표현 기술의 중요성이 커지고 있다. 최근 딥러닝 기반의 종단 간 학습 방식은 임베딩 추출과 스코어링을 통합하여 전반적인 화자인증 성능을 크게 향상시키고 있다. 이와 같은 배경에서, 본 논문은 딥러닝 기반의 화자인증 시스템의 강인성, 정확도, 결정 신뢰도를 향상시키기 위한 세 가지 새로운 기법을 제안한다.
    우선, 프레임 수준과 문장 수준 간의 상호 정보를 최대화하는 정보 보존 풀링 기법을 통해, 일반적인 풀링 과정에서 발생할 수 있는 정보 손실 문제를 완화하고 보다 정밀한 화자 표현을 가능하게 한다.
    둘째, 검증 목적에 최적화된 일반화된 점수 비교 학습 목적(GSC loss)을 도입하여, 동일/비동일 화자 음성 간 점수 분리를 효과적으로 학습한다.
    셋째, 예측 불확실성을 정량화할 수 있는 증거 기반 스코어링 네트워크(ESN)를 제안하여 OOD(분포 바깥) 환경에서도 신뢰도 높은 판단을 가능하게 한다.
    제안된 기법들은 ECAPA-TDNN, MFA-Conformer 등 다양한 모델과 결합하여, VoxCeleb 및 CN-Celeb 데이터셋에서 수행된 다양한 실험을 통해 일관된 성능 향상과 시스템의 신뢰도 개선을 보였다. 본 연구는 기존 화자 임베딩 시스템에 유연하게 통합될 수 있으며, 향후 불확실성 인식 및 검증 중심 학습 방식 연구의 기반을 제공한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Figures v
    • List of Tables viii
    • 1 Introduction 1
    • Abstract i
    • Contents ii
    • List of Figures v
    • List of Tables viii
    • 1 Introduction 1
    • 1.1 Speaker Verification 1
    • 1.2 Main Challenges in Deep Learning-based Speaker Verification 2
    • 1.3 Outline of the thesis 3
    • 2 Information Preservation Pooling for Deep Speaker Embedding 5
    • 2.1 Introduction 5
    • 2.2 Related Works 8
    • 2.2.1 X-vector Baseline System 8
    • 2.2.2 Mutual Information Estimation and Maximization 9
    • 2.3 Proposed Method 12
    • 2.3.1 Global Mutual Information Maximization 13
    • 2.3.2 Local Mutual Information Maximization 14
    • 2.3.3 Information Preservation Pooling 15
    • 2.4 Experiments 15
    • 2.4.1 Experimental Settings 15
    • 2.4.2 Results 18
    • 2.5 Conclusion 21
    • 3 Generalized Score Comparison-based Objective for Deep SpeakerEmbedding 23
    • 3.1 Introduction 23
    • 3.2 Review on Learning Objectives for Deep Speaker Embedding 28
    • 3.2.1 General Deep Speaker Embedding Framework 28
    • 3.2.2 Loss Functions for Speaker Classification 29
    • 3.2.3 Loss Functions for Inter-speaker Separability 30
    • 3.3 Proposed Method 32
    • 3.3.1 Motivation 32
    • 3.3.2 Score Comparison-based Loss 33
    • 3.3.3 Generalized Score Comparison Loss 34
    • 3.4 Experiments 39
    • 3.4.1 Datasets 39
    • 3.4.2 Experimental Settings 41
    • 3.4.3 Results Based on Training Losses 44
    • 3.4.4 Computational Complexity and GPU Memory Usage 46
    • 3.4.5 Effects of Batch Configuration and Hard Score Mining for GSC loss 47
    • 3.4.6 Ablation Study on Score Components in GSC Loss 51
    • 3.4.7 Results on State-of-the-Art Speaker Embedding Networks 52
    • 3.4.8 Results on Out-domain Datasets 55
    • 3.4.9 Multilingual Performance Analysis 57
    • 3.5 Conclusion 57
    • 4 Uncertainty-aware End-to-end Speaker Verification via Evidential Deep Learning 59
    • 4.1 Introduction 59
    • 4.2 Proposed Method 62
    • 4.2.1 Motivation 62
    • 4.2.2 Front-end Encoder Network and Speaker Classifier 63
    • 4.2.3 Evidential Scoring Network 64
    • 4.2.4 Training Loss Functions 65
    • 4.3 Experiments 67
    • 4.3.1 Datasets 67
    • 4.3.2 Network Settings 67
    • 4.3.3 Baseline Back-End Scoring Methods 68
    • 4.3.4 Speaker Verification Results 69
    • 4.3.5 Predictive Uncertainty 71
    • 4.4 Conclusion 72
    • 5 Conclusions 73
    • Bibliography 75
    • 요 약 87
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼