RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Deep Generative Models for Personalized Speech and Spoken Dialog Modeling = 개인화 음성 및 음성 대화 합성을 위한 생성 모델

    한글로보기

    https://www.riss.kr/link?id=T17314560

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Speech is one of the most fundamental means of human communication and plays an indispensable role in our daily lives. In tasks such as audiobook narration or voice-based interaction with AI systems, collecting large-scale speech data from human speakers is often impractical and costly. As a result, deep learning-based speech synthesis has emerged as a viable and scalable solution. While the core objective is to generate natural and high-quality speech, the specific characteristics required vary by application—for example, capturing emotion, mimicking a particular speaker’s timbre, or producing conversational and affective speech that resembles real-world dialog. Among these, speaker-adaptive speech synthesis, also referred to as personalized speech synthesis, has received considerable attention due to its diverse applications.

    This dissertation presents a comprehensive methodology and analysis of personalized speech synthesis across multiple tasks, including text-to-speech (TTS), voice conversion (VC), and spoken dialog modeling. We first introduce a unified framework for building personalized TTS and VC systems, which requires only 5–10 seconds of untranscribed speech by fine-tuning a multi-speaker model. We then enhance the framework’s parameter efficiency by identifying key parameters for speaker adaptation and exploring various strategies to strengthen speaker conditioning. Finally, we propose a spoken dialog model capable of generating affective speech responses in a target speaker's voice, guided by a short reference audio sample.

    To construct a unified framework for personalized TTS and VC, we propose UnitSpeech, a fine-tuning-based, speaker-adaptive speech synthesis model that enables adaptation from minimal untranscribed speech by replacing textual input with phonetic units. These units, a type of semantic token, are self-supervised speech representations known to capture the linguistic content of speech. In this framework, a unit encoder replaces the conventional text encoder to process these semantic tokens, thereby eliminating the need for transcripts during model fine-tuning. The personalized decoder, fine-tuned on unit-speech pairs, supports both speaker-adaptive TTS through integration with the text encoder and any-to-any VC via the unit encoder. UnitSpeech achieves performance comparable to or surpassing strong baselines and demonstrates notable robustness on real-world data.

    To reduce the storage required per speaker in the fine-tuning-based, speaker-adaptive speech synthesis model, we further propose VoiceTailor, a model with parameter-efficient fine-tuning strategy that identifies and adapts only key parameters, specifically the linear layers in attention modules, through low-rank adapters. In addition, to achieve strong speaker adaptation performance with a minimal number of trainable parameters, we explore various guiding strategies that enhance speaker information during speech synthesis, resulting in optimal performance. This approach reduces the fine-tuned parameter count to just 0.25% of the full model, while maintaining comparable speaker similarity and audio quality.

    To extend speaker adaptation capabilities to a speech large language model (speech LLM) for spoken dialog, we introduce the Unified Spoken Dialog Model (USDM). This end-to-end model generates coherent and prosodically natural responses in the voice of a target speaker, without relying on explicit automatic speech recognition (ASR) or TTS modules. Initially, we demonstrate that our selected speech tokens preserve both semantic content and detailed prosodic features. We then modify a well-established personalized speech decoder framework and incorporate it into our speech LLM. Within this integrated model, in conjunction with the proposed speech LLM pretraining method, our model generates expressive and personalized spoken responses.
    번역하기

    Speech is one of the most fundamental means of human communication and plays an indispensable role in our daily lives. In tasks such as audiobook narration or voice-based interaction with AI systems, collecting large-scale speech data from human speak...

    Speech is one of the most fundamental means of human communication and plays an indispensable role in our daily lives. In tasks such as audiobook narration or voice-based interaction with AI systems, collecting large-scale speech data from human speakers is often impractical and costly. As a result, deep learning-based speech synthesis has emerged as a viable and scalable solution. While the core objective is to generate natural and high-quality speech, the specific characteristics required vary by application—for example, capturing emotion, mimicking a particular speaker’s timbre, or producing conversational and affective speech that resembles real-world dialog. Among these, speaker-adaptive speech synthesis, also referred to as personalized speech synthesis, has received considerable attention due to its diverse applications.

    This dissertation presents a comprehensive methodology and analysis of personalized speech synthesis across multiple tasks, including text-to-speech (TTS), voice conversion (VC), and spoken dialog modeling. We first introduce a unified framework for building personalized TTS and VC systems, which requires only 5–10 seconds of untranscribed speech by fine-tuning a multi-speaker model. We then enhance the framework’s parameter efficiency by identifying key parameters for speaker adaptation and exploring various strategies to strengthen speaker conditioning. Finally, we propose a spoken dialog model capable of generating affective speech responses in a target speaker's voice, guided by a short reference audio sample.

    To construct a unified framework for personalized TTS and VC, we propose UnitSpeech, a fine-tuning-based, speaker-adaptive speech synthesis model that enables adaptation from minimal untranscribed speech by replacing textual input with phonetic units. These units, a type of semantic token, are self-supervised speech representations known to capture the linguistic content of speech. In this framework, a unit encoder replaces the conventional text encoder to process these semantic tokens, thereby eliminating the need for transcripts during model fine-tuning. The personalized decoder, fine-tuned on unit-speech pairs, supports both speaker-adaptive TTS through integration with the text encoder and any-to-any VC via the unit encoder. UnitSpeech achieves performance comparable to or surpassing strong baselines and demonstrates notable robustness on real-world data.

    To reduce the storage required per speaker in the fine-tuning-based, speaker-adaptive speech synthesis model, we further propose VoiceTailor, a model with parameter-efficient fine-tuning strategy that identifies and adapts only key parameters, specifically the linear layers in attention modules, through low-rank adapters. In addition, to achieve strong speaker adaptation performance with a minimal number of trainable parameters, we explore various guiding strategies that enhance speaker information during speech synthesis, resulting in optimal performance. This approach reduces the fine-tuned parameter count to just 0.25% of the full model, while maintaining comparable speaker similarity and audio quality.

    To extend speaker adaptation capabilities to a speech large language model (speech LLM) for spoken dialog, we introduce the Unified Spoken Dialog Model (USDM). This end-to-end model generates coherent and prosodically natural responses in the voice of a target speaker, without relying on explicit automatic speech recognition (ASR) or TTS modules. Initially, we demonstrate that our selected speech tokens preserve both semantic content and detailed prosodic features. We then modify a well-established personalized speech decoder framework and incorporate it into our speech LLM. Within this integrated model, in conjunction with the proposed speech LLM pretraining method, our model generates expressive and personalized spoken responses.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    음성은 사람 간 소통의 가장 근본적인 수단 중 하나로, 일상생활에서 없어서는 안 될 중요한 역할을 합니다. 오디오북 내레이션이나 AI 모델과의 음성 기반 상호작용과 같은 응용 사례에서는, 대규모 음성 데이터를 성우로부터 수집하는 것이 시간과 비용 측면에서 매우 비효율적일 수 있습니다. 이러한 한계를 극복하기 위해, 딥러닝 기반 음성 합성 기술이 실용적이며 확장 가능한 대안으로 주목받고 있습니다. 음성 합성의 핵심 목표는 자연스럽고 고품질의 음성을 생성하는 것이지만, 적용 분야나 task에 따라 요구되는 부가적인 특성이 달라질 수 있습니다. 예를 들어, 감정 표현, 특정 화자의 목소리 모방, 또는 실제 대화처럼 자연스럽고 정서적인 발화 생성을 요구하는 경우도 있습니다. 그중에서도 화자 적응형 음성 합성, 또는 개인화 음성 합성은 다양한 활용 가능성 덕분에 특히 많은 관심을 받고 있습니다.

    본 학위논문에서는 음성 합성 (Text-to-Speech, TTS), 음성 변조 (Voice Conversion, VC), 그리고 음성 대화 모델링 (Spoken Dialog Modeling)을 아우르는 개인화 음성 합성에 대한 종합적인 방법론과 분석을 제시합니다. 먼저, 다화자 음성 합성 모델을 기반으로 5~10초 분량의 소량 음성만으로 개인화된 TTS 및 VC 시스템을 구축할 수 있는 통합 프레임워크를 제안합니다. 이후, 화자 적응을 위한 핵심 파라미터를 식별하고, 화자 정보를 효과적으로 강화할 수 있는 다양한 전략을 탐색함으로써, fine-tuning 기반 개인화 방법의 파라미터 효율성을 향상시킵니다. 마지막으로, 짧은 참조 음성만을 활용하여 목표 화자의 목소리로 자연스럽고 감정적인 응답을 생성할 수 있는 음성 대화 모델을 제안합니다.

    개인화된 TTS 및 VC를 위한 통합 프레임워크를 구축하기 위해, 본 논문에서는 UnitSpeech를 제안합니다. UnitSpeech는 10초 이내의 소량 음성만으로도 구축 가능한 fine-tuning 기반 화자 적응형 음성 합성 및 변조 모델입니다. 본 모델은 기존 음성 합성 시스템에서 입력으로 사용되던 텍스트를, 음성의 언어적 정보를 내포한 self-supervised 특성인 phonetic unit으로 대체함으로써, 전사 텍스트 없이도 unit-음성 쌍만으로 모델을 fine-tuning할 수 있도록 설계되었습니다. 이를 위해 기존의 텍스트 인코더 대신 unit 인코더를 도입하여, 모델이 unit을 입력으로 처리할 수 있도록 구조를 확장하였습니다. 이후, unit 인코더와 사전 학습된 디코더를 unit-음성 쌍으로 fine-tuning하여 개인화된 디코더를 구축합니다. 이렇게 fine-tuning된 디코더는 텍스트 인코더와 결합할 경우 화자 적응형 TTS를, unit 인코더와 결합할 경우 any-to-any VC를 지원합니다. UnitSpeech는 기존 베이스라인 모델들과 동등하거나 그 이상의 성능을 달성하였으며, 실제 환경에서도 다양한 목소리에 대해 강건한 성능을 보였습니다.

    Fine-tuning 기반 화자 적응 음성 합성 모델의 파라미터 효율성을 극대화하기 위해, 본 논문에서는 VoiceTailor를 제안합니다. 우리는 디코더의 전체 파라미터를 학습하지 않고도 효과적인 화자 개인화가 가능하며, 어텐션 모듈 내 선형 계층이 화자 정보 적응에 핵심적인 역할을 한다는 것을 분석 및 발견하였습니다. 이에 따라 해당 모듈에 low-rank adapter를 삽입하고, 이 어댑터만을 선택적으로 학습함으로써 효율적인 fine-tuning을 가능하게 하였습니다. 또한, 소수의 파라미터만을 학습하는 상황에서도 높은 수준의 화자 개인화 성능을 유지할 수 있도록, 음성 생성 과정에서 화자 정보를 더욱 풍부하게 반영하는 다양한 가이딩 전략을 함께 적용하였습니다. 제안하는 접근법은 전체 모델 파라미터의 단 0.25%만을 학습함으로써, 화자 유사성과 음질 측면에서 기존 full fine-tuning과 동등한 수준의 성능을 유지하면서도, 파라미터 및 저장 공간의 효율성을 크게 향상시켰습니다.

    이러한 화자 적응 능력을 음성 LLM (Speech Large Language Model), 특히 음성 대화 모델링 (Spoken Dialog Modeling)으로 확장하고자, 본 논문에서는 Unified Spoken Dialog Model (USDM)을 제안합니다. USDM은 별도의 음성 인식 (ASR)이나 음성 합성 (TTS) 모듈 없이도, 목표 화자의 목소리로 일관되고 자연스러운 응답을 생성할 수 있는 end-to-end 음성 대화 모델입니다. 우리는 USDM에서 사용하는 음성 토큰 (speech token)이 의미 정보뿐 아니라 운율 (prosody) 정보도 포함하고 있음을 확인하였으며, 이를 활용해 화자의 감정 상태를 반영한 정서적인 응답 생성이 가능함을 보였습니다. 또한, 음성 LLM의 출력인 speech token을 원하는 화자의 목소리로 복원할 수 있는 개인화된 음성 디코더와 결합함으로써, USDM은 다중 턴 대화에서도 목표 화자의 음색을 일관되게 유지하면서 감정이 담긴 자연스러운 응답을 생성할 수 있습니다. 제안하는 모델은 다양한 평가 지표에서 기존 기법들을 능가하는 성능을 달성하였으며, 학습에 사용되지 않은 화자에 대해서도 멀티턴 대화 상황에서 우수한 화자 개인화 성능을 입증하였습니다.
    번역하기

    음성은 사람 간 소통의 가장 근본적인 수단 중 하나로, 일상생활에서 없어서는 안 될 중요한 역할을 합니다. 오디오북 내레이션이나 AI 모델과의 음성 기반 상호작용과 같은 응용 사례에서는...

    음성은 사람 간 소통의 가장 근본적인 수단 중 하나로, 일상생활에서 없어서는 안 될 중요한 역할을 합니다. 오디오북 내레이션이나 AI 모델과의 음성 기반 상호작용과 같은 응용 사례에서는, 대규모 음성 데이터를 성우로부터 수집하는 것이 시간과 비용 측면에서 매우 비효율적일 수 있습니다. 이러한 한계를 극복하기 위해, 딥러닝 기반 음성 합성 기술이 실용적이며 확장 가능한 대안으로 주목받고 있습니다. 음성 합성의 핵심 목표는 자연스럽고 고품질의 음성을 생성하는 것이지만, 적용 분야나 task에 따라 요구되는 부가적인 특성이 달라질 수 있습니다. 예를 들어, 감정 표현, 특정 화자의 목소리 모방, 또는 실제 대화처럼 자연스럽고 정서적인 발화 생성을 요구하는 경우도 있습니다. 그중에서도 화자 적응형 음성 합성, 또는 개인화 음성 합성은 다양한 활용 가능성 덕분에 특히 많은 관심을 받고 있습니다.

    본 학위논문에서는 음성 합성 (Text-to-Speech, TTS), 음성 변조 (Voice Conversion, VC), 그리고 음성 대화 모델링 (Spoken Dialog Modeling)을 아우르는 개인화 음성 합성에 대한 종합적인 방법론과 분석을 제시합니다. 먼저, 다화자 음성 합성 모델을 기반으로 5~10초 분량의 소량 음성만으로 개인화된 TTS 및 VC 시스템을 구축할 수 있는 통합 프레임워크를 제안합니다. 이후, 화자 적응을 위한 핵심 파라미터를 식별하고, 화자 정보를 효과적으로 강화할 수 있는 다양한 전략을 탐색함으로써, fine-tuning 기반 개인화 방법의 파라미터 효율성을 향상시킵니다. 마지막으로, 짧은 참조 음성만을 활용하여 목표 화자의 목소리로 자연스럽고 감정적인 응답을 생성할 수 있는 음성 대화 모델을 제안합니다.

    개인화된 TTS 및 VC를 위한 통합 프레임워크를 구축하기 위해, 본 논문에서는 UnitSpeech를 제안합니다. UnitSpeech는 10초 이내의 소량 음성만으로도 구축 가능한 fine-tuning 기반 화자 적응형 음성 합성 및 변조 모델입니다. 본 모델은 기존 음성 합성 시스템에서 입력으로 사용되던 텍스트를, 음성의 언어적 정보를 내포한 self-supervised 특성인 phonetic unit으로 대체함으로써, 전사 텍스트 없이도 unit-음성 쌍만으로 모델을 fine-tuning할 수 있도록 설계되었습니다. 이를 위해 기존의 텍스트 인코더 대신 unit 인코더를 도입하여, 모델이 unit을 입력으로 처리할 수 있도록 구조를 확장하였습니다. 이후, unit 인코더와 사전 학습된 디코더를 unit-음성 쌍으로 fine-tuning하여 개인화된 디코더를 구축합니다. 이렇게 fine-tuning된 디코더는 텍스트 인코더와 결합할 경우 화자 적응형 TTS를, unit 인코더와 결합할 경우 any-to-any VC를 지원합니다. UnitSpeech는 기존 베이스라인 모델들과 동등하거나 그 이상의 성능을 달성하였으며, 실제 환경에서도 다양한 목소리에 대해 강건한 성능을 보였습니다.

    Fine-tuning 기반 화자 적응 음성 합성 모델의 파라미터 효율성을 극대화하기 위해, 본 논문에서는 VoiceTailor를 제안합니다. 우리는 디코더의 전체 파라미터를 학습하지 않고도 효과적인 화자 개인화가 가능하며, 어텐션 모듈 내 선형 계층이 화자 정보 적응에 핵심적인 역할을 한다는 것을 분석 및 발견하였습니다. 이에 따라 해당 모듈에 low-rank adapter를 삽입하고, 이 어댑터만을 선택적으로 학습함으로써 효율적인 fine-tuning을 가능하게 하였습니다. 또한, 소수의 파라미터만을 학습하는 상황에서도 높은 수준의 화자 개인화 성능을 유지할 수 있도록, 음성 생성 과정에서 화자 정보를 더욱 풍부하게 반영하는 다양한 가이딩 전략을 함께 적용하였습니다. 제안하는 접근법은 전체 모델 파라미터의 단 0.25%만을 학습함으로써, 화자 유사성과 음질 측면에서 기존 full fine-tuning과 동등한 수준의 성능을 유지하면서도, 파라미터 및 저장 공간의 효율성을 크게 향상시켰습니다.

    이러한 화자 적응 능력을 음성 LLM (Speech Large Language Model), 특히 음성 대화 모델링 (Spoken Dialog Modeling)으로 확장하고자, 본 논문에서는 Unified Spoken Dialog Model (USDM)을 제안합니다. USDM은 별도의 음성 인식 (ASR)이나 음성 합성 (TTS) 모듈 없이도, 목표 화자의 목소리로 일관되고 자연스러운 응답을 생성할 수 있는 end-to-end 음성 대화 모델입니다. 우리는 USDM에서 사용하는 음성 토큰 (speech token)이 의미 정보뿐 아니라 운율 (prosody) 정보도 포함하고 있음을 확인하였으며, 이를 활용해 화자의 감정 상태를 반영한 정서적인 응답 생성이 가능함을 보였습니다. 또한, 음성 LLM의 출력인 speech token을 원하는 화자의 목소리로 복원할 수 있는 개인화된 음성 디코더와 결합함으로써, USDM은 다중 턴 대화에서도 목표 화자의 음색을 일관되게 유지하면서 감정이 담긴 자연스러운 응답을 생성할 수 있습니다. 제안하는 모델은 다양한 평가 지표에서 기존 기법들을 능가하는 성능을 달성하였으며, 학습에 사용되지 않은 화자에 대해서도 멀티턴 대화 상황에서 우수한 화자 개인화 성능을 입증하였습니다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • List of Tables vii
    • List of Figures ix
    • 1 Introduction 1
    • Abstract i
    • Contents iii
    • List of Tables vii
    • List of Figures ix
    • 1 Introduction 1
    • 1.1 Motivation for Personalized Speech Synthesis 1
    • 1.2 Overview of Speech and Personalized Speech Synthesis 3
    • 1.2.1 Speech Synthesis: Overall Pipelines 5
    • 1.2.2 Approaches for Personalized Speech Synthesis 7
    • 1.3 Applications of Speaker-Adaptive Speech Synthesis 10
    • 1.3.1 Text-to-Speech 10
    • 1.3.2 Voice Conversion 11
    • 1.3.3 Speech Editing 12
    • 1.3.4 Speech-to-Speech Translation 13
    • 1.3.5 Spoken Dialog Modeling 13
    • 1.4 Scope of Dissertation 14
    • 2 Background 17
    • 2.1 Generative Models 18
    • 2.1.1 Autoregressive Models 18
    • 2.1.2 Diffusion Models 21
    • 2.1.3 Conditional Flow Matching 26
    • 2.2 Representative Models for Personalized Speech Synthesis 29
    • 2.2.1 One-Shot Speaker Adaptation: Guided-TTS 2 29
    • 2.2.2 Zero-Shot Speaker Adaptation: Voicebox 34
    • 3 Speaker-Adaptive Text-to-Speech and Voice Conversion 37
    • 3.1 Introduction 37
    • 3.2 Method 39
    • 3.2.1 Diffusion-based Text-to-Speech Model 40
    • 3.2.2 Unit Encoder Training 41
    • 3.2.3 Speaker-Adaptive Speech Synthesis 42
    • 3.3 Experimental Results 44
    • 3.3.1 Experimental Setup 44
    • 3.3.2 Results 48
    • 3.4 Concluding Remark 51
    • 4 Speaker-Adaptive Speech Synthesis with Parameter Efficiency 53
    • 4.1 Introduction 53
    • 4.2 Method 56
    • 4.2.1 UnitSpeech 57
    • 4.2.2 Parameter-Efficient Speaker Adaptation 58
    • 4.2.3 Speaker Information Strengthening Strategies 60
    • 4.3 Experimental Results 61
    • 4.3.1 Experimental Setup 61
    • 4.3.2 Results 63
    • 4.4 Concluding Remark 67
    • 5 Speaker-Adaptive Spoken Dialog Model with Paralinguistic Awareness 69
    • 5.1 Introduction 69
    • 5.2 Comparison to Related Work 73
    • 5.3 Method 75
    • 5.3.1 Speech-to-Unit Encoder 75
    • 5.3.2 Unified Speech-Text Pretraining 78
    • 5.3.3 Unified Spoken Dialog Model 80
    • 5.3.4 Speaker-Adaptive Unit-to-Speech Decoder 81
    • 5.4 Additional Details for Our Method 81
    • 5.4.1 Emotional Cues in Semantic Tokens 81
    • 5.4.2 Templates for Fine-tuning 84
    • 5.4.3 Voicebox 84
    • 5.5 Experimental Results 85
    • 5.5.1 Model Comparisons 85
    • 5.5.2 Evaluation of Speaker Adaptation 90
    • 5.5.3 Ablation Studies 95
    • 5.5.4 Analysis on Input Modality 97
    • 5.6 Additional Details for Experimental Results 98
    • 5.6.1 Models, Datasets, Training Details 98
    • 5.6.2 Human Evaluation and GPT-4 Judge 99
    • 5.7 Additional Experimental Results 103
    • 5.7.1 Additional Results for DailyTalk 103
    • 5.7.2 Ablation Studies for Pretraining Scheme 104
    • 5.7.3 Per-Task Training Dynamic Analysis 106
    • 5.7.4 Additional Attention Maps 107
    • 5.8 Samples and Demo 109
    • 5.9 Concluding Remark 110
    • 6 Concluding Remark 113
    • 6.1 Summary of Dissertation 113
    • 6.2 Discussion and Outlook 114
    • 6.2.1 Challenges in One-Shot Personalized Speech Synthesis 114
    • 6.2.2 Challenges in Zero-Shot Personalized Speech Synthesis 115
    • 6.2.3 Ethical Considerations 115
    • 6.2.4 Challenges in Speech Synthesis Beyond Personalization 116
    • Abstract (In Korean) 153
    • 감사의 글 156
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼