RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Efficient Neural Text-to-Speech via Semantic Representations = 의미 기반 표현을 활용한 효율적인 음성 합성

    한글로보기

    https://www.riss.kr/link?id=T17314673

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 몇 년간, 신경망 기반 음성합성(Text-to-Speech, TTS) 시스템의 도입은 합성 음성의 자연스러움과 표현력을 크게 향상시켰다. 그럼에도 불구하고, 현대 TTS 모델은 여전히 대규모로 레이블링된 음성 데이터셋과 상당한 수준의 연산 자원을 필요로 하며, 이는 TTS의 특수성에 맞춘 보다 효율적인 모델링 접근의 필요성을 시사한다. 본 학위논문에서는 의미 정보를 반영한 최신 음성 표현 기술을 활용하여, TTS 시스템의 학습 과정과 아키텍처 설계를 개선하기 위한 효율적인 모델링 및 학습 방법을 제안한다.

    첫째, 레이블링되지 않은 비지도 음성 데이터를 TTS 학습 과정에 통합하는 전이 학습 프레임워크를 제안한다. 제안하는 모델은 비지도 음성에서 추출한 의사 텍스트 정보를 기반으로 음향 특징을 생성하도록 사전학습되고, 이후 소량의 레이블링된 데이터를 사용하여 정확한 언어적 매핑을 학습하기 위해 미세조정된다. 이 프레임워크는 대규모 레이블링 데이터에 대한 의존도를 줄이면서도 높은 합성 품질을 유지한다.

    둘째, 텍스트와 음성 사이에 의미 토큰을 중간 표현으로 도입하는 2단계 모델링 프레임워크의 두 가지 형태를 제안한다. 이 의미 표현은 언어 모델링과 음향 모델링을 분리하여 정렬 학습을 단순화하고, 아키텍처 설계의 유연성을 높인다. 첫 번째 형태에서는 단조 정렬 제약 하에 안정성과 계산 효율성을 동시에 확보할 수 있는 트랜스듀서 기반의 구조를 활용한다. 두 번째 형태에서는 SegINR이라 명명한 새로운 구조를 제안한다. 이는 내재적 신경 표현(Implicit Neural Representation, INR)을 기반으로 하며, 구간 단위의 연속 함수로 음성을 모델링함으로써 명시적인 발음 길이 예측 없이도 출력 길이 확장을 가능하게 한다. SegINR은 가변 길이 음성 모델링에 있어 효율적이고 유연한 대안을 제공한다.
    번역하기

    최근 몇 년간, 신경망 기반 음성합성(Text-to-Speech, TTS) 시스템의 도입은 합성 음성의 자연스러움과 표현력을 크게 향상시켰다. 그럼에도 불구하고, 현대 TTS 모델은 여전히 대규모로 레이블링...

    최근 몇 년간, 신경망 기반 음성합성(Text-to-Speech, TTS) 시스템의 도입은 합성 음성의 자연스러움과 표현력을 크게 향상시켰다. 그럼에도 불구하고, 현대 TTS 모델은 여전히 대규모로 레이블링된 음성 데이터셋과 상당한 수준의 연산 자원을 필요로 하며, 이는 TTS의 특수성에 맞춘 보다 효율적인 모델링 접근의 필요성을 시사한다. 본 학위논문에서는 의미 정보를 반영한 최신 음성 표현 기술을 활용하여, TTS 시스템의 학습 과정과 아키텍처 설계를 개선하기 위한 효율적인 모델링 및 학습 방법을 제안한다.

    첫째, 레이블링되지 않은 비지도 음성 데이터를 TTS 학습 과정에 통합하는 전이 학습 프레임워크를 제안한다. 제안하는 모델은 비지도 음성에서 추출한 의사 텍스트 정보를 기반으로 음향 특징을 생성하도록 사전학습되고, 이후 소량의 레이블링된 데이터를 사용하여 정확한 언어적 매핑을 학습하기 위해 미세조정된다. 이 프레임워크는 대규모 레이블링 데이터에 대한 의존도를 줄이면서도 높은 합성 품질을 유지한다.

    둘째, 텍스트와 음성 사이에 의미 토큰을 중간 표현으로 도입하는 2단계 모델링 프레임워크의 두 가지 형태를 제안한다. 이 의미 표현은 언어 모델링과 음향 모델링을 분리하여 정렬 학습을 단순화하고, 아키텍처 설계의 유연성을 높인다. 첫 번째 형태에서는 단조 정렬 제약 하에 안정성과 계산 효율성을 동시에 확보할 수 있는 트랜스듀서 기반의 구조를 활용한다. 두 번째 형태에서는 SegINR이라 명명한 새로운 구조를 제안한다. 이는 내재적 신경 표현(Implicit Neural Representation, INR)을 기반으로 하며, 구간 단위의 연속 함수로 음성을 모델링함으로써 명시적인 발음 길이 예측 없이도 출력 길이 확장을 가능하게 한다. SegINR은 가변 길이 음성 모델링에 있어 효율적이고 유연한 대안을 제공한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Over recent years, the adoption of neural networks in Text-to-Speech (TTS) systems has significantly enhanced the naturalness and expressiveness of synthesized speech. Despite these advancements, modern TTS models still require substantial resources, including large-scale labeled speech corpora and considerable computational power. This underscores the need for more efficient modeling approaches tailored to the specific characteristics of TTS. In this thesis, I aim to design efficient modeling and training methods for TTS systems by leveraging recent developments in speech representations that capture semantic information. To this end, I propose three approaches that utilize semantic representations to improve both the training process and architectural design of TTS systems.

    First, I introduce a transfer learning framework that incorporates unlabeled speech corpora into the TTS training pipeline. The model is initially pretrained to generate acoustic features from pseudo-textual information extracted from unlabeled speech. It is then fine-tuned using a small amount of labeled data to learn accurate linguistic mappings. This framework significantly reduces the reliance on large labeled datasets while preserving synthesis quality.

    Second, I propose two variants of a two-stage modeling framework that introduces semantic tokens as an intermediate representation between text and speech. These tokens serve to decouple linguistic modeling from acoustic modeling, simplifying alignment learning and enabling greater architectural flexibility. In the first variant, I employ a transducer-based architecture for robust and efficient sequence-to-sequence (seq2seq) generation under monotonic alignment constraints, improving stability and reducing computational complexity. In the second variant, I present a novel architecture named SegINR, which leverages implicit neural representations (INRs) for speech synthesis. SegINR models speech as a continuous function using segment-wise implicit functions, allowing output length expansion without explicit duration modeling. This design provides a flexible and efficient approach to modeling variable-length speech, reducing the need for explicit duration modeling.
    번역하기

    Over recent years, the adoption of neural networks in Text-to-Speech (TTS) systems has significantly enhanced the naturalness and expressiveness of synthesized speech. Despite these advancements, modern TTS models still require substantial resources, ...

    Over recent years, the adoption of neural networks in Text-to-Speech (TTS) systems has significantly enhanced the naturalness and expressiveness of synthesized speech. Despite these advancements, modern TTS models still require substantial resources, including large-scale labeled speech corpora and considerable computational power. This underscores the need for more efficient modeling approaches tailored to the specific characteristics of TTS. In this thesis, I aim to design efficient modeling and training methods for TTS systems by leveraging recent developments in speech representations that capture semantic information. To this end, I propose three approaches that utilize semantic representations to improve both the training process and architectural design of TTS systems.

    First, I introduce a transfer learning framework that incorporates unlabeled speech corpora into the TTS training pipeline. The model is initially pretrained to generate acoustic features from pseudo-textual information extracted from unlabeled speech. It is then fine-tuned using a small amount of labeled data to learn accurate linguistic mappings. This framework significantly reduces the reliance on large labeled datasets while preserving synthesis quality.

    Second, I propose two variants of a two-stage modeling framework that introduces semantic tokens as an intermediate representation between text and speech. These tokens serve to decouple linguistic modeling from acoustic modeling, simplifying alignment learning and enabling greater architectural flexibility. In the first variant, I employ a transducer-based architecture for robust and efficient sequence-to-sequence (seq2seq) generation under monotonic alignment constraints, improving stability and reducing computational complexity. In the second variant, I present a novel architecture named SegINR, which leverages implicit neural representations (INRs) for speech synthesis. SegINR models speech as a continuous function using segment-wise implicit functions, allowing output length expansion without explicit duration modeling. This design provides a flexible and efficient approach to modeling variable-length speech, reducing the need for explicit duration modeling.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • List of Figures v
    • List of Tables viii
    • 1 Introduction 1
    • Abstract i
    • Contents iii
    • List of Figures v
    • List of Tables viii
    • 1 Introduction 1
    • 1.1 Neural Text-to-Speech 1
    • 1.2 Semantic Speech Representation 2
    • 1.3 Thesis Overview 3
    • 2 Transfer Learning Framework for Low-Resource Text-to-Speech Using a Large-Scale Unlabeled Speech Corpus 7
    • 2.1 Introduction 7
    • 2.2 Proposed Method 9
    • 2.2.1 Pseudo Phoneme 10
    • 2.2.2 Transfer Learning Framework for TTS 11
    • 2.3 Experiments 13
    • 2.3.1 Single-Speaker TTS 13
    • 2.3.2 zero-shot TTS 14
    • 2.4 Conclusions 17
    • 3 Utilizing Neural Transducers for Two-Stage Text-to-Speech via Semantic Token Prediction 19
    • 3.1 Introduction 19
    • 3.2 Backgrounds 22
    • 3.2.1 Self-Supervised Representations for TTS 22
    • 3.2.2 Alignment Modeling in Neural TTS 23
    • 3.2.3 Neural Transducer 25
    • 3.3 Proposed Method 29
    • 3.3.1 Token Transducer 29
    • 3.3.2 Speech Generator 32
    • 3.4 Experimental Settings 35
    • 3.4.1 Dataset 35
    • 3.4.2 Implementation Details 35
    • 3.4.3 Baselines 37
    • 3.5 Experimental Results 37
    • 3.5.1 Main Results: Overall Performance 37
    • 3.5.2 Analysis: Alignment of Token Transducer 39
    • 3.5.3 Analysis: Inference Speed 42
    • 3.5.4 Analysis: Paralinguistic Controllability 43
    • 3.5.5 Ablation: Semantic Token Configurations 44
    • 3.5.6 Ablation: Reducing Decoding Steps for Transducer 48
    • 3.5.7 Ablation: Cropped Reference Speech for Token Transducer 49
    • 3.6 Conclusion 50
    • 4 Segment-Wise Implicit Neural Representation for Sequence Alignment in Neural Text-to-Speech 53
    • 4.1 Introduction 53
    • 4.2 Backgrounds 56
    • 4.2.1 Length Regulation in TTS 56
    • 4.2.2 Implicit Neural Representation (INR) 57
    • 4.3 Proposed Method 58
    • 4.3.1 Segment-wise Implicit Neural Representation (SegINR) 58
    • 4.3.2 Application 60
    • 4.4 Experiments 63
    • 4.4.1 Experimental Setting 63
    • 4.4.2 Results: Zero-Shot TTS 65
    • 4.4.3 Ablation: Training and Inference Schemes 66
    • 4.5 Conclusion 68
    • 5 Conclusions 69
    • Bibliography 72
    • 요약 88
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼