RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Novel VAE-based Framework to Infer Complex Speaking Style from Arbitrary Speaker and Emotion State = 임의의 화자 및 감정 상태로부터 복잡한 발화 스타일 추론을 위한 새로운 VAE 기반 프레임워크

    한글로보기

    https://www.riss.kr/link?id=T16999523

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    음성 변환 분야에 있어 타겟 화자의 감정적 발화 스타일을 전달하기 위한 이전의 연구들은, 복잡한 음성 특징을 화자 정보와 언어적 정보 등의 특정한 속성만을 담은 잠재 공간으로 분리하는 방법을 선택해 왔다. 그러나 이 방법은 타겟 화자의 정체성 또는 원하는 감정 상태를 정확하게 반영하지 못하는 등 발화 스타일의 복잡한 특성을 변환된 음성에 반영하기 어려운 경우가 있다. 해당 화자 정보 및 감정 상태에 따른 발화 스타일을 변환된 음성에 제대로 반영하기 위해선, 화자 특유의 음성 특성 및 여러 감정 표현을 포괄하는 복잡한 발화 스타일이 반영된 잠재 공간을 모델링하는 것이 중요하다. 본 연구에서는 간단한 모델로도 화자 정보와 감정 상태로부터 복잡한 발화 스타일이 반영된 잠재 공간을 모델링할 수 있는 새로운 VAE 기반 프레임워크를 소개한다. 이 프레임워크는 주어진 감정 상태의 감정적 발화 스타일과 함께 화자의 고유한 음성 특징도 일관되게 전달하는 것을 목표로, 복잡한 발화 스타일을 표현해야 하는 디코더의 학습 부담을 줄이는 동시에 표현력 있고 감정적 뉘앙스가 풍부한 자연스러운 음성 생성을 가능하게 한다. Adain-VC 모델을 음성 데이터로부터 발화 스타일을 추출하는 사후 네트워크로 사용되고, 화자 정보와 감정 상태가 컨디셔닝된 Normalizing Flows 으로 구성된 사전 네트워크가 그 사후 분포를 포착하는 방향으로 함께 훈련된다. 본 프레임워크가 화자 정보와 감정 상태가 주어진 상황에서 변환된 음성에서 복잡한 말하기 스타일을 포착할 수 있는지를 검증하기 위해, 제로-샷 음성 변환 및 감정 음성 변환을 동시에 수행하는 실험을 진행하였다.
    번역하기

    음성 변환 분야에 있어 타겟 화자의 감정적 발화 스타일을 전달하기 위한 이전의 연구들은, 복잡한 음성 특징을 화자 정보와 언어적 정보 등의 특정한 속성만을 담은 잠재 공간으로 분리하...

    음성 변환 분야에 있어 타겟 화자의 감정적 발화 스타일을 전달하기 위한 이전의 연구들은, 복잡한 음성 특징을 화자 정보와 언어적 정보 등의 특정한 속성만을 담은 잠재 공간으로 분리하는 방법을 선택해 왔다. 그러나 이 방법은 타겟 화자의 정체성 또는 원하는 감정 상태를 정확하게 반영하지 못하는 등 발화 스타일의 복잡한 특성을 변환된 음성에 반영하기 어려운 경우가 있다. 해당 화자 정보 및 감정 상태에 따른 발화 스타일을 변환된 음성에 제대로 반영하기 위해선, 화자 특유의 음성 특성 및 여러 감정 표현을 포괄하는 복잡한 발화 스타일이 반영된 잠재 공간을 모델링하는 것이 중요하다. 본 연구에서는 간단한 모델로도 화자 정보와 감정 상태로부터 복잡한 발화 스타일이 반영된 잠재 공간을 모델링할 수 있는 새로운 VAE 기반 프레임워크를 소개한다. 이 프레임워크는 주어진 감정 상태의 감정적 발화 스타일과 함께 화자의 고유한 음성 특징도 일관되게 전달하는 것을 목표로, 복잡한 발화 스타일을 표현해야 하는 디코더의 학습 부담을 줄이는 동시에 표현력 있고 감정적 뉘앙스가 풍부한 자연스러운 음성 생성을 가능하게 한다. Adain-VC 모델을 음성 데이터로부터 발화 스타일을 추출하는 사후 네트워크로 사용되고, 화자 정보와 감정 상태가 컨디셔닝된 Normalizing Flows 으로 구성된 사전 네트워크가 그 사후 분포를 포착하는 방향으로 함께 훈련된다. 본 프레임워크가 화자 정보와 감정 상태가 주어진 상황에서 변환된 음성에서 복잡한 말하기 스타일을 포착할 수 있는지를 검증하기 위해, 제로-샷 음성 변환 및 감정 음성 변환을 동시에 수행하는 실험을 진행하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In the field of Voice Conversion aimed at conveying the emotional speaking style of the target speaker, many previous studies have opted to disentangle the speech features into several latent spaces containing only specific attributes such as content information or speaker identity. However, often fails to accurately reflect the identity of the target speaker or emotional states, and it struggles to adequately represent the complex characteristics of speaking styles, leading to poor quality in voice conversion. In contrast to conventional methods that separate features, it is crucial to effectively model a latent space that encompasses various speaker-specific speech characteristics and emotional expressions. This approach aims to reduce the training burden on the decoder to reflect complex nature of speaking style, while enabling the efficient and natural generation of speech that is not only more expressive but also rich in emotional nuances. This study introduces a novel VAE-based framework capable of modeling a latent space from speaker information and an emotional state, which consistently transfers any speaker identity along with its emotional speaking style of a given categorical emotion state in the converted speech, even with a simple model. The Adain-VC model is used as a posterior network capable of capturing speaking styles, and to effectively capture the distribution of speaking styles formed by this posterior network, it is employed in conjunction with a prior network composed of Normalizing Flows. To validate whether this latent space can effectively facilitate a complex emotional speaking style in converted speech, this framework was tested with two tasks - zero-shot Voice Conversion and Emotional Voice Conversion - conducted simultaneously.
    번역하기

    In the field of Voice Conversion aimed at conveying the emotional speaking style of the target speaker, many previous studies have opted to disentangle the speech features into several latent spaces containing only specific attributes such as content ...

    In the field of Voice Conversion aimed at conveying the emotional speaking style of the target speaker, many previous studies have opted to disentangle the speech features into several latent spaces containing only specific attributes such as content information or speaker identity. However, often fails to accurately reflect the identity of the target speaker or emotional states, and it struggles to adequately represent the complex characteristics of speaking styles, leading to poor quality in voice conversion. In contrast to conventional methods that separate features, it is crucial to effectively model a latent space that encompasses various speaker-specific speech characteristics and emotional expressions. This approach aims to reduce the training burden on the decoder to reflect complex nature of speaking style, while enabling the efficient and natural generation of speech that is not only more expressive but also rich in emotional nuances. This study introduces a novel VAE-based framework capable of modeling a latent space from speaker information and an emotional state, which consistently transfers any speaker identity along with its emotional speaking style of a given categorical emotion state in the converted speech, even with a simple model. The Adain-VC model is used as a posterior network capable of capturing speaking styles, and to effectively capture the distribution of speaking styles formed by this posterior network, it is employed in conjunction with a prior network composed of Normalizing Flows. To validate whether this latent space can effectively facilitate a complex emotional speaking style in converted speech, this framework was tested with two tasks - zero-shot Voice Conversion and Emotional Voice Conversion - conducted simultaneously.

    더보기

    목차 (Table of Contents)

    • 제 1장 서 론 10
    • 1.1 연구 동기 15
    • 1.2 연구 목적 16
    • 1.3 연구 전략 19
    • 제 1장 서 론 10
    • 1.1 연구 동기 15
    • 1.2 연구 목적 16
    • 1.3 연구 전략 19
    • 제 2장 관련 연구 21
    • 2.1 Variational Autoencoder: VAE 21
    • 2.2 Normalizing Flows 24
    • 2.2.1 Continuous Normalizing Flows 25
    • 2.3 Adaptative Instance Normalization 26
    • 2.4 음성 특징과 감정의 연관성 27
    • 2.5 관련 연구 29
    • 2.5.1 음성 변환 및 잠재 공간 분리 29
    • 2.5.2 음성 변환 및 VAE 30
    • 제 3장 제안 기법 32
    • 3.1 데이터 전처리 32
    • 3.1.1 화자 인코더 (ECAPA-TDNN) 33
    • 3.1.2 음성-유닛 인코더 (ContentVec) 35
    • 3.2 VAE 기반 모델 구조 제안 36
    • 3.2.1 사후 인코더 및 디코더 37
    • 3.2.2 Post-Net 39
    • 3.2.3 사전 네트워크 41
    • 3.2.4 손실 함수 42
    • 3.2.5 보코더 (HiFi-GAN) 43
    • 제 4장 실 험 44
    • 4.1 데이터셋 44
    • 4.1.1 ESD 데이터셋 45
    • 4.1.2 Emov-DB 데이터셋 45
    • 4.1.3 JL Corpus 데이터셋 46
    • 4.2 검증 방법 49
    • 4.2.1 문자 오류율 및 단어 오류율 49
    • 4.2.2 감정 인식 정확도 49
    • 4.2.3 성별 인식 정확도 & 화자 검증 정확도 50
    • 4.3 모델 상세 구조 및 실험 설정 51
    • 제 5장 결과 및 논의 53
    • 5.1 결과 53
    • 5.2 논의 57
    • 5.2.1 음소 정보 57
    • 5.2.2 감정 상태 변화 58
    • 5.2.3 화자 정보 59
    • 5.2.4 한계점 및 고찰 59
    • 제 6장 결론 61
    • 6.1 Contributions 61
    • ABSTRACT 80
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼