RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Multi-modal speech synthesis using face-based speaker identity with enhanced pitch consistency = 얼굴 기반 화자 정보를 활용한 멀티모달 음성 합성 및 피치 일관성 향상 기법

    한글로보기

    https://www.riss.kr/link?id=T17314989

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This thesis explores novel approaches to multi-modal speech synthesis using face-based speaker identity, with a particular focus on enhancing pitch consistency and enabling expressive voice generation from non-auditory inputs. Specifically, we investigate two major tasks: electromyography (EMG)-to-speech synthesis and lip-to-speech synthesis.

    First, we present a framework for extracting face-based speaker identity that enhances pitch consistency by estimating the average fundamental frequency (F0) of the target speaker solely from facial images. We evaluate the proposed method through face-based voice conversion tasks. Beyond conventional evaluation metrics such as speaker embedding similarity, we also propose analyses using explicit voice features—such as global pitch—and further examine voice attributes using our newly developed explainable speaker identity evaluation metric.

    Next, we propose a novel framework for multi-speaker EMG-to-speech synthesis using face-based speaker identity. In this setting, EMG signals are used to capture linguistic content, while facial images provide speaker identity. To bridge the modality gap between EMG and speech, we introduce a pitch-disentangled content embedding that improves linguistic representation and enhances content-related local pitch consistency.

    Finally, we present a prosody-consistency-enhanced lip-to-speech synthesis framework. Building on a state-of-the-art diffusion-based model, we incorporate pitch and energy supervision during training and estimate these prosodic features from silent video input at inference time. This approach significantly improves prosodic expressiveness—capturing both pitch and energy dynamics—while maintaining intelligibility.

    Together, these contributions enable pitch-consistent and expressively rich multi-modal speech synthesis driven by face-based speaker identity. Departing from conventional frameworks that focus solely on intelligibility, our methods place greater emphasis on imprinting speaker-specific vocal characteristics using only non-auditory input. By generating personalized speech from silent articulatory inputs or visual lip movements, the proposed methods offer a promising communication pathway for individuals who are unable to produce audible speech—preserving both their vocal identity and expressive intent.
    번역하기

    This thesis explores novel approaches to multi-modal speech synthesis using face-based speaker identity, with a particular focus on enhancing pitch consistency and enabling expressive voice generation from non-auditory inputs. Specifically, we investi...

    This thesis explores novel approaches to multi-modal speech synthesis using face-based speaker identity, with a particular focus on enhancing pitch consistency and enabling expressive voice generation from non-auditory inputs. Specifically, we investigate two major tasks: electromyography (EMG)-to-speech synthesis and lip-to-speech synthesis.

    First, we present a framework for extracting face-based speaker identity that enhances pitch consistency by estimating the average fundamental frequency (F0) of the target speaker solely from facial images. We evaluate the proposed method through face-based voice conversion tasks. Beyond conventional evaluation metrics such as speaker embedding similarity, we also propose analyses using explicit voice features—such as global pitch—and further examine voice attributes using our newly developed explainable speaker identity evaluation metric.

    Next, we propose a novel framework for multi-speaker EMG-to-speech synthesis using face-based speaker identity. In this setting, EMG signals are used to capture linguistic content, while facial images provide speaker identity. To bridge the modality gap between EMG and speech, we introduce a pitch-disentangled content embedding that improves linguistic representation and enhances content-related local pitch consistency.

    Finally, we present a prosody-consistency-enhanced lip-to-speech synthesis framework. Building on a state-of-the-art diffusion-based model, we incorporate pitch and energy supervision during training and estimate these prosodic features from silent video input at inference time. This approach significantly improves prosodic expressiveness—capturing both pitch and energy dynamics—while maintaining intelligibility.

    Together, these contributions enable pitch-consistent and expressively rich multi-modal speech synthesis driven by face-based speaker identity. Departing from conventional frameworks that focus solely on intelligibility, our methods place greater emphasis on imprinting speaker-specific vocal characteristics using only non-auditory input. By generating personalized speech from silent articulatory inputs or visual lip movements, the proposed methods offer a promising communication pathway for individuals who are unable to produce audible speech—preserving both their vocal identity and expressive intent.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문은 얼굴 기반 화자 정보를 활용한 다중모달 음성 합성의 새로운 접근 방식을 탐구하며, 특히 음고(pitch)의 일관성 향상과 표현력 있고 개인화된 음성 생성을 목표로 한다. 이를 위해 얼굴 이미지로부터 음성을 생성하는 기법을 제안하고, 이를 근전도(EMG) 기반 음성 합성 및 입술 영상 기반 음성 합성이라는 두 가지 대표적인 과제에 적용하였다.

    먼저, 얼굴 이미지로부터 화자의 평균 기본 주파수($F_0$)를 추정하여 전역 음고 일관성을 향상시키는 얼굴 기반 음성 변환 프레임워크를 제안한다. 본 방법은 시각적 얼굴 정보와 음성 특성 간의 내재된 상관관계를 학습함으로써, 대상 화자의 음성을 요구하지 않고도 일관된 음고 특성과 화자 고유의 음색을 반영한 음성 생성을 가능하게 한다. 또한, 기존 화자 구분 모델에 의존한 목소리 평가의 해석력 한계를 극복하기 위해 해석 가능한 화자 평가 벡터를 새롭게 제안하였으며, 이를 통해 합성 음성의 화자 분석에 효과적으로 활용할 수 있음을 보였다.

    다음으로, 얼굴 이미지를 활용한 다화자 무음성 생체신호 기반 음성 합성 프레임워크를 소개한다. 제안하는 구조에서는 EMG 신호를 통해 언어적 콘텐츠 정보를 추출하고, 얼굴 이미지를 통해 화자 정보를 부여한다. 특히, EMG와 음성 간의 모달리티 불일치를 해소하기 위해 음고 평탄화 모듈을 도입하여, 언어 표현력과 콘텐츠 기반의 지역 음고 일관성을 개선하였으며, 이를 통해 기존에 없던 다화자 무음성 음성 합성이 가능함을 입증하였다.

    마지막으로, 운율 일관성을 향상시킨 입술 영상 기반 음성 합성 프레임워크를 제안한다. 최신 확산 기반 모델을 기반으로, 학습 단계에서는 음고 및 에너지 정보를 지도 학습하고, 추론 단계에서는 무음성 영상 입력으로부터 해당 운율 정보를 추정한다. 제안된 방법은 음고와 에너지의 시간적 변화를 효과적으로 포착함으로써 운율 표현력을 크게 향상시키면서도 높은 명료도의 음성을 생성할 수 있음을 보였다.

    이러한 일련의 연구를 통해, 본 논문은 얼굴 이미지 기반 다중모달 음성 합성 분야에서 전역 음고의 일관성을 확보함과 동시에 화자 개인화 및 표현력 있는 음성 생성을 실현할 수 있음을 실증하였다. 제안된 방법은 무음성 생체신호 및 시각 정보만으로도 화자의 정체성을 보존한 음성 합성이 가능함을 보여주며, 향후 발화에 어려움을 겪는 사용자에게 자연스럽고 개인화된 음성 전달 수단으로 활용될 수 있을 것으로 기대된다.
    번역하기

    본 논문은 얼굴 기반 화자 정보를 활용한 다중모달 음성 합성의 새로운 접근 방식을 탐구하며, 특히 음고(pitch)의 일관성 향상과 표현력 있고 개인화된 음성 생성을 목표로 한다. 이를 위해 ...

    본 논문은 얼굴 기반 화자 정보를 활용한 다중모달 음성 합성의 새로운 접근 방식을 탐구하며, 특히 음고(pitch)의 일관성 향상과 표현력 있고 개인화된 음성 생성을 목표로 한다. 이를 위해 얼굴 이미지로부터 음성을 생성하는 기법을 제안하고, 이를 근전도(EMG) 기반 음성 합성 및 입술 영상 기반 음성 합성이라는 두 가지 대표적인 과제에 적용하였다.

    먼저, 얼굴 이미지로부터 화자의 평균 기본 주파수($F_0$)를 추정하여 전역 음고 일관성을 향상시키는 얼굴 기반 음성 변환 프레임워크를 제안한다. 본 방법은 시각적 얼굴 정보와 음성 특성 간의 내재된 상관관계를 학습함으로써, 대상 화자의 음성을 요구하지 않고도 일관된 음고 특성과 화자 고유의 음색을 반영한 음성 생성을 가능하게 한다. 또한, 기존 화자 구분 모델에 의존한 목소리 평가의 해석력 한계를 극복하기 위해 해석 가능한 화자 평가 벡터를 새롭게 제안하였으며, 이를 통해 합성 음성의 화자 분석에 효과적으로 활용할 수 있음을 보였다.

    다음으로, 얼굴 이미지를 활용한 다화자 무음성 생체신호 기반 음성 합성 프레임워크를 소개한다. 제안하는 구조에서는 EMG 신호를 통해 언어적 콘텐츠 정보를 추출하고, 얼굴 이미지를 통해 화자 정보를 부여한다. 특히, EMG와 음성 간의 모달리티 불일치를 해소하기 위해 음고 평탄화 모듈을 도입하여, 언어 표현력과 콘텐츠 기반의 지역 음고 일관성을 개선하였으며, 이를 통해 기존에 없던 다화자 무음성 음성 합성이 가능함을 입증하였다.

    마지막으로, 운율 일관성을 향상시킨 입술 영상 기반 음성 합성 프레임워크를 제안한다. 최신 확산 기반 모델을 기반으로, 학습 단계에서는 음고 및 에너지 정보를 지도 학습하고, 추론 단계에서는 무음성 영상 입력으로부터 해당 운율 정보를 추정한다. 제안된 방법은 음고와 에너지의 시간적 변화를 효과적으로 포착함으로써 운율 표현력을 크게 향상시키면서도 높은 명료도의 음성을 생성할 수 있음을 보였다.

    이러한 일련의 연구를 통해, 본 논문은 얼굴 이미지 기반 다중모달 음성 합성 분야에서 전역 음고의 일관성을 확보함과 동시에 화자 개인화 및 표현력 있는 음성 생성을 실현할 수 있음을 실증하였다. 제안된 방법은 무음성 생체신호 및 시각 정보만으로도 화자의 정체성을 보존한 음성 합성이 가능함을 보여주며, 향후 발화에 어려움을 겪는 사용자에게 자연스럽고 개인화된 음성 전달 수단으로 활용될 수 있을 것으로 기대된다.

    더보기

    목차 (Table of Contents)

    • Chapter 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Contribution 3
    • 1.3 Structure and organization 3
    • Chapter 1 Introduction 1
    • 1.1 Motivation 1
    • 1.2 Contribution 3
    • 1.3 Structure and organization 3
    • Chapter 2 Background 5
    • 2.1 Conventional speech synthesis 5
    • 2.1.1 Text-to-speech 5
    • 2.1.2 Voice conversion 7
    • 2.2 Multi-modal speech synthesis 10
    • 2.2.1 Face-based voice conversion 10
    • 2.2.2 Electromyography-to-speech 11
    • 2.2.3 Lip-to-speech 13
    • Chapter 3 Face-based speaker identity with enhanced pitch consistency 15
    • 3.1 Overview 15
    • 3.2 Methods 15
    • 3.2.1 HYFace 16
    • 3.2.2 Model architecture 19
    • 3.3 Experiments 20
    • 3.3.1 Dataset 20
    • 3.3.2 Comparison Systems 21
    • 3.3.3 Metrics 21
    • 3.3.4 Results 23
    • 3.3.5 Additional experiments 27
    • 3.3.6 Demo 28
    • 3.4 Extended analysis 28
    • 3.4.1 Vo-Ve attribute comparison 28
    • 3.4.2 General comparison of voice attributes between ground truth and synthesized speech 29
    • 3.4.3 Clip-wise comparison of voice attributes between ground truth and synthesized speech 33
    • 3.5 Conclusion 35
    • Chapter 4 EMG-to-speech synthesis with face-based speaker identity 36
    • 4.1 Overview 36
    • 4.2 Methods 37
    • 4.2.1 SWiS 37
    • 4.2.2 Model architecture 41
    • 4.3 Experiments 42
    • 4.3.1 Dataset 42
    • 4.3.2 Metrics 43
    • 4.3.3 Results 45
    • 4.3.4 Demo 47
    • 4.4 Discussion 47
    • 4.5 Conclusion 49
    • Chapter 5 Lip-to-speech synthesis with enhanced prosody consistency 50
    • 5.1 Overview 50
    • 5.2 Preliminaries 51
    • 5.2.1 Diffusion-based speech generation 51
    • 5.3 Methods 52
    • 5.3.1 Diffusion-based lip-to-speech network 52
    • 5.3.2 LipSody 53
    • 5.3.3 Prosody prediction 55
    • 5.3.4 Model architecture 56
    • 5.4 Experiments 57
    • 5.4.1 Datasets 57
    • 5.4.2 Comparison systems 58
    • 5.4.3 Implementation details 58
    • 5.4.4 Metrics 58
    • 5.5 Results 60
    • 5.5.1 Conventional metric evaluation 60
    • 5.5.2 Prosody consistency evaluation 61
    • 5.5.3 Performance according to prosody information 62
    • 5.5.4 Comparison between various pitch measurement schemes 63
    • 5.6 Conclusion 63
    • Chapter 6 Conclusion 65
    • 6.1 Summary 65
    • 6.2 Future Work 66
    • 6.2.1 Interpretable face-to-voice relationship modeling 66
    • 6.2.2 Character faces-driven voice synthesis 68
    • 6.2.3 Combine with text-driven voice synthesis 68
    • Appendix A Vo-Ve: an explainable voice-vector for speaker identity evaluation 70
    • A.1 Voice attribute dataset 70
    • A.2 Multi-label classification 71
    • A.2.1 Classification performance 72
    • A.3 Leveraging Vo-Ve: from evaluation to practical applications 73
    • A.3.1 Speaker embedding similarity evaluation 73
    • A.3.2 Interpretable application of Vo-Ve 76
    • Appendix B Crowdsource Evaluation 80
    • B.1 Subjective evalutions 80
    • Acknowledgements 101
    • 요약 102
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼