RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    발화자 특징과 감정 차원 표현을 활용한 wav2vec 2.0 기반 음성 감정 인식

    한글로보기

    https://www.riss.kr/link?id=T16954243

    • 저자
    • 발행사항

      서울 : 중앙대학교 대학원, 2024

    • 학위논문사항
    • 발행연도

      2024

    • 작성언어

      한국어

    • 발행국(도시)

      서울

    • 기타서명

      Using speaker features and emotion dimensional representation in Wav2vec 2.0-based modules for speech emotion recognition

    • 형태사항

      iv, 54장 : 삽화, 도표 ; 26 cm

    • 일반주기명

      중앙대학교 논문은 저작권에 의해 보호받습니다
      지도교수: 홍현기
      참고문헌수록

    • UCI식별코드

      I804:11052-000000240174

    • DOI식별코드
    • 소장기관
      • 중앙대학교 서울캠퍼스 학술정보원 소장기관정보
    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    음성 감정 인식은 컴퓨터가 사람의 반언어적인 표현을 이해하고 상호작용하기 위해 중요한 요소다. 하지만 문화, 언어, 성별과 성격과 같은 개인의 음성 특성의 다양성은 음성 감정 인식의 성능을 저하할 수 있다. 그리고 감정은 범주적으로 분류될 뿐만 아니라 연속적인 감정 공간 위치로도 표현할 수 있다. 그렇기에 본 논문은 발화자 특성과 최종 감정 분류하기 전에 감정 차원 분류의 중요성을 강조한다. 제안된 방법은 세 개의 wav2vec 2.0 모듈을 기반으로 확장된 네트워크(발화자 식별, 감정 분류, 감정 차원 네트워크)다. 이 네트워크는 Arcface 혹은 교차-엔트로피 손실함수를 사용하여 학습한다. 발화자 식별 네트워크는 입력 음성 파형을 발화자 특성으로 인코딩하기 위해 하나의 어텐션 블록을 사용한다. 감정 분류 네트워크 또한 wav2vec 2.0과 네 개의 어텐션 블록을 활용하여 같은 음성 파형을 감정 특징값으로 인코딩한다. 이 두 개의 특징값을 발화자 특성과 감정을 담은 특징값으로 융합한다. 감정 차원 기반 감정 분류 네트워크 또한 wav2vec 2.0 기반에 두 개의 어텐션 블록을 사용한다. 어텐션 블록에서 나온 벡터는 각각 정서가와 각성도를 나타내어 하나의 감정 특징값으로 융합한다. 감정 차원 기반 감정 분류 네트워크 또한 발화자 식별 네트워크와 결합하여 발화자와 감정 차원 기반 감정 특징값을 만든다. 발화자 특성과 감정 차원은 각각 음성 감정 인식의 성능을 올린다. 하지만 이 두 개의 특징을 함께 사용한다고 성능이 좋아지지는 않는다. Arcface 손실함수를 사용하여, 같은 클래스의 분산도를 줄이고 다른 클래스와의 경계가 분명해지는 것을 t-분포 확률적 임베딩 (t-SNE)을 통해 시각화하여 확인한다. 제안된 방법은 Interactive Emotional Dynamic Motion Capture (IEMOCAP) 데이터 셋에서 weighted accuracy (WA)가 72.14%, unweighted accuracy (UA)가 72.97%를 달성하며 이전 방법들보다 좋은 성능을 보인다. 이러한 실험 결과는 제안된 방법의 효과성과 음성 감정 인식을 활용한 사람과 컴퓨터 사이의 상호작용 능력을 향상할 수 있는 잠재력을 보였다.
    번역하기

    음성 감정 인식은 컴퓨터가 사람의 반언어적인 표현을 이해하고 상호작용하기 위해 중요한 요소다. 하지만 문화, 언어, 성별과 성격과 같은 개인의 음성 특성의 다양성은 음성 감정 인식의 ...

    음성 감정 인식은 컴퓨터가 사람의 반언어적인 표현을 이해하고 상호작용하기 위해 중요한 요소다. 하지만 문화, 언어, 성별과 성격과 같은 개인의 음성 특성의 다양성은 음성 감정 인식의 성능을 저하할 수 있다. 그리고 감정은 범주적으로 분류될 뿐만 아니라 연속적인 감정 공간 위치로도 표현할 수 있다. 그렇기에 본 논문은 발화자 특성과 최종 감정 분류하기 전에 감정 차원 분류의 중요성을 강조한다. 제안된 방법은 세 개의 wav2vec 2.0 모듈을 기반으로 확장된 네트워크(발화자 식별, 감정 분류, 감정 차원 네트워크)다. 이 네트워크는 Arcface 혹은 교차-엔트로피 손실함수를 사용하여 학습한다. 발화자 식별 네트워크는 입력 음성 파형을 발화자 특성으로 인코딩하기 위해 하나의 어텐션 블록을 사용한다. 감정 분류 네트워크 또한 wav2vec 2.0과 네 개의 어텐션 블록을 활용하여 같은 음성 파형을 감정 특징값으로 인코딩한다. 이 두 개의 특징값을 발화자 특성과 감정을 담은 특징값으로 융합한다. 감정 차원 기반 감정 분류 네트워크 또한 wav2vec 2.0 기반에 두 개의 어텐션 블록을 사용한다. 어텐션 블록에서 나온 벡터는 각각 정서가와 각성도를 나타내어 하나의 감정 특징값으로 융합한다. 감정 차원 기반 감정 분류 네트워크 또한 발화자 식별 네트워크와 결합하여 발화자와 감정 차원 기반 감정 특징값을 만든다. 발화자 특성과 감정 차원은 각각 음성 감정 인식의 성능을 올린다. 하지만 이 두 개의 특징을 함께 사용한다고 성능이 좋아지지는 않는다. Arcface 손실함수를 사용하여, 같은 클래스의 분산도를 줄이고 다른 클래스와의 경계가 분명해지는 것을 t-분포 확률적 임베딩 (t-SNE)을 통해 시각화하여 확인한다. 제안된 방법은 Interactive Emotional Dynamic Motion Capture (IEMOCAP) 데이터 셋에서 weighted accuracy (WA)가 72.14%, unweighted accuracy (UA)가 72.97%를 달성하며 이전 방법들보다 좋은 성능을 보인다. 이러한 실험 결과는 제안된 방법의 효과성과 음성 감정 인식을 활용한 사람과 컴퓨터 사이의 상호작용 능력을 향상할 수 있는 잠재력을 보였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In human-machine interaction, speech emotion recognition (SER) plays a crucial role as machines respond to humans using paralanguage expressions. However, variations in speaker-specific properties, such as culture, language, gender, and personality, can hinder the performance of standard representations in SER tasks. Additionally, emotions are not only classified into categories but also positioned within a continuous emotion dimensional space. This study highlights the significance of speaker-specific speech characteristics and the importance of emotion dimensional classification prior to final emotion classification, aiming to enhance the performance of SER models. The proposed approach consists of three wav2vec-based modules: a speaker-identification network, an emotion classification network, and an emotion dimensional network. These modules are trained using either Arcface or cross-entropy loss. The speaker-identification network employs a single attention block to encode an input audio waveform into a speaker-specific representation. The emotion classification network, utilizing a wav2vec 2.0-backbone and four attention blocks, encodes the same input audio waveform into an emotion representation. These representations are fused into a single vector that encapsulates both emotion and speaker-specific information. The emotion dimensional network, equipped with wav2vec 2.0 and two attention blocks, encodes valence and arousal separately, merging them later to form an emotion representation. Additionally, the emotion dimensional network, combined with speaker identification, generates speaker-dimensional specific emotion representations. Experimental results demonstrate that incorporating speaker-specific characteristics and emotion dimensional classification enhances SER performance. However, combining these two features did not exhibit significant improvement. Furthermore, incorporating an angular marginal loss, such as the Arcface loss, improves intra-class compactness and inter-class separability, as shown by t-distributed stochastic neighbor embeddings (t-SNE) plots. The proposed approach surpasses previous methods, achieving a weighted accuracy (WA) of 72.14% and an unweighted accuracy (UA) of 72.97% on the Interactive Emotional Dynamic Motion Capture (IEMOCAP) dataset. The experimental findings confirm the effectiveness of the proposed method and its potential to enhance human-machine interaction through more accurate emotion recognition in speech.
    번역하기

    In human-machine interaction, speech emotion recognition (SER) plays a crucial role as machines respond to humans using paralanguage expressions. However, variations in speaker-specific properties, such as culture, language, gender, and personality, c...

    In human-machine interaction, speech emotion recognition (SER) plays a crucial role as machines respond to humans using paralanguage expressions. However, variations in speaker-specific properties, such as culture, language, gender, and personality, can hinder the performance of standard representations in SER tasks. Additionally, emotions are not only classified into categories but also positioned within a continuous emotion dimensional space. This study highlights the significance of speaker-specific speech characteristics and the importance of emotion dimensional classification prior to final emotion classification, aiming to enhance the performance of SER models. The proposed approach consists of three wav2vec-based modules: a speaker-identification network, an emotion classification network, and an emotion dimensional network. These modules are trained using either Arcface or cross-entropy loss. The speaker-identification network employs a single attention block to encode an input audio waveform into a speaker-specific representation. The emotion classification network, utilizing a wav2vec 2.0-backbone and four attention blocks, encodes the same input audio waveform into an emotion representation. These representations are fused into a single vector that encapsulates both emotion and speaker-specific information. The emotion dimensional network, equipped with wav2vec 2.0 and two attention blocks, encodes valence and arousal separately, merging them later to form an emotion representation. Additionally, the emotion dimensional network, combined with speaker identification, generates speaker-dimensional specific emotion representations. Experimental results demonstrate that incorporating speaker-specific characteristics and emotion dimensional classification enhances SER performance. However, combining these two features did not exhibit significant improvement. Furthermore, incorporating an angular marginal loss, such as the Arcface loss, improves intra-class compactness and inter-class separability, as shown by t-distributed stochastic neighbor embeddings (t-SNE) plots. The proposed approach surpasses previous methods, achieving a weighted accuracy (WA) of 72.14% and an unweighted accuracy (UA) of 72.97% on the Interactive Emotional Dynamic Motion Capture (IEMOCAP) dataset. The experimental findings confirm the effectiveness of the proposed method and its potential to enhance human-machine interaction through more accurate emotion recognition in speech.

    더보기

    목차 (Table of Contents)

    • 1. 서론 1
    • 1.1 연구 배경 2
    • 1.2 연구 목적 및 내용 4
    • 1.3 논문 구성 5
    • 2. 관련 연구 6
    • 1. 서론 1
    • 1.1 연구 배경 2
    • 1.2 연구 목적 및 내용 4
    • 1.3 논문 구성 5
    • 2. 관련 연구 6
    • 2.1 오디오 특징 추출 7
    • 2.1.1 자기 지도 학습을 이용한 오디오 특징 추출 7
    • 2.1.2 Wav2vec 2.0 8
    • 2.2 Additive Angular Margin 손실함수 11
    • 2.3 발화자 특성과 감정 차원 기반 음성 감정 인식 13
    • 2.3.1 발화자 특성 기반 음성 감정 인식 13
    • 2.3.2 감정 차원 기반 음성 감정 인식 13
    • 3. 제안된 방법 15
    • 3.1 발화자 특성 기반 음성 감정 인식 16
    • 3.1.1 발화자 식별과 감정 분류 네트워크 16
    • 3.1.2 발화자 특성 기반 감정 분류 네트워크 18
    • 3.2 감정 차원 기반 음성 감정 인식 20
    • 3.2.1 감정 차원 범주화 20
    • 3.2.2 감정 차원 기반 감정 분류 네트워크 20
    • 3.2.3 감정 차원 구성과 손실함수에 따른 실험 22
    • 3.2.4 발화자와 감정 차원 기반 음성 감정 인식 22
    • 4. 실험 결과 24
    • 4.1 데이터 셋 25
    • 4.1.1 IEMOCAP 데이터 셋 25
    • 4.1.2 감정 차원과 러셀의 감정 차원 모형 비교 26
    • 4.1.3 VoxCeleb1 데이터 셋 27
    • 4.2 세부적인 실험 조건 29
    • 4.3 평가 지표 (Metric) 30
    • 4.4 실험 결과 31
    • 4.4.1 발화자 식별과 감정 분류 네트워크 분석 31
    • 4.4.2 발화자 특성 기반 감정 분류 네트워크 미세 조정 분석 34
    • 4.4.3 감정 차원 기반 감정 분류 네트워크 분석 37
    • 4.4.4 발화자와 감정 차원 기반 감정 분류 네트워크 분석 41
    • 4.4.5 기존 연구와 비교 42
    • 5. 결론 44
    • 참고문헌 46
    • 국문초록 51
    • ABSTRACT 53
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼