RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    유아동 음절특성을 고려한 Whisper 성능 분석과 에러 유형화 연구 = Whisper Performance Evaluation and Error Categorization with Consideration of Child Syllable Characteristics

    한글로보기

    https://www.riss.kr/link?id=T17405624

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Child speech recognition remains one of the major challenges in the field of automatic speech recognition (ASR). Most existing ASR systems have been trained predominantly on large-scale adult speech corpora, learning acoustic patterns, pronunciation characteristics, and linguistic contexts that are typical of adult speakers. This adult-oriented training paradigm has significantly improved the generalization performance of modern ASR models. However, when applied to child speech, these systems exhibit substantially degraded accuracy.
    This performance gap arises from fundamental physiological and developmental differences between children and adults. Children produce speech with higher pitch, variable speaking rates, unstable articulation, and immature phonetic structures. Their acoustic features—such as spectral distribution, formant patterns, syllable duration, and pronunciation variability—differ markedly from those of adults. As a result, ASR models trained solely on adult speech face severe difficulties in processing these unique characteristics, leading to a pronounced domain mismatch problem.
    Such domain mismatch is not merely an issue of insufficient training data; it represents a deeper domain gap in acoustic and linguistic properties, making the adaptation of existing models to child speech a demanding task across the AI community. Therefore, developing child-specific models, constructing appropriate datasets, extracting features tailored to child speech, and designing effective fine-tuning strategies are essential research directions.
    Advances in this area are crucial for improving AI services targeted at children across various applications, including education, healthcare, and safety, ultimately enhancing the inclusiveness and reliability of modern speech technologies.
    번역하기

    Child speech recognition remains one of the major challenges in the field of automatic speech recognition (ASR). Most existing ASR systems have been trained predominantly on large-scale adult speech corpora, learning acoustic patterns, pronunciation c...

    Child speech recognition remains one of the major challenges in the field of automatic speech recognition (ASR). Most existing ASR systems have been trained predominantly on large-scale adult speech corpora, learning acoustic patterns, pronunciation characteristics, and linguistic contexts that are typical of adult speakers. This adult-oriented training paradigm has significantly improved the generalization performance of modern ASR models. However, when applied to child speech, these systems exhibit substantially degraded accuracy.
    This performance gap arises from fundamental physiological and developmental differences between children and adults. Children produce speech with higher pitch, variable speaking rates, unstable articulation, and immature phonetic structures. Their acoustic features—such as spectral distribution, formant patterns, syllable duration, and pronunciation variability—differ markedly from those of adults. As a result, ASR models trained solely on adult speech face severe difficulties in processing these unique characteristics, leading to a pronounced domain mismatch problem.
    Such domain mismatch is not merely an issue of insufficient training data; it represents a deeper domain gap in acoustic and linguistic properties, making the adaptation of existing models to child speech a demanding task across the AI community. Therefore, developing child-specific models, constructing appropriate datasets, extracting features tailored to child speech, and designing effective fine-tuning strategies are essential research directions.
    Advances in this area are crucial for improving AI services targeted at children across various applications, including education, healthcare, and safety, ultimately enhancing the inclusiveness and reliability of modern speech technologies.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    유아동 음성인식 분야는 음성인식 연구 전반에서 여전히 해결해야 할 중요한 과제로 남아 있다. 현재 대부분의 자동 음성인식(ASR) 시스템은 방대한 성인 음성 기반 데이터를 중심으로 학습되어 있으며, 성인 화자의 발화 특성, 발음 패턴, 문맥 구조가 모델의 일반화 성능을 크게 향상시키는 데 기여해왔다. 그러나 유아동의 경우 성인과는 명확히 구분되는 발성 기관의 생리적 차이, 불안정한 발음, 높은 피치, 빠른 발화 속도 변화, 제한된 문장 구조 등 여러 요인으로 인해 기존 ASR 모델은 성인 대비 현저히 낮은 성능을 보인다.
    특히 유아동 음성은 스펙트럼 분포, 포먼트 구조, 음절 길이, 발음 변이 등 음향학적·운율적 특성에서 큰 차이를 보이기 때문에, 성인 음성을 기반으로 학습한 모델이 이를 그대로 처리하기 어렵다. 이러한 도메인 불일치는 단순한 데이터 부족의 문제가 아니라, 음성 특징 자체가 다른 도메인 갭(domain gap) 문제로 볼 수 있어, AI 분야에서도 해결하기 어려운 도전적 과제로 평가된다.
    실험 결과, 피치 기반 변조는 제한적인 성능 개선만을 보였으며, 과도한 조정 시 인식 성능 저하가 발생하였다. 반면 발화 속도를 완만하게 감소시키는 tempo 변조와 성도 길이 보정 기반의 VTLN 변조는 유아동 음성의 음향적·구조적 특성을 효과적으로 반영하여 오류율 감소에 기여하였다. 특히 VTLN 변조는 다양한 조건에서 가장 안정적인 성능 향상을 보여, 유아동 음성 인식 성능 개선을 위한 유망한 접근 방식임을 확인하였다.
    따라서 유아동 음성에 특화된 모델 학습, 적합한 데이터 구성, 음향·언어적 특성을 반영한 특징 추출, 적절한 파인튜닝 전략에 대한 연구가 필수적이며, 이는 실제 교육·헬스케어·안전 등 다양한 응용 분야에서 어린이를 위한 AI 서비스의 품질을 크게 좌우하는 핵심 요소가 된다.
    번역하기

    유아동 음성인식 분야는 음성인식 연구 전반에서 여전히 해결해야 할 중요한 과제로 남아 있다. 현재 대부분의 자동 음성인식(ASR) 시스템은 방대한 성인 음성 기반 데이터를 중심으로 학습...

    유아동 음성인식 분야는 음성인식 연구 전반에서 여전히 해결해야 할 중요한 과제로 남아 있다. 현재 대부분의 자동 음성인식(ASR) 시스템은 방대한 성인 음성 기반 데이터를 중심으로 학습되어 있으며, 성인 화자의 발화 특성, 발음 패턴, 문맥 구조가 모델의 일반화 성능을 크게 향상시키는 데 기여해왔다. 그러나 유아동의 경우 성인과는 명확히 구분되는 발성 기관의 생리적 차이, 불안정한 발음, 높은 피치, 빠른 발화 속도 변화, 제한된 문장 구조 등 여러 요인으로 인해 기존 ASR 모델은 성인 대비 현저히 낮은 성능을 보인다.
    특히 유아동 음성은 스펙트럼 분포, 포먼트 구조, 음절 길이, 발음 변이 등 음향학적·운율적 특성에서 큰 차이를 보이기 때문에, 성인 음성을 기반으로 학습한 모델이 이를 그대로 처리하기 어렵다. 이러한 도메인 불일치는 단순한 데이터 부족의 문제가 아니라, 음성 특징 자체가 다른 도메인 갭(domain gap) 문제로 볼 수 있어, AI 분야에서도 해결하기 어려운 도전적 과제로 평가된다.
    실험 결과, 피치 기반 변조는 제한적인 성능 개선만을 보였으며, 과도한 조정 시 인식 성능 저하가 발생하였다. 반면 발화 속도를 완만하게 감소시키는 tempo 변조와 성도 길이 보정 기반의 VTLN 변조는 유아동 음성의 음향적·구조적 특성을 효과적으로 반영하여 오류율 감소에 기여하였다. 특히 VTLN 변조는 다양한 조건에서 가장 안정적인 성능 향상을 보여, 유아동 음성 인식 성능 개선을 위한 유망한 접근 방식임을 확인하였다.
    따라서 유아동 음성에 특화된 모델 학습, 적합한 데이터 구성, 음향·언어적 특성을 반영한 특징 추출, 적절한 파인튜닝 전략에 대한 연구가 필수적이며, 이는 실제 교육·헬스케어·안전 등 다양한 응용 분야에서 어린이를 위한 AI 서비스의 품질을 크게 좌우하는 핵심 요소가 된다.

    더보기

    목차 (Table of Contents)

    • 국문초록ⅰ
    • 목 차ⅲ
    • 그림목차ⅴ
    • 도표목차ⅵ
    • 제1장 서론 1
    • 국문초록ⅰ
    • 목 차ⅲ
    • 그림목차ⅴ
    • 도표목차ⅵ
    • 제1장 서론 1
    • 1.1 유아동 발화의 발달적 특성과 음성 인식 과제 1
    • 1.2 연구배경2
    • 1.3 Whisper 모델 등장과 한계4
    • 제2장 유아동 및 성인 음성 데이터 분석 8
    • 2.1 유아동 및 성인 음성 데이터셋 구성 8
    • 2.2 음절 단위 음성 데이터 전처리11
    • 2.3 고빈도 음절 기반 성인 음성 데이터셋 분석15
    • 2.4 유아동 음성 데이터셋의 발화 특성 및 한계 분석 17
    • 2.5 유아동과 성인 공통 음절 분석18
    • 2.6 유아동 음성과 성인 음성의 비교 분석 19
    • 2.7 공통 음절 기반 음향 특징 분석 방법22
    • 2.8 공통 음절 기반 분석 결과 시각화23
    • 제3장 유아동 음성의 에러유형 27
    • 3.1 유아동 음성의 에러유형 종류 27
    • 3.2 유아동 발화의 음향적 특성과 MFCC 차이점 29
    • 제4장 성인 데이터 유아동 특징 적용 31
    • 4.1 유아동 음성 특성 반영 변조 전략 설계 31
    • 4.2 주파수 변조(Pitch Perturbation)32
    • 4.3 속도 변조(Tempo Perturbation) 33
    • 4.4 기본주파수 변조 + 속도 변조 35
    • 4.5 성대길이 정규화(VTLN : vocal tract length normalization) 36
    • 4.6 오류 유형 분석 기반 변조 전략 설계 근거38
    • 제5장 Whisper 실험 방법과 실험 결과 40
    • 5.1 실험 환경 및 모델 설정40
    • 5.2 Whisper Normalization 42
    • 5.3 성능 도출 기준 43
    • 5.4 Whisper 실험 결과44
    • 제6장 결론 46
    • 6.1 Conclusion46
    • 6.2 Limitations47
    • 6.3 Future Work 48
    • 참고문헌 51
    • 영문초록 54
    • 감사의 글(Acknowledgement) 56
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼