RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    주파수 변화율을 이용한 음성과 음악의 구분 = Speech and music discrimination using spectral transition rate

    한글로보기

    https://www.riss.kr/link?id=T11733548

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Automatic speech recognition (ASR) is becoming an indispensable technology in many application areas such as telephonic system, ubiquitous system, robot system, and telematics. In real-world environments, ASR faces many kinds of sound sources and they should be discriminated to improve ASR performance. In ASR systems, speech is usually detected from the input signal by voice activity detection (VAD) scheme. Speech and music, however, are not easily discriminated by the VAD because they share similar characteristics such as periodicity and frequency. Speech and music discrimination (SMD) has gained much popularity in recent years for efficient coding and automatic retrieval of multimedia sources and automated speech recognition. Many kinds of approaches have previously been taken to the problem of SMD using spectral energy changes, harmonics, delta cepstral energy, power spectrum deviation, cepstral distance, and etc. However, those feature parameters are not so efficient to get high performance with fast output. The mean of minimum cepstral distance (MMCD) has recently showed high performance. However, the duration used in MMCD was 1 second. Considering the length of a sentence, too much long duration is not proper for the practical system. In this paper, we propose the spectral transition rate (STR) as a novel feature for SMD. We found that the spectral peaks of speech are gradually changing. However, the sound of musical instruments in general tends to keep the peak frequencies and energies unchanged for relatively long period of time compared to speech. The STR is based on this observation. The STR of speech is much higher than that of music. Comparing to other algorithms, the STR based SMD gives relatively fast output without losing its performance. The experimental result shows that the STR based SMD method outperforms a conventional method.
    번역하기

    Automatic speech recognition (ASR) is becoming an indispensable technology in many application areas such as telephonic system, ubiquitous system, robot system, and telematics. In real-world environments, ASR faces many kinds of sound sources and they...

    Automatic speech recognition (ASR) is becoming an indispensable technology in many application areas such as telephonic system, ubiquitous system, robot system, and telematics. In real-world environments, ASR faces many kinds of sound sources and they should be discriminated to improve ASR performance. In ASR systems, speech is usually detected from the input signal by voice activity detection (VAD) scheme. Speech and music, however, are not easily discriminated by the VAD because they share similar characteristics such as periodicity and frequency. Speech and music discrimination (SMD) has gained much popularity in recent years for efficient coding and automatic retrieval of multimedia sources and automated speech recognition. Many kinds of approaches have previously been taken to the problem of SMD using spectral energy changes, harmonics, delta cepstral energy, power spectrum deviation, cepstral distance, and etc. However, those feature parameters are not so efficient to get high performance with fast output. The mean of minimum cepstral distance (MMCD) has recently showed high performance. However, the duration used in MMCD was 1 second. Considering the length of a sentence, too much long duration is not proper for the practical system. In this paper, we propose the spectral transition rate (STR) as a novel feature for SMD. We found that the spectral peaks of speech are gradually changing. However, the sound of musical instruments in general tends to keep the peak frequencies and energies unchanged for relatively long period of time compared to speech. The STR is based on this observation. The STR of speech is much higher than that of music. Comparing to other algorithms, the STR based SMD gives relatively fast output without losing its performance. The experimental result shows that the STR based SMD method outperforms a conventional method.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 제 2 장 음악과 음성의 주파수 특성 3
    • 2.1 음성의 스펙트로그램 3
    • 2.2 음악의 스펙트로그램 7
    • 2.3 음성과 음악의 주파수 특성 비교 10
    • 제 1 장 서론 1
    • 제 2 장 음악과 음성의 주파수 특성 3
    • 2.1 음성의 스펙트로그램 3
    • 2.2 음악의 스펙트로그램 7
    • 2.3 음성과 음악의 주파수 특성 비교 10
    • 제 3 장 Spectral Transition Rate 11
    • 3.1 STR 11
    • 3.2 음성의 STR 13
    • 3.3 음악의 STR 16
    • 3.4 STR기반의 SMD 19
    • 제 4 장 기존의 음성과 음악 검출 알고리즘 22
    • 4.1 The mean of minimum cepstral distance 22
    • 제 5 장 실험 및 결과 분석 26
    • 5.1 실험 환경 26
    • 5.2 음악 데이터의 종류 27
    • 5.3 실험 결과 27
    • 제 6 장 결론 및 향후 과제 38
    • 참고문헌 39
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼