RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재 SCI SCIE SCOPUS

    Speaker Tracking Using Eigendecomposition and an Index Tree of Reference Models

    한글로보기

    https://www.riss.kr/link?id=A103370877

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This paper focuses on online speaker tracking for telephone conversations and broadcast news. Since the online applicability imposes some limitations on the tracking strategy, such as data insufficiency, a reliable approach should be applied to compensate for this shortage. In this framework, a set of reference speaker models are used as side information to facilitate online tracking. To improve the indexing accuracy, adaptation approaches in eigenvoice decomposition space are proposed in this paper. We believe that the eigenvoice adaptation techniques would help to embed the speaker space in the models and hence enrich the generality of the selected speaker models. Also, an index structure of the reference models is proposed to speed up the search in the model space. The proposed framework is evaluated on 2002 Rich Transcription Broadcast News and Conversational Telephone Speech corpus as well as a synthetic dataset. The indexing errors of the proposed framework on telephone conversations, broadcast news,and synthetic dataset are 8.77%, 9.36%, and 12.4%,respectively. Using the index tree structure approach, the run time of the proposed framework is improved by 22%.
    번역하기

    This paper focuses on online speaker tracking for telephone conversations and broadcast news. Since the online applicability imposes some limitations on the tracking strategy, such as data insufficiency, a reliable approach should be applied to compen...

    This paper focuses on online speaker tracking for telephone conversations and broadcast news. Since the online applicability imposes some limitations on the tracking strategy, such as data insufficiency, a reliable approach should be applied to compensate for this shortage. In this framework, a set of reference speaker models are used as side information to facilitate online tracking. To improve the indexing accuracy, adaptation approaches in eigenvoice decomposition space are proposed in this paper. We believe that the eigenvoice adaptation techniques would help to embed the speaker space in the models and hence enrich the generality of the selected speaker models. Also, an index structure of the reference models is proposed to speed up the search in the model space. The proposed framework is evaluated on 2002 Rich Transcription Broadcast News and Conversational Telephone Speech corpus as well as a synthetic dataset. The indexing errors of the proposed framework on telephone conversations, broadcast news,and synthetic dataset are 8.77%, 9.36%, and 12.4%,respectively. Using the index tree structure approach, the run time of the proposed framework is improved by 22%.

    더보기

    참고문헌 (Reference)

    1 S. Kwon, "Unsupervised Speaker Indexing Using Generic Models" 13 : 1004-1013, 2004

    2 Y.K. Muthusamy, "The OGI Multi-language Telephone Speech Corpus" 895-898, 1992

    3 C. Wooters, "The ICSI RT07s Speaker Diarization System" 2007

    4 "The 2009 (RT-09) Rich Transcription Evaluation Plan"

    5 M.Bijankhan,Great Farsdat Database, "Technical report" Iran Research Center on Intelligent Signal Processing 2002

    6 M. Davy, "Supervised Classification Using MCMC Methods" 33-36, 2000

    7 S. Meignier, "Step-by-Step and Integrated Approaches in Broadcast News Speaker Diarization" 20 (20): 303-330, 2006

    8 D.A. Reynolds, "Speaker Verification Using Adapted Gaussian Mixture Models" 10 (10): 19-41, 2000

    9 M. Kotti, "Speaker Segmentation and Clustering" 88 (88): 1091-1124, 2008

    10 A. Martin, "Speaker Recognition in a Multispeaker Environment" 787-790, 2001

    1 S. Kwon, "Unsupervised Speaker Indexing Using Generic Models" 13 : 1004-1013, 2004

    2 Y.K. Muthusamy, "The OGI Multi-language Telephone Speech Corpus" 895-898, 1992

    3 C. Wooters, "The ICSI RT07s Speaker Diarization System" 2007

    4 "The 2009 (RT-09) Rich Transcription Evaluation Plan"

    5 M.Bijankhan,Great Farsdat Database, "Technical report" Iran Research Center on Intelligent Signal Processing 2002

    6 M. Davy, "Supervised Classification Using MCMC Methods" 33-36, 2000

    7 S. Meignier, "Step-by-Step and Integrated Approaches in Broadcast News Speaker Diarization" 20 (20): 303-330, 2006

    8 D.A. Reynolds, "Speaker Verification Using Adapted Gaussian Mixture Models" 10 (10): 19-41, 2000

    9 M. Kotti, "Speaker Segmentation and Clustering" 88 (88): 1091-1124, 2008

    10 A. Martin, "Speaker Recognition in a Multispeaker Environment" 787-790, 2001

    11 D.A. Reynolds, "Speaker Identification and Verification Using Gaussian Mixture Speaker Models" 17 (17): 91-108, 1995

    12 K. Iso, "Speaker Clustering Using Vector Quantization and Spectral Clustering" Proc. ICASSP 4986-4989, 2010

    13 P. Zezula, "Similarity Search: The Metric Space Approach" 32 : 23-38, 2006

    14 S. Berrani, "Robust Content-Based Image Searches for Copyright Protection" 70-77, 2003

    15 R. Kuhn et, "Rapid Speaker Adaptation in Eigenvoice Space" 8 (8): 695-707, 2000

    16 I.T. Jolliffe, "Principal Component Analysis" Springer-Verlag 1986

    17 A.K. Noulas, "Online Multimodal Speaker Diarization" 350-357, 2007

    18 S. Kullback, "On Information and Sufficiency" 22 (22): 79-86, 1951

    19 K. Markov, "Never-Ending Learning with Dynamic Hidden Markov Network" 1437-1440, 2007

    20 J. Garofolo, "NIST Rich Transcription 2002 Evaluation: A Preview" 2002

    21 A. Dempster, "Maximum Likelihood from Incomplete Data via the EM Algorithm" 39 (39): 1-38, 1977

    22 M. Zamalloa, "Low Latency Online Speaker Tracking on the AMI Corpus of Meeting Conversations" 4962-4965, 2010

    23 K. Markov, "Improved Novelty Detection for Online GMM Based Speaker Diarization" 363-366, 2008

    24 J. Schmalenstroeer, "Fusing Audio and Video Information for Online Speaker Diarization" 1163-1166, 2007

    25 L.R. Rabiner, "Fundamentals of Speech Recognition" Prentice-Hall 1993

    26 X. Anguera, "Frame Purification for Cluster Comparison in Speaker Diarization" 135-139, 2006

    27 K. Chen, "Fast Speaker Adaptation Using Eigenspace-Based Maximum Likelihood Linear Regression" 742-745, 2000

    28 T.H. Nguyen, "Cluster Criterion Functions in Spectral Subspace and Their Application in Speaker Clustering" 4085-4088, 2009

    29 Mohammad Hossein Moattar, "A Weighted Feature Voting Approach for Robust and Real-Time Voice Activity Detection" 한국전자통신연구원 33 (33): 99-109, 2011

    30 B. Mak, "A Study of Various Composite Kernels for Kernel Eigenvoice Speaker Adaptation" 1 : 325-328, 2004

    31 C.H. Huang, "A New Eigenvoice Approach to Speaker Adaptation" 109-112, 2004

    32 C. Vaquero, "A Hybrid Approach to Online Speaker Diarization" 2638-2631, 2010

    33 W. Wang, "A Decision-Tree-Based Online Speaker Clustering" 4477 : 555-562, 2007

    더보기

    동일학술지(권/호) 다른 논문

    동일학술지 더보기

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2023 평가 해외DB학술지평가 신청대상 (해외등재 학술지 평가)
    2020-01-01 등재 등재학술지 유지 (해외등재 학술지 평가) KCI등재
    2005-09-27 학술지등록 한글명 : ETRI Journal
    외국어명 : ETRI Journal
    KCI등재
    2003-01-01 등재 SCI 등재 (신규평가) KCI등재
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 0.78 0.28 0.57
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    0.47 0.42 0.4 0.06
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼