RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    악기 종류에 무관한 보컬 분리를 위한 디퓨전 기반 반사실적 생성 기법 = A Diffusion based Counterfactual Generation Method for Instrument-Independent Vocal Separation

    한글로보기

    https://www.riss.kr/link?id=T17243917

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구에서는 디퓨전 기반 반사실적 생성 기법을 이용하여 악기 종류에 관계없이 일관된 성능을 유지할 수 있는 보컬 분리 방법론을 제안한다. 기존의 보컬 분리 연구는 대부분 특정 악기에만 적합한 데이터셋을 활용하여 모델을 학습함으로써, 데이터셋에 존재하지 않는 새로운 악기가 사용된 음악에서는 성능 저하를 보이는 경향이 있었다. 이를 해결하기 위해 본 연구에서는 반사실적 생성 기법을 활용하여 혼합 음원에서 보컬 스템을 분리하는 새로운 방식을 제안하였다. 반사실적 생성은 ‘만약 반주가 없었다면?’ 이라는 가정 하에 주어진 음악으로부터 가정에 부합하는 가상의 데이터셋을 생성하도록 모델을 학습시키는 기법으로, 데이터 속 특성이 데이터의 상태에 미치는 영향을 효과 적으로 학습할 수 있도록 돕는다. 본 연구는 디퓨전 기반의 음성 생성 모델에 반사실적 생성 기법을 적용한 형태의 방법론을 제안하였고, 기존 보컬 분리 연구에서 사용하는 벤치마크 데이터셋, 해당 데이터셋에 존재하지 않는 새로운 악기가 존재하는 데이터셋을 이용해 평가를 수행하였다. 실험 결과, 제안된 방법론은 학습 과정에서 벤치마크 데이 터셋의 학습 데이터만을 사용했음에도 불구하고 추가적인 데이터셋을 활용한 기존의 최신 성능 모델들보다 보컬 분리 측면에서 높은 성능을 기록하였다. 비록 음질 측면에서 기존의 최신 성능 모델보다 좋지 못한 보컬 스템을 생성했다는 한계점이 존재하지만, 제안 기법을 기존 데이터셋에 존재하지 않는 새로운 악기가 사용되는 음악의 보컬 분리 작업 및 다양한 종류의 악기 스템에 적용해 확장할 수 있는 가능성을 확인하였다.
    번역하기

    본 연구에서는 디퓨전 기반 반사실적 생성 기법을 이용하여 악기 종류에 관계없이 일관된 성능을 유지할 수 있는 보컬 분리 방법론을 제안한다. 기존의 보컬 분리 연구는 대부분 특정 악기...

    본 연구에서는 디퓨전 기반 반사실적 생성 기법을 이용하여 악기 종류에 관계없이 일관된 성능을 유지할 수 있는 보컬 분리 방법론을 제안한다. 기존의 보컬 분리 연구는 대부분 특정 악기에만 적합한 데이터셋을 활용하여 모델을 학습함으로써, 데이터셋에 존재하지 않는 새로운 악기가 사용된 음악에서는 성능 저하를 보이는 경향이 있었다. 이를 해결하기 위해 본 연구에서는 반사실적 생성 기법을 활용하여 혼합 음원에서 보컬 스템을 분리하는 새로운 방식을 제안하였다. 반사실적 생성은 ‘만약 반주가 없었다면?’ 이라는 가정 하에 주어진 음악으로부터 가정에 부합하는 가상의 데이터셋을 생성하도록 모델을 학습시키는 기법으로, 데이터 속 특성이 데이터의 상태에 미치는 영향을 효과 적으로 학습할 수 있도록 돕는다. 본 연구는 디퓨전 기반의 음성 생성 모델에 반사실적 생성 기법을 적용한 형태의 방법론을 제안하였고, 기존 보컬 분리 연구에서 사용하는 벤치마크 데이터셋, 해당 데이터셋에 존재하지 않는 새로운 악기가 존재하는 데이터셋을 이용해 평가를 수행하였다. 실험 결과, 제안된 방법론은 학습 과정에서 벤치마크 데이 터셋의 학습 데이터만을 사용했음에도 불구하고 추가적인 데이터셋을 활용한 기존의 최신 성능 모델들보다 보컬 분리 측면에서 높은 성능을 기록하였다. 비록 음질 측면에서 기존의 최신 성능 모델보다 좋지 못한 보컬 스템을 생성했다는 한계점이 존재하지만, 제안 기법을 기존 데이터셋에 존재하지 않는 새로운 악기가 사용되는 음악의 보컬 분리 작업 및 다양한 종류의 악기 스템에 적용해 확장할 수 있는 가능성을 확인하였다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to specific instruments, resulting in a performance decline when the music contains new instruments not present in the dataset. To address this, we propose a novel approach that utilizes counterfactual generation to separate vocal stems from mixed audio. Counterfactual generation is a technique that trains the model to generate hypothetical datasets based on assumptions like ”What if there were no accompaniment?” This approach helps the model learn the effect of certain features on the state of the data effectively. In this study, we propose a methodology that ap- plies the counterfactual generation technique to a diffusion-based audio generation model. We evaluate the proposed approach using benchmark datasets commonly used in vocal separation research, as well as datasets containing new instruments not present in the benchmark datasets. Experimental results show that the proposed approach outperformed existing state-of-the-art models in terms of vocal separation, despite only using the benchmark dataset for training without leveraging additional datasets. Although there is a limitation in generating vocal stems with lower audio quality compared to state-of-the-art models, we confirm the potential to extend the proposed method to vocal separation tasks for music containing new instruments not present in the original dataset and to various types of instrument stems.
    번역하기

    This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to speci...

    This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to specific instruments, resulting in a performance decline when the music contains new instruments not present in the dataset. To address this, we propose a novel approach that utilizes counterfactual generation to separate vocal stems from mixed audio. Counterfactual generation is a technique that trains the model to generate hypothetical datasets based on assumptions like ”What if there were no accompaniment?” This approach helps the model learn the effect of certain features on the state of the data effectively. In this study, we propose a methodology that ap- plies the counterfactual generation technique to a diffusion-based audio generation model. We evaluate the proposed approach using benchmark datasets commonly used in vocal separation research, as well as datasets containing new instruments not present in the benchmark datasets. Experimental results show that the proposed approach outperformed existing state-of-the-art models in terms of vocal separation, despite only using the benchmark dataset for training without leveraging additional datasets. Although there is a limitation in generating vocal stems with lower audio quality compared to state-of-the-art models, we confirm the potential to extend the proposed method to vocal separation tasks for music containing new instruments not present in the original dataset and to various types of instrument stems.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to specific instruments, resulting in a performance decline when the music contains new instruments not present in the dataset. To address this, we propose a novel approach that utilizes counterfactual generation to separate vocal stems from mixed audio. Counterfactual generation is a technique that trains the model to generate hypothetical datasets based on assumptions like ”What if there were no accompaniment?” This approach helps the model learn the effect of certain features on the state of the data effectively. In this study, we propose a methodology that ap- plies the counterfactual generation technique to a diffusion-based audio generation model. We evaluate the proposed approach using benchmark datasets commonly used in vocal separation research, as well as datasets containing new instruments not present in the benchmark datasets. Experimental results show that the proposed approach outperformed existing state-of-the-art models in terms of vocal separation, despite only using the benchmark dataset for training without leveraging additional datasets. Although there is a limitation in generating vocal stems with lower audio quality compared to state-of-the-art models, we confirm the potential to extend the proposed method to vocal separation tasks for music containing new instruments not present in the original dataset and to various types of instrument stems.
    번역하기

    This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to speci...

    This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to specific instruments, resulting in a performance decline when the music contains new instruments not present in the dataset. To address this, we propose a novel approach that utilizes counterfactual generation to separate vocal stems from mixed audio. Counterfactual generation is a technique that trains the model to generate hypothetical datasets based on assumptions like ”What if there were no accompaniment?” This approach helps the model learn the effect of certain features on the state of the data effectively. In this study, we propose a methodology that ap- plies the counterfactual generation technique to a diffusion-based audio generation model. We evaluate the proposed approach using benchmark datasets commonly used in vocal separation research, as well as datasets containing new instruments not present in the benchmark datasets. Experimental results show that the proposed approach outperformed existing state-of-the-art models in terms of vocal separation, despite only using the benchmark dataset for training without leveraging additional datasets. Although there is a limitation in generating vocal stems with lower audio quality compared to state-of-the-art models, we confirm the potential to extend the proposed method to vocal separation tasks for music containing new instruments not present in the original dataset and to various types of instrument stems.

    더보기

    목차 (Table of Contents)

    • 초록 i
    • 목차 iii
    • 표 목차 iv
    • 그림 목차 v
    • 제 1 장 서론 1
    • 초록 i
    • 목차 iii
    • 표 목차 iv
    • 그림 목차 v
    • 제 1 장 서론 1
    • 1.1 연구배경및동기 1
    • 1.2 연구목적 4
    • 1.3 문제정의 5
    • 1.4 논문구성 6
    • 제 2 장 배경이론및관련연구 7
    • 2.1 배경이론 7
    • 2.1.1 딥러닝을활용한보컬분리 7
    • 2.2 관련연구 15
    • 2.2.1 악기종류에대한강건성확보를위한연구 15
    • 2.2.2 디퓨전기반보컬분리 17
    • 2.2.3 반사실적생성기법 20
    • 제 3 장 반사실적생성기법을적용한보컬분리 23
    • 3.1 문제정의 23
    • 3.2 적용모델및비교모델 26
    • 3.2.1 반사실적생성기법을적용한모델 26
    • 3.2.2 비교모델 28
    • 3.3 제안기법 31
    • 3.4 평가척도 34
    • 3.4.1 신호처리기반보컬분리성능지표 34
    • 3.4.2 MeanOpinionScore 36
    • 제 4 장 실험결과 38
    • 4.1 데이터셋 38
    • 4.2 실험세팅 41
    • 4.2.1 데이터처리및증강 41
    • 4.2.2 성능평가 43
    • 4.3 실험결과 45
    • 4.3.1 정량적평가 45
    • 4.3.2 정성적평가 47
    • 4.3.3 결과분석 49
    • 제 5 장 결론 54
    • 5.1 결론 54
    • 5.2 향후연구 57
    • 참고문헌 59
    • Abstract 66
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼