본 연구에서는 디퓨전 기반 반사실적 생성 기법을 이용하여 악기 종류에 관계없이 일관된 성능을 유지할 수 있는 보컬 분리 방법론을 제안한다. 기존의 보컬 분리 연구는 대부분 특정 악기...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 연구에서는 디퓨전 기반 반사실적 생성 기법을 이용하여 악기 종류에 관계없이 일관된 성능을 유지할 수 있는 보컬 분리 방법론을 제안한다. 기존의 보컬 분리 연구는 대부분 특정 악기...
본 연구에서는 디퓨전 기반 반사실적 생성 기법을 이용하여 악기 종류에 관계없이 일관된 성능을 유지할 수 있는 보컬 분리 방법론을 제안한다. 기존의 보컬 분리 연구는 대부분 특정 악기에만 적합한 데이터셋을 활용하여 모델을 학습함으로써, 데이터셋에 존재하지 않는 새로운 악기가 사용된 음악에서는 성능 저하를 보이는 경향이 있었다. 이를 해결하기 위해 본 연구에서는 반사실적 생성 기법을 활용하여 혼합 음원에서 보컬 스템을 분리하는 새로운 방식을 제안하였다. 반사실적 생성은 ‘만약 반주가 없었다면?’ 이라는 가정 하에 주어진 음악으로부터 가정에 부합하는 가상의 데이터셋을 생성하도록 모델을 학습시키는 기법으로, 데이터 속 특성이 데이터의 상태에 미치는 영향을 효과 적으로 학습할 수 있도록 돕는다. 본 연구는 디퓨전 기반의 음성 생성 모델에 반사실적 생성 기법을 적용한 형태의 방법론을 제안하였고, 기존 보컬 분리 연구에서 사용하는 벤치마크 데이터셋, 해당 데이터셋에 존재하지 않는 새로운 악기가 존재하는 데이터셋을 이용해 평가를 수행하였다. 실험 결과, 제안된 방법론은 학습 과정에서 벤치마크 데이 터셋의 학습 데이터만을 사용했음에도 불구하고 추가적인 데이터셋을 활용한 기존의 최신 성능 모델들보다 보컬 분리 측면에서 높은 성능을 기록하였다. 비록 음질 측면에서 기존의 최신 성능 모델보다 좋지 못한 보컬 스템을 생성했다는 한계점이 존재하지만, 제안 기법을 기존 데이터셋에 존재하지 않는 새로운 악기가 사용되는 음악의 보컬 분리 작업 및 다양한 종류의 악기 스템에 적용해 확장할 수 있는 가능성을 확인하였다.
다국어 초록 (Multilingual Abstract)
This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to speci...
This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to specific instruments, resulting in a performance decline when the music contains new instruments not present in the dataset. To address this, we propose a novel approach that utilizes counterfactual generation to separate vocal stems from mixed audio. Counterfactual generation is a technique that trains the model to generate hypothetical datasets based on assumptions like ”What if there were no accompaniment?” This approach helps the model learn the effect of certain features on the state of the data effectively. In this study, we propose a methodology that ap- plies the counterfactual generation technique to a diffusion-based audio generation model. We evaluate the proposed approach using benchmark datasets commonly used in vocal separation research, as well as datasets containing new instruments not present in the benchmark datasets. Experimental results show that the proposed approach outperformed existing state-of-the-art models in terms of vocal separation, despite only using the benchmark dataset for training without leveraging additional datasets. Although there is a limitation in generating vocal stems with lower audio quality compared to state-of-the-art models, we confirm the potential to extend the proposed method to vocal separation tasks for music containing new instruments not present in the original dataset and to various types of instrument stems.
다국어 초록 (Multilingual Abstract)
This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to speci...
This study proposes a vocal separation methodology that is independent to various instrument types using a diffusion-based counterfactual generation technique. Conventional vocal separation research has typically used datasets that are suited to specific instruments, resulting in a performance decline when the music contains new instruments not present in the dataset. To address this, we propose a novel approach that utilizes counterfactual generation to separate vocal stems from mixed audio. Counterfactual generation is a technique that trains the model to generate hypothetical datasets based on assumptions like ”What if there were no accompaniment?” This approach helps the model learn the effect of certain features on the state of the data effectively. In this study, we propose a methodology that ap- plies the counterfactual generation technique to a diffusion-based audio generation model. We evaluate the proposed approach using benchmark datasets commonly used in vocal separation research, as well as datasets containing new instruments not present in the benchmark datasets. Experimental results show that the proposed approach outperformed existing state-of-the-art models in terms of vocal separation, despite only using the benchmark dataset for training without leveraging additional datasets. Although there is a limitation in generating vocal stems with lower audio quality compared to state-of-the-art models, we confirm the potential to extend the proposed method to vocal separation tasks for music containing new instruments not present in the original dataset and to various types of instrument stems.
목차 (Table of Contents)