첫 가창 음성합성 제작물이 등장한 1962년 이후로 60여 년이 지난 현재에 이르기까지, 인공신경망을 적용하여 2019년 Synthesizer V가 등장하고 나서야 인공신경망, 딥러닝 등을 결합한 가창 음성...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17208974
서울: 상명대학교 일반대학원, 2025
학위논문(석사) -- 상명대학교 일반대학원 , 뉴미디어음악학과 , 2025. 2
2025
한국어
음소 ; 포먼트 ; 가창음성합성 ; 음성합성 ; Synthesizer V
786.76 판사항(23)
서울
Study on Korean Utilization Methods through Phoneme Notation Analysis in Synthesizer V
98 p.; 26 cm
상명대학교 논문은 저작권에 의해 보호받습니다.
지도교수:장민호
참고문헌 수록
I804:11028-200000852599
0
상세조회0
다운로드첫 가창 음성합성 제작물이 등장한 1962년 이후로 60여 년이 지난 현재에 이르기까지, 인공신경망을 적용하여 2019년 Synthesizer V가 등장하고 나서야 인공신경망, 딥러닝 등을 결합한 가창 음성...
첫 가창 음성합성 제작물이 등장한 1962년 이후로 60여 년이 지난 현재에 이르기까지, 인공신경망을 적용하여 2019년 Synthesizer V가 등장하고 나서야 인공신경망, 딥러닝 등을 결합한 가창 음성합성 프로그램들이 출시되기 시작했다. 그러나 현재까지 한국어의 원활한 사용이 가능한 프로그램은 없었고, 그 결과로 외국어 가사의 경우에만 한정적으로 사용하는 등 한국의 음악 제작자들은 가창 음성합성 프로그램을 유용하게 활용할 수 없었다. 때문에, 음악 제작 활동의 인적, 비용적인 효율을 위해 가창 음성합성 프로그램을 한국어로 활용하는 방안이 요구된다 하겠다.
본 논문은 이 목적을 달성하기 위해 Synthesizer V의 음소를 활용하여 해당 음소의 포먼트 분석값을 한국어와 비교하였는데, 그 과정은 다음과 같다.
첫 번째로, 가창 음성합성의 발전양상과 음소 단위 편집을 지원하는 여러 가창 음성합성 프로그램을 조사, 최종적으로 연구에 활용할 가창 음성합성 프로그램을 Synthesizer V로 결정하였다.
두 번째로, Synthesizer V에 사용되는 3가지 음소 표기법들에 대하여 알아보고 음소 표기 체계를 통일하기 위해 Amazon Polly에서 제공하는 음소 표와 비교하여 표기법을 통일시키고 표기법상 한국어와 일치하는 음소와 유사한 발음의 음소들을 목록화하였다.
세 번째로, 정리된 목록의 음소를 각각 Synthesizer V로 샘플을 생성하여 포먼트 분석을 실행한 뒤 한국어 포먼트 측정값과 비교하여 가장 유사한 음소를 선택하여 정리하게 되었다.
마지막으로는 Synthesizer V의 피치 파라미터를 조절하여 유사 발음이 존재하지 않았던 된소리의 해결과 종성에 대한 해결방안을 제시하였다.
이상과 같은 과정을 통하여 Synthesizer V에서 사용되는 음소 중 한국어와 포먼트 측정값이 가장 유사한 음소들에 대한 목록을 완성하게 되었다. 해당 분석을 통해 가창 음성합성 프로그램의 한국어 활용을 위한 비교값으로 활용함과 동시에 한국어 활용을 위한 방안 중의 하나로 기대할 수 있을 것이다.
다국어 초록 (Multilingual Abstract)
Since the debut of the first singing voice synthesis in 1962, it took more than 60 years, with advancements in neural networks and the launch of Synthesizer V in 2019, for programs integrating neural networks and deep learning to emerge in the field o...
Since the debut of the first singing voice synthesis in 1962, it took more than 60 years, with advancements in neural networks and the launch of Synthesizer V in 2019, for programs integrating neural networks and deep learning to emerge in the field of singing voice synthesis. However, no program has yet provided seamless support for Korean, limiting its use to foreign-language lyrics. Consequently, Korean music producers have not been able to utilize these singing synthesis programs effectively. Therefore, a method to use these programs in Korean is needed to improve both human and cost efficiency in music production.
This study aims to achieve this by comparing the phoneme formants in Synthesizer V to those of the Korean language through the following process. First, the development of singing voice synthesis and various programs supporting phoneme-level editing were reviewed, ultimately selecting Synthesizer V for the study. Second, an analysis of the three phonetic transcription systems used in Synthesizer V was conducted, and for standardization, comparisons were made with the phonetic chart provided by Amazon Polly. Phonemes matching or resembling Korean sounds were then cataloged. Third, samples of each phoneme from the organized list were generated in Synthesizer V, analyzed for formant values, and compared with Korean formants to identify the closest matches. Finally, pitch parameter adjustments were applied to address aspirated and unreleased consonants that lacked similar sounds in Synthesizer V.
Through these steps, a list of phonemes most closely resembling Korean was completed, serving as a reference for Korean compatibility in singing synthesis and offering a potential solution for enhancing Korean usage in such programs.
목차 (Table of Contents)