RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Improving Conditional Generation of Musical Components: Focusing on Chord and Expression = 음악적 요소에 대한 조건부 생성의 개선에 관한 연구: 화음과 표현을 중심으로

    한글로보기

    https://www.riss.kr/link?id=T16750055

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Conditional generation of musical components (CGMC) creates a part of music based on partial musical components such as melody or chord. CGMC is beneficial for discovering complex relationships among musical attributes. It can also assist non-experts who face difficulties in making music. However, recent studies for CGMC are still facing two challenges in terms of generation quality and model controllability. First, the structure of the generated music is not robust. Second, only limited ranges of musical factors and tasks have been examined as targets for flexible control of generation. In this thesis, we aim to mitigate these two challenges to improve the CGMC systems. For musical structure, we focus on intuitive modeling of musical hierarchy to help the model explicitly learn musically meaningful dependency. To this end, we utilize alignment paths between the raw music data and the musical units such as notes or chords. For musical creativity, we facilitate smooth control of novel musical attributes using latent representations. We attempt to achieve disentangled representations of the intended factors by regularizing them with data-driven inductive bias. This thesis verifies the proposed approaches particularly in two representative CGMC tasks, melody harmonization and expressive performance rendering. A variety of experimental results show the possibility of the proposed approaches to expand musical creativity under stable generation quality.
    번역하기

    Conditional generation of musical components (CGMC) creates a part of music based on partial musical components such as melody or chord. CGMC is beneficial for discovering complex relationships among musical attributes. It can also assist non-experts ...

    Conditional generation of musical components (CGMC) creates a part of music based on partial musical components such as melody or chord. CGMC is beneficial for discovering complex relationships among musical attributes. It can also assist non-experts who face difficulties in making music. However, recent studies for CGMC are still facing two challenges in terms of generation quality and model controllability. First, the structure of the generated music is not robust. Second, only limited ranges of musical factors and tasks have been examined as targets for flexible control of generation. In this thesis, we aim to mitigate these two challenges to improve the CGMC systems. For musical structure, we focus on intuitive modeling of musical hierarchy to help the model explicitly learn musically meaningful dependency. To this end, we utilize alignment paths between the raw music data and the musical units such as notes or chords. For musical creativity, we facilitate smooth control of novel musical attributes using latent representations. We attempt to achieve disentangled representations of the intended factors by regularizing them with data-driven inductive bias. This thesis verifies the proposed approaches particularly in two representative CGMC tasks, melody harmonization and expressive performance rendering. A variety of experimental results show the possibility of the proposed approaches to expand musical creativity under stable generation quality.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    음악적 요소를 조건부 생성하는 분야인 CGMC는 멜로디나 화음과 같은 음악의 일부분을 기반으로 나머지 부분을 생성하는 것을 목표로 한다. 이 분야는 음악적 요소 간 복잡한 관계를 탐구하는 데 용이하고, 음악을 만드는 데 어려움을 겪는 비전문가들을 도울 수 있다. 최근 연구들은 딥러닝 기술을 활용하여 CGMC 시스템의 성능을 높여왔다. 하지만, 이러한 연구들에는 아직 생성 품질과 제어가능성 측면에서 두 가지의 한계점이 있다. 먼저, 생성된 음악의 음악적 구조가 명확하지 않다. 또한, 아직 좁은 범위의 음악적 요소 및 테스크만이 유연한 제어의 대상으로서 탐구되었다. 이에 본 학위논문에서는 CGMC의 개선을 위해 위 두 가지의 한계점을 해결하고자 한다. 첫 번째로, 음악 구조를 이루는 음악적 위계를 직관적으로 모델링하는 데 집중하고자 한다. 본래 데이터와 음, 화음과 같은 음악적 단위 간 정렬 경로를 사용하여 모델이 음악적으로 의미있는 종속성을 명확하게 배울 수 있도록 한다. 두 번째로, 잠재 표상을 활용하여 새로운 음악적 요소들을 유연하게 제어하고자 한다. 특히 잠재 표상이 의도된 요소에 대해 풀리도록 훈련하기 위해서 비지도 혹은 자가지도 학습 프레임워크을 사용하여 잠재 표상을 제한하도록 한다. 본 학위논문에서는 CGMC 분야의 대표적인 두 테스크인 멜로디 하모나이제이션 및 표현적 연주 렌더링 테스크에 대해 위의 두 가지 방법론을 검증한다. 다양한 실험적 결과들을 통해 제안한 방법론이 CGMC 시스템의 음악적 창의성을 안정적인 생성 품질로 확장할 수 있다는 가능성을 시사한다.
    번역하기

    음악적 요소를 조건부 생성하는 분야인 CGMC는 멜로디나 화음과 같은 음악의 일부분을 기반으로 나머지 부분을 생성하는 것을 목표로 한다. 이 분야는 음악적 요소 간 복잡한 관계를 탐구하...

    음악적 요소를 조건부 생성하는 분야인 CGMC는 멜로디나 화음과 같은 음악의 일부분을 기반으로 나머지 부분을 생성하는 것을 목표로 한다. 이 분야는 음악적 요소 간 복잡한 관계를 탐구하는 데 용이하고, 음악을 만드는 데 어려움을 겪는 비전문가들을 도울 수 있다. 최근 연구들은 딥러닝 기술을 활용하여 CGMC 시스템의 성능을 높여왔다. 하지만, 이러한 연구들에는 아직 생성 품질과 제어가능성 측면에서 두 가지의 한계점이 있다. 먼저, 생성된 음악의 음악적 구조가 명확하지 않다. 또한, 아직 좁은 범위의 음악적 요소 및 테스크만이 유연한 제어의 대상으로서 탐구되었다. 이에 본 학위논문에서는 CGMC의 개선을 위해 위 두 가지의 한계점을 해결하고자 한다. 첫 번째로, 음악 구조를 이루는 음악적 위계를 직관적으로 모델링하는 데 집중하고자 한다. 본래 데이터와 음, 화음과 같은 음악적 단위 간 정렬 경로를 사용하여 모델이 음악적으로 의미있는 종속성을 명확하게 배울 수 있도록 한다. 두 번째로, 잠재 표상을 활용하여 새로운 음악적 요소들을 유연하게 제어하고자 한다. 특히 잠재 표상이 의도된 요소에 대해 풀리도록 훈련하기 위해서 비지도 혹은 자가지도 학습 프레임워크을 사용하여 잠재 표상을 제한하도록 한다. 본 학위논문에서는 CGMC 분야의 대표적인 두 테스크인 멜로디 하모나이제이션 및 표현적 연주 렌더링 테스크에 대해 위의 두 가지 방법론을 검증한다. 다양한 실험적 결과들을 통해 제안한 방법론이 CGMC 시스템의 음악적 창의성을 안정적인 생성 품질로 확장할 수 있다는 가능성을 시사한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1 Introduction 1
    • 1.1 Motivation 5
    • 1.2 Definitions 8
    • 1.3 Tasks of Interest 10
    • 1.3.1 Generation Quality 10
    • Chapter 1 Introduction 1
    • 1.1 Motivation 5
    • 1.2 Definitions 8
    • 1.3 Tasks of Interest 10
    • 1.3.1 Generation Quality 10
    • 1.3.2 Controllability 12
    • 1.4 Approaches 13
    • 1.4.1 Modeling Musical Hierarchy 14
    • 1.4.2 Regularizing Latent Representations 16
    • 1.4.3 Target Tasks 18
    • 1.5 Outline of the Thesis 19
    • Chapter 2 Background 22
    • 2.1 Music Generation Tasks 23
    • 2.1.1 Melody Harmonization 23
    • 2.1.2 Expressive Performance Rendering 25
    • 2.2 Structure-enhanced Music Generation 27
    • 2.2.1 Hierarchical Music Generation 27
    • 2.2.2 Transformer-based Music Generation 28
    • 2.3 Disentanglement Learning 29
    • 2.3.1 Unsupervised Approaches 30
    • 2.3.2 Supervised Approaches 30
    • 2.3.3 Self-supervised Approaches 31
    • 2.4 Controllable Music Generation 32
    • 2.4.1 Score Generation 32
    • 2.4.2 Performance Rendering 33
    • 2.5 Summary 34
    • Chapter 3 Translating Melody to Chord: Structured and Flexible Harmonization of Melody with Transformer 36
    • 3.1 Introduction 36
    • 3.2 Proposed Methods 41
    • 3.2.1 Standard Transformer Model (STHarm) 41
    • 3.2.2 Variational Transformer Model (VTHarm) 44
    • 3.2.3 Regularized Variational Transformer Model (rVTHarm) 46
    • 3.2.4 Training Objectives 47
    • 3.3 Experimental Settings 48
    • 3.3.1 Datasets 49
    • 3.3.2 Comparative Methods 50
    • 3.3.3 Training 50
    • 3.3.4 Metrics 51
    • 3.4 Evaluation 56
    • 3.4.1 Chord Coherence and Diversity 57
    • 3.4.2 Harmonic Similarity to Human 59
    • 3.4.3 Controlling Chord Complexity 60
    • 3.4.4 Subjective Evaluation 62
    • 3.4.5 Qualitative Results 67
    • 3.4.6 Ablation Study 73
    • 3.5 Conclusion and Future Work 74
    • Chapter 4 Sketching the Expression: Flexible Rendering of Expressive Piano Performance with Self-supervised Learning 76
    • 4.1 Introduction 76
    • 4.2 Proposed Methods 79
    • 4.2.1 Data Representation 79
    • 4.2.2 Modeling Musical Hierarchy 80
    • 4.2.3 Overall Network Architecture 81
    • 4.2.4 Regularizing the Latent Variables 84
    • 4.2.5 Overall Objective 86
    • 4.3 Experimental Settings 87
    • 4.3.1 Dataset and Implementation 87
    • 4.3.2 Comparative Methods 88
    • 4.4 Evaluation 88
    • 4.4.1 Generation Quality 89
    • 4.4.2 Disentangling Latent Representations 90
    • 4.4.3 Controllability of Expressive Attributes 91
    • 4.4.4 KL Divergence 93
    • 4.4.5 Ablation Study 94
    • 4.4.6 Subjective Evaluation 95
    • 4.4.7 Qualitative Examples 97
    • 4.4.8 Extent of Control 100
    • 4.5 Conclusion 102
    • Chapter 5 Conclusion and Future Work 103
    • 5.1 Conclusion 103
    • 5.2 Future Work 106
    • 5.2.1 Deeper Investigation of Controllable Factors 106
    • 5.2.2 More Analysis of Qualitative Evaluation Results 107
    • 5.2.3 Improving Diversity and Scale of Dataset 108
    • Bibliography 109
    • 초 록 137
    더보기

    참고문헌 (Reference)

    1. Attention is all you need, A. N. Gomez, J. Uszkoreit, L. Jones, N. Parmar, A. Vaswani, N. Shazeer, L. Kaiser, I. Polosukhin, Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, United States, , 2017

    2. Modelling musical dynamics, T. H. Axel Berndt, Proceedings of the 5th Audio Mostly Conference, New York, United States, , 2010

    3. Counterpoint by convolution,, T. Cooijmans, C.-Z. A. Huang, A. Roberts, A. Courville, D. Eck, Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, , 2017

    4. Disentangling by factorising, H. Kim, A. Mnih, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018

    5. Disentangled sequential autoencoder, Y. Li, S. Mandt, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018

    6. Music performance analysis: A survey, C. Arthur, A. Pati, S. Gururani, A. Lerch, Proceedings of the 20st International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019

    7. Toward controlled generation of text, Z. Yang, X. Liang, R. Salakhutdinov, E. P. Xing, Z. Hu, Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, , 2017

    8. DeepJ: Style-specific music generation,, T. Shin, G. W. Cottrell, H. H. Mao, Proceedings of the 12th International Conference on Semantic Computing, Laguna Hills, United States, , 2018

    9. Detecting harmonic change in musical audio, C. A. Harte, M. B. Sandler, M. Gasser, Proceedings of the 1st ACM International Conference on Multimedia, Santa Barbara, United States, , 2006

    10. Chord2Vec: Learning musical chord embeddings, C. Walder, S. Madjiheurem, L. Qu, Proceedings of the 30th Conference on Neural Information Processing Systems, Barcelona, Spain, , 2016

    1. Attention is all you need, A. N. Gomez, J. Uszkoreit, L. Jones, N. Parmar, A. Vaswani, N. Shazeer, L. Kaiser, I. Polosukhin, Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, United States, , 2017

    2. Modelling musical dynamics, T. H. Axel Berndt, Proceedings of the 5th Audio Mostly Conference, New York, United States, , 2010

    3. Counterpoint by convolution,, T. Cooijmans, C.-Z. A. Huang, A. Roberts, A. Courville, D. Eck, Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, , 2017

    4. Disentangling by factorising, H. Kim, A. Mnih, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018

    5. Disentangled sequential autoencoder, Y. Li, S. Mandt, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018

    6. Music performance analysis: A survey, C. Arthur, A. Pati, S. Gururani, A. Lerch, Proceedings of the 20st International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019

    7. Toward controlled generation of text, Z. Yang, X. Liang, R. Salakhutdinov, E. P. Xing, Z. Hu, Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, , 2017

    8. DeepJ: Style-specific music generation,, T. Shin, G. W. Cottrell, H. H. Mao, Proceedings of the 12th International Conference on Semantic Computing, Laguna Hills, United States, , 2018

    9. Detecting harmonic change in musical audio, C. A. Harte, M. B. Sandler, M. Gasser, Proceedings of the 1st ACM International Conference on Multimedia, Santa Barbara, United States, , 2006

    10. Chord2Vec: Learning musical chord embeddings, C. Walder, S. Madjiheurem, L. Qu, Proceedings of the 30th Conference on Neural Information Processing Systems, Barcelona, Spain, , 2016

    11. Self-supervised audio-visual co-segmentation, H. Zhao, A. Torralba, A. Rouditchenko, J. McDermott, H. Gan, Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing, Brighton, United Kingdom, , 2019

    12. Learning a latent space of multitrack measures, I. Simon, A. Roberts, C. Raffel, D. Eck, J. Engel, C. Hawthorne, Proceedings of the 32nd Conference on Neural Information Processing Systems, Montréal, Canada, , 2018

    13. Hierarchical variational autoencoders for music, J. Engel, D. Eck, A. Roberts, Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, United States, , 2017

    14. Neural speech synthesis with Transformer network, M. Liu, Y. Liu, S. Liu, S. Zhao, M. Zhou, N. Li, Proceedings of the 33rd AAAI Conference on Artificial Intelligence, Honolulu, United States, , 2019

    15. Music perception, pitch, and the auditory system,, A. J. Oxenham, J. H. McDermott, vol. 18, no. 4, pp. 452–463, , 2008

    16. Harmonic, melodic, and functional automatic analysis, D. Rizo, P. R. Illescas, J. M. Iñesta, Proceedings of the 33rd International Computer Music Conference, Copenhagen, Denmark, , 2007

    17. A recurrent latent variable model for sequential data, K. Kastner, A. C. Courville, Y. Bengio, L. Dinh, J. Chung, K. Goel, Proceedings of the 29th Conference on Neural Information Processing Systems, Montréal, Canada, , 2015

    18. Encoding musical style with Transformer autoencoders,, J. Engel, C. Hawthorne, M. Dinculescu, K. Choi, I. Simon, Proceedings of the 37th International Conference on Machine Learning, Online, , 2020

    19. Schenkerian analysis by computer: A proof of concept,, A. Marsden, vol. 39, no. 3, pp. 269–289, , 2010

    20. Colorization as a proxy task for visual understandings,, G. Larsson, M. Maire, G. Shakhnarovich, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, United States, , 2017

    21. Functional harmonic analysis using probabilistic models,, J. Stoddard, C. Raphael, vol. 28, no. 3, pp. 45–52, , 2004

    22. From time to time: The representation of timing and tempo, H. Honing, vol. 25, , 2001

    23. Chord generation from symbolic melody using BLSTM networks,, H. Lim, S. Rhyu, K. Lee, Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, , 2017

    24. Deep music analogy via latent representation disentanglement, D. Wang, J. Jiang, G. Xia, Z. Wang, R. Yang, T. Chen, Proceedings of the 20th International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019

    25. Melody harmonization with interpolated probabilistic models,, E. Vincent, S. Fukayama, S. A. Raczyński, vol. 42, no. 3, pp. 223–235, , 2013

    26. Emotional coloring of computer-controlled music performances,, R. Bresin, A. Friberg, vol. 24, , 2000

    27. MySong: Automatic accompaniment generation for vocal melodies, D. Morris, I. Simon, S. Basu, Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, Florence, Italy, , 2008

    28. AI methods in algorithmic composition: A comprehensive survey,, F. Vico, J. D. Fernández, vol. 48, pp. 513–582, , 2013

    29. Learning to traverse latent spaces for musical score inpainting, A. Lerch, G. Hadjeres, A. Pati, Proceedings of the 20th International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019

    30. A directional interval class representation of chord transitions, E. Cambouropoulos, Proceedings of the 12th International Conference on Music Perception and Cognition, Thessaloniki, Greece, , 2012

    31. Melody lead in piano performance: Expressive device or artifact?, W. Goebl, vol. 110, , 2001

    32. Unsupervised visual representation learning by context prediction, C. Doersch, A. A. Efros, A. Gupta, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Santiago, Chile, , 2015

    33. Automatic melody harmonization with triad chords: A comparative study,, H.-M. Liu, H.-W. Dong, Y. Chen, T. Kitahara, B. Genchel, Y.-H. Yang, T. Leong, W.-Y. Hsiao, Y.-C. Yeh, S. Fukayama, vol. 50, no. 1, pp. 37–51, 2020., , 2020

    34. Linear basis models for prediction and analysis of musical expression,, G. Widmer, M. Grachten, vol. 41, no. 4, pp. 311–322, , 2012

    35. CollageNet: Fusing arbitrary melody and accompaniment into a coherent song, C. Benetatos, C. Z. Zhiyao Duan, A. Wuerkaixi, Proceedings of the 22nd International Society for Music Information Retrieval Conference, Online, , 2021

    36. Computational analysis and modeling of expressive timing in Chopin Mazurkas, Z. Shi, Proceedings of the 22nd International Society for Music Information Retrieval Conference, Online, , 2021

    37. Computational models of expressive music performance: The state of the art,, G. Widmer, W. Goebl, vol. 33, no. 3, pp. 203–216, , 2004

    38. Principles for learning controllable TTS from annotated and latent variation, G. Henter, X. Wang, J. Yamagishi, J. Lorenzo-Trueba, Proceedings of the 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, , 2017

    39. Realtime chord recognition of musical sound: A system using Common Lisp Music, T. Fujishima, Proceedings of the 25th International Computer Music Conference, Beijing, China, , 1999

    40. Automatic melodic harmonization: An overview, challenges and future directions, I. Karydis, D. Makris, S. Sioutas, Trends in Music Information Seeking, Behavior, and Retrieval for Creativity, , 2016

    41. Learning structured output representation using deep conditional generative models,, H. Lee, X. Yan, K. Sohn, Proceedings of the 29th Conference on Neural Information Processing Systems, Montréal, Canada, , 2015

    42. Multi-view perceptron: A deep model for learning face identity and view representations, P. Luo, X. Wang, Z. Zhu, X. Tang, Proceedings of the 28th Conference on Neural Information Processing Systems, Montréal, Canada, , 2014

    43. X2Face: A network for controlling face generation by using images, audio, and pose codes, O. Wiles, A. S. Koepke, A. Zisserman, Proceedings of the European Conference on Computer Vision, Munich, Germany, , 2018

    44. Changing musical emotion: A computational rule system for modifying score and performance,, A. R. Brown, R. Muhlberger, W. F. Thompson, S. R. Livingstone, vol. 34, , 2010

    45. Deep convolutional inverse graphics network. in advances in neural information processing systems, T. D. Kulkarni, J. Tenenbaum, W. F. Whitney, P. Kohli, Proceedings of the 29th Conference on Neural Information Processing Systems, Montréal, Canada, , 2015

    46. InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets, R. Houthooft, P. Abbeel, X. Chen, Y. Duan, I. Sutskever, J. Schulman, Proceedings of the 30th Conference on Neural Information Processing Systems, Barcelona, Spain, , 2016

    47. Laminae: A stochastic modeling-based autonomous performance rendering system that elucidates performer characteristics, K. Okumura, Proceedings of the 40th International Computer Music Conference, Athens, Greece, , 2014

    48. The relationship between explicit planning and expressive performance of dynamic variations in an aural modeling task,, R. H. Woody, vol. 47, no. 4, pp. 331–342, , 1999

    49. A microcosm of musical expression: II. Quantitative analysis of pianists’ dynamics in the initial measures of Chopin’s Etude in E major,, B. H. Repp, vol. 105, pp. 1972–1988, , 1999

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼