1. Attention is all you need, A. N. Gomez, J. Uszkoreit, L. Jones, N. Parmar, A. Vaswani, N. Shazeer, L. Kaiser, I. Polosukhin, Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, United States, , 2017
2. Modelling musical dynamics, T. H. Axel Berndt, Proceedings of the 5th Audio Mostly Conference, New York, United States, , 2010
3. Counterpoint by convolution,, T. Cooijmans, C.-Z. A. Huang, A. Roberts, A. Courville, D. Eck, Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, , 2017
4. Disentangling by factorising, H. Kim, A. Mnih, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018
5. Disentangled sequential autoencoder, Y. Li, S. Mandt, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018
6. Music performance analysis: A survey, C. Arthur, A. Pati, S. Gururani, A. Lerch, Proceedings of the 20st International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019
7. Toward controlled generation of text, Z. Yang, X. Liang, R. Salakhutdinov, E. P. Xing, Z. Hu, Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, , 2017
8. DeepJ: Style-specific music generation,, T. Shin, G. W. Cottrell, H. H. Mao, Proceedings of the 12th International Conference on Semantic Computing, Laguna Hills, United States, , 2018
9. Detecting harmonic change in musical audio, C. A. Harte, M. B. Sandler, M. Gasser, Proceedings of the 1st ACM International Conference on Multimedia, Santa Barbara, United States, , 2006
10. Chord2Vec: Learning musical chord embeddings, C. Walder, S. Madjiheurem, L. Qu, Proceedings of the 30th Conference on Neural Information Processing Systems, Barcelona, Spain, , 2016
1. Attention is all you need, A. N. Gomez, J. Uszkoreit, L. Jones, N. Parmar, A. Vaswani, N. Shazeer, L. Kaiser, I. Polosukhin, Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, United States, , 2017
2. Modelling musical dynamics, T. H. Axel Berndt, Proceedings of the 5th Audio Mostly Conference, New York, United States, , 2010
3. Counterpoint by convolution,, T. Cooijmans, C.-Z. A. Huang, A. Roberts, A. Courville, D. Eck, Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, , 2017
4. Disentangling by factorising, H. Kim, A. Mnih, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018
5. Disentangled sequential autoencoder, Y. Li, S. Mandt, Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, , 2018
6. Music performance analysis: A survey, C. Arthur, A. Pati, S. Gururani, A. Lerch, Proceedings of the 20st International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019
7. Toward controlled generation of text, Z. Yang, X. Liang, R. Salakhutdinov, E. P. Xing, Z. Hu, Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, , 2017
8. DeepJ: Style-specific music generation,, T. Shin, G. W. Cottrell, H. H. Mao, Proceedings of the 12th International Conference on Semantic Computing, Laguna Hills, United States, , 2018
9. Detecting harmonic change in musical audio, C. A. Harte, M. B. Sandler, M. Gasser, Proceedings of the 1st ACM International Conference on Multimedia, Santa Barbara, United States, , 2006
10. Chord2Vec: Learning musical chord embeddings, C. Walder, S. Madjiheurem, L. Qu, Proceedings of the 30th Conference on Neural Information Processing Systems, Barcelona, Spain, , 2016
11. Self-supervised audio-visual co-segmentation, H. Zhao, A. Torralba, A. Rouditchenko, J. McDermott, H. Gan, Proceedings of the 44th IEEE International Conference on Acoustics, Speech and Signal Processing, Brighton, United Kingdom, , 2019
12. Learning a latent space of multitrack measures, I. Simon, A. Roberts, C. Raffel, D. Eck, J. Engel, C. Hawthorne, Proceedings of the 32nd Conference on Neural Information Processing Systems, Montréal, Canada, , 2018
13. Hierarchical variational autoencoders for music, J. Engel, D. Eck, A. Roberts, Proceedings of the 31st Conference on Neural Information Processing Systems, Long Beach, United States, , 2017
14. Neural speech synthesis with Transformer network, M. Liu, Y. Liu, S. Liu, S. Zhao, M. Zhou, N. Li, Proceedings of the 33rd AAAI Conference on Artificial Intelligence, Honolulu, United States, , 2019
15. Music perception, pitch, and the auditory system,, A. J. Oxenham, J. H. McDermott, vol. 18, no. 4, pp. 452–463, , 2008
16. Harmonic, melodic, and functional automatic analysis, D. Rizo, P. R. Illescas, J. M. Iñesta, Proceedings of the 33rd International Computer Music Conference, Copenhagen, Denmark, , 2007
17. A recurrent latent variable model for sequential data, K. Kastner, A. C. Courville, Y. Bengio, L. Dinh, J. Chung, K. Goel, Proceedings of the 29th Conference on Neural Information Processing Systems, Montréal, Canada, , 2015
18. Encoding musical style with Transformer autoencoders,, J. Engel, C. Hawthorne, M. Dinculescu, K. Choi, I. Simon, Proceedings of the 37th International Conference on Machine Learning, Online, , 2020
19. Schenkerian analysis by computer: A proof of concept,, A. Marsden, vol. 39, no. 3, pp. 269–289, , 2010
20. Colorization as a proxy task for visual understandings,, G. Larsson, M. Maire, G. Shakhnarovich, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Honolulu, United States, , 2017
21. Functional harmonic analysis using probabilistic models,, J. Stoddard, C. Raphael, vol. 28, no. 3, pp. 45–52, , 2004
22. From time to time: The representation of timing and tempo, H. Honing, vol. 25, , 2001
23. Chord generation from symbolic melody using BLSTM networks,, H. Lim, S. Rhyu, K. Lee, Proceedings of the 18th International Society for Music Information Retrieval Conference, Suzhou, China, , 2017
24. Deep music analogy via latent representation disentanglement, D. Wang, J. Jiang, G. Xia, Z. Wang, R. Yang, T. Chen, Proceedings of the 20th International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019
25. Melody harmonization with interpolated probabilistic models,, E. Vincent, S. Fukayama, S. A. Raczyński, vol. 42, no. 3, pp. 223–235, , 2013
26. Emotional coloring of computer-controlled music performances,, R. Bresin, A. Friberg, vol. 24, , 2000
27. MySong: Automatic accompaniment generation for vocal melodies, D. Morris, I. Simon, S. Basu, Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, Florence, Italy, , 2008
28. AI methods in algorithmic composition: A comprehensive survey,, F. Vico, J. D. Fernández, vol. 48, pp. 513–582, , 2013
29. Learning to traverse latent spaces for musical score inpainting, A. Lerch, G. Hadjeres, A. Pati, Proceedings of the 20th International Society for Music Information Retrieval Conference, Delft, The Netherlands, , 2019
30. A directional interval class representation of chord transitions, E. Cambouropoulos, Proceedings of the 12th International Conference on Music Perception and Cognition, Thessaloniki, Greece, , 2012
31. Melody lead in piano performance: Expressive device or artifact?, W. Goebl, vol. 110, , 2001
32. Unsupervised visual representation learning by context prediction, C. Doersch, A. A. Efros, A. Gupta, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Santiago, Chile, , 2015
33. Automatic melody harmonization with triad chords: A comparative study,, H.-M. Liu, H.-W. Dong, Y. Chen, T. Kitahara, B. Genchel, Y.-H. Yang, T. Leong, W.-Y. Hsiao, Y.-C. Yeh, S. Fukayama, vol. 50, no. 1, pp. 37–51, 2020., , 2020
34. Linear basis models for prediction and analysis of musical expression,, G. Widmer, M. Grachten, vol. 41, no. 4, pp. 311–322, , 2012
35. CollageNet: Fusing arbitrary melody and accompaniment into a coherent song, C. Benetatos, C. Z. Zhiyao Duan, A. Wuerkaixi, Proceedings of the 22nd International Society for Music Information Retrieval Conference, Online, , 2021
36. Computational analysis and modeling of expressive timing in Chopin Mazurkas, Z. Shi, Proceedings of the 22nd International Society for Music Information Retrieval Conference, Online, , 2021
37. Computational models of expressive music performance: The state of the art,, G. Widmer, W. Goebl, vol. 33, no. 3, pp. 203–216, , 2004
38. Principles for learning controllable TTS from annotated and latent variation, G. Henter, X. Wang, J. Yamagishi, J. Lorenzo-Trueba, Proceedings of the 18th Annual Conference of the International Speech Communication Association, Stockholm, Sweden, , 2017
39. Realtime chord recognition of musical sound: A system using Common Lisp Music, T. Fujishima, Proceedings of the 25th International Computer Music Conference, Beijing, China, , 1999
40. Automatic melodic harmonization: An overview, challenges and future directions, I. Karydis, D. Makris, S. Sioutas, Trends in Music Information Seeking, Behavior, and Retrieval for Creativity, , 2016
41. Learning structured output representation using deep conditional generative models,, H. Lee, X. Yan, K. Sohn, Proceedings of the 29th Conference on Neural Information Processing Systems, Montréal, Canada, , 2015
42. Multi-view perceptron: A deep model for learning face identity and view representations, P. Luo, X. Wang, Z. Zhu, X. Tang, Proceedings of the 28th Conference on Neural Information Processing Systems, Montréal, Canada, , 2014
43. X2Face: A network for controlling face generation by using images, audio, and pose codes, O. Wiles, A. S. Koepke, A. Zisserman, Proceedings of the European Conference on Computer Vision, Munich, Germany, , 2018
44. Changing musical emotion: A computational rule system for modifying score and performance,, A. R. Brown, R. Muhlberger, W. F. Thompson, S. R. Livingstone, vol. 34, , 2010
45. Deep convolutional inverse graphics network. in advances in neural information processing systems, T. D. Kulkarni, J. Tenenbaum, W. F. Whitney, P. Kohli, Proceedings of the 29th Conference on Neural Information Processing Systems, Montréal, Canada, , 2015
46. InfoGAN: Interpretable representation learning by information maximizing generative adversarial nets, R. Houthooft, P. Abbeel, X. Chen, Y. Duan, I. Sutskever, J. Schulman, Proceedings of the 30th Conference on Neural Information Processing Systems, Barcelona, Spain, , 2016
47. Laminae: A stochastic modeling-based autonomous performance rendering system that elucidates performer characteristics, K. Okumura, Proceedings of the 40th International Computer Music Conference, Athens, Greece, , 2014
48. The relationship between explicit planning and expressive performance of dynamic variations in an aural modeling task,, R. H. Woody, vol. 47, no. 4, pp. 331–342, , 1999
49. A microcosm of musical expression: II. Quantitative analysis of pianists’ dynamics in the initial measures of Chopin’s Etude in E major,, B. H. Repp, vol. 105, pp. 1972–1988, , 1999