欢迎访问中国科学院大学学报,今天是
电子信息与计算机科学

面向丰富曲调要素的影视配乐生成模型

  • 赵冰爽 ,
  • 罗铁坚 ,
  • 王承杰
展开
  • 中国科学院大学计算机科学与技术学院,北京 101408
E-mail:tjluo@ucas.ac.cn

收稿日期: 2024-03-04

  修回日期: 2024-05-06

  网络出版日期: 2024-05-29

基金资助

中国科学院战略先导项目(E0421104)

BM-Transformer: a generative model for film and television soundtracks enriched with melodic elements

  • Bingshuang ZHAO ,
  • Tiejian LUO ,
  • Chengjie WANG
Expand
  • School of Computer Science and Technology,University of Chinese Academy of Sciences,Beijing 101408,China

Received date: 2024-03-04

  Revised date: 2024-05-06

  Online published: 2024-05-29

摘要

多模态模型在生成语言、视频和乐曲等任务上表现出极大的潜力,然而在面向丰富曲调要素的背景音乐生成任务上仍面临着情感一致性和专业引导等问题。本文提出多维交互引导和时间比例编码,在音乐表达复合词中嵌入情感标签、韵律密度和韵律强度,生成具有多维交互特性的音乐向量表示。设计具备视频与音乐的节奏和谐对应的影视配乐生成模型,并给出相应的非配对数据驱动的模型网络训练方法。提出的“创作者-观众”双视角评估策略,可以更全面地评估模型的交互性和配乐效果。实验结果表明,该模型在客观评价指标上提高20.6%,推断效率由0.125提升至14.370,并且在主观评价指标上也有明显优势。

本文引用格式

赵冰爽 , 罗铁坚 , 王承杰 . 面向丰富曲调要素的影视配乐生成模型[J]. 中国科学院大学学报, 2026 , 43(4) : 566 -575 . DOI: 10.7523/j.ucas.2024.040

Abstract

Multimodal models show great potential in tasks such as generating language, video, and musical scores. However, they still face problems such as emotional consistency and professional guidance for background music generation tasks oriented toward rich melodic elements. In this paper, we propose multidimensional interaction guidance and temporal scaling encoding to generate music vector representations with multidimensional interaction properties by embedding emotion labels, rhythmic density, and rhythmic intensity into music expression compound words. We design the background music transformer (BMT), a model for generating film and television soundtracks with rhythmic and harmonic correspondence between video and music, and provide a corresponding unpaired data-driven network training method for the model. We propose a “creator-audience” dual-view evaluation strategy to evaluate the interactivity and soundtrack effect of the model in a more comprehensive way. The experimental results show that the BMT model improves the objective evaluation index by 20.6%, and the model inference efficiency increases from 0.125 to 14.370, and also has obvious advantages in subjective evaluation index.

参考文献

[1] Brooks T, Peebles B, Homes C, et al. Video generation models as world simulators[EB/OL]. (2024-02-15)[2024-02-26]. .
[2] Agostinelli A, Denk T I, Borsos Z, et al. MusicLM: generating music from text[EB/OL]. arXiv 2023:2301.11325.(2023-01-26)[2024-02-26]. .
[3] Huang Q Q, Park D S, Wang T, et al. Noise2Music: text-conditioned music generation with diffusion models[EB/OL]. arXiv 2023:2302.03917. (2023-02-08)[2024-02-26]. .
[4] Copet J, Kreuk F, Gat I, et al. Simple and controllable music generation[EB/OL]. arXiv 2023:2306.05284. (2023-06-08)[2024-02-26]. .
[5] OpenAI. Introducing ChatGPT[EB/OL]. (2022-11-30)[2024-02-26]. .
[6] Kaliakatsos-Papakostas M, Floros A, Vrahatis M N. Artificial intelligence methods for music generation: a review and future perspectives[M]//Nature-Inspired Computation and Swarm Intelligence. Amsterdam: Elsevier, 2020: 217-245. DOI: 10.1016/b978-0-12-819714-1.00024-5 .
[7] Tesoriero M, Rickard N S. Music-enhanced recall: an effect of mood congruence, emotion arousal or emotion function? [J]. Musicae Scientiae201216(3): 340-356. DOI: 10.1177/1029864912459046 .
[8] Millet B, Chattah J, Ahn S. Soundtrack design: the impact of music on visual attention and affective responses[J]. Applied Ergonomics202193: 103301. DOI: 10.1016/j.apergo.2020.103301 .
[9] V?stfj?ll D. Emotion induction through music: a review of the musical mood induction procedure[J]. Musicae Scientiae20015(): 173-211. DOI: 10.1177/10298649020050S107 .
[10] Bachorik J P, Bangert M, Loui P, et al. Emotion in motion: investigating the time-course of emotional judgments of musical stimuli[J]. Music Perception200926(4): 355-364. DOI: 10.1525/mp.2009.26.4.355 .
[11] Di S Z, Jiang Z R, Liu S, et al. Video background music generation with controllable music transformer[C]//Proceedings of the 29th ACM International Conference on Multimedia. October 20 - 24, 2021, Virtual Event, China. ACM, 2021: 2037-2045. DOI: 10.1145/3474085.3475195 .
[12] Ferreira L N, Mou L L, Whitehead J, et al. Controlling perceived emotion in symbolic music generation with Monte Carlo Tree search[C]//Proceedings of the Eighteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment. 202218(1): 163-170. ACM, 2022: 163-170. DOI: 10.1609/aiide.v18i1.21960 .
[13] Gómez-Ca?ón J S, Cano E, Eerola T, et al. Music emotion recognition: toward new, robust standards in personalized and context-sensitive applications[J]. IEEE Signal Processing Magazine202138(6): 106-114. DOI: 10.1109/MSP.2021.3106232 .
[14] Wu X P. An analysis of the origin, integration and development of contemporary music composition and artificial intelligence and human-computer interaction[C]//International Conference on Human-Computer Interaction. Cham: Springer Nature Switzerland, 2023: 259-268. DOI: 10.1007/978-3-031-34609-5_19 .
[15] Briot J P. From artificial neural networks to deep learning for music generation: history, concepts and trends[J]. Neural Computing and Applications202133(1): 39-65. DOI: 10.1007/s00521-020-05399-0 .
[16] Wang L, Zhao Z Y, Liu H W, et al. A review of intelligent music generation systems[J]. Neural Computing and Applications202436(12): 6381-6401. DOI: 10.1007/s00521-024-09418-2 .
[17] Yang X T, Yu Y, Wu X Y. Double linear transformer for background music generation from videos[J]. Applied Sciences202212(10): 5050. DOI: 10.3390/app12105050 .
[18] Karlin F, Wright R. On the track: a guide to contemporary film scoring[M]. New York: Schirmer Books, 1990. DOI: 10.4324/9780203643907 .
[19] Wu S L, Yang Y H. The Jazz Transformer on the front line: exploring the shortcomings of AI-composed music through quantitative measures[EB/OL]. arXiv 2020:2008.01307. (2020-08-04)[2024-02-26]. .
[20] Hsiao W Y, Liu J Y, Yeh Y C, et al. Compound word transformer: learning to compose full-song music over dynamic directed hypergraphs[J]. Proceedings of the AAAI Conference on Artificial Intelligence202135(1): 178-186. DOI: 10.1609/aaai.v35i1.16091 .
[21] Davis A, Agrawala M. Visual rhythm and beat[J]. ACM Transactions on Graphics201837(4): 122. DOI: 10.1145/3197517.3201371 .
[22] Hernandez-Olivan C, Beltr??n J R. Music composition with deep learning: a review [M]//. Advances in Speech and Music Technology . Cham: Springer, 2023: 25-50. DOI: 10.1007/978-3-031-18444-4_2 .
[23] Huang Y S, Yang Y H. Pop music transformer: beat-based modeling and generation of expressive pop piano compositions[C]//Proceedings of the 28th ACM International Conference on Multimedia. Seattle WA USA. ACM, 2020: 1180-1188. DOI: 10.1145/3394171.3413671 .
[24] Eerola T, Toiviainen P. MIDI toolbox: MATLAB tools for music research [EB/OL]. University of Jyv?skyl?Finland.(2017-05-17)[2024-02-26]. .
[25] Hung H T, Ching J, Doh S, et al. EMOPIA: a multi-modal pop piano dataset for emotion recognition and emotion-based music generation [EB/OL]. arXiv 2021:2108.01374. (2021-08-03)[2024-02-26]. .
[26] Sulun S, Davies M E P, Viana P. Symbolic music generation conditioned on continuous-valued emotions[J]. IEEE Access202210: 44617-44626. DOI: 10.1109/ACCESS.2022.3169744 .
[27] Pesek M, Strle G, Kav?i? A, et al. The Moodo dataset: integrating user context with emotional and color perception of music for affective music information retrieval[J]. Journal of New Music Research201746(3): 246-260. DOI: 10.1080/09298215.2017.1333518 .
[28] Russell J A. A circumplex model of affect[J]. Journal of Personality and Social Psychology198039(6): 1161-1178. DOI: 10.1037/H0077714 .
[29] Katharopoulos A, Vyas A, Pappas N, et al. Transformers are RNNs: fast autoregressive transformers with linear attention[EB/OL]. arXiv 2020: 2006.16236. (2020-06-29) [2024-02-26]. .
[30] Rae J W, Potapenko A, Jayakumar S M, et al. Compressive transformers for long-range sequence modelling[EB/OL]. arXiv 2019:1911.05507. (2019-11-13)[2024-02-26]. .
[31] Dai Z H, Yang Z L, Yang Y M, et al. Transformer-XL: attentive language models beyond a fixed-length context[C]//proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy. Stroudsburg, PA, USA: Association for Computational Linguistics, 2019: 2978-2988. DOI: 10.18653/v1/p19-1285 .
[32] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need [EB/OL]. arXiv 2017:1706.03762.(2017-06-12) [2024-02-26]. .
文章导航

/