Welcome to Journal of University of Chinese Academy of Sciences,Today is
Electronics and Computer Science

BM-Transformer: a generative model for film and television soundtracks enriched with melodic elements

  • Bingshuang ZHAO ,
  • Tiejian LUO ,
  • Chengjie WANG
Expand
  • School of Computer Science and Technology,University of Chinese Academy of Sciences,Beijing 101408,China

Received date: 2024-03-04

  Revised date: 2024-05-06

  Online published: 2024-05-29

Abstract

Multimodal models show great potential in tasks such as generating language, video, and musical scores. However, they still face problems such as emotional consistency and professional guidance for background music generation tasks oriented toward rich melodic elements. In this paper, we propose multidimensional interaction guidance and temporal scaling encoding to generate music vector representations with multidimensional interaction properties by embedding emotion labels, rhythmic density, and rhythmic intensity into music expression compound words. We design the background music transformer (BMT), a model for generating film and television soundtracks with rhythmic and harmonic correspondence between video and music, and provide a corresponding unpaired data-driven network training method for the model. We propose a “creator-audience” dual-view evaluation strategy to evaluate the interactivity and soundtrack effect of the model in a more comprehensive way. The experimental results show that the BMT model improves the objective evaluation index by 20.6%, and the model inference efficiency increases from 0.125 to 14.370, and also has obvious advantages in subjective evaluation index.

Cite this article

Bingshuang ZHAO , Tiejian LUO , Chengjie WANG . BM-Transformer: a generative model for film and television soundtracks enriched with melodic elements[J]. Journal of University of Chinese Academy of Sciences, 2026 , 43(4) : 566 -575 . DOI: 10.7523/j.ucas.2024.040

References

[1] Brooks T, Peebles B, Homes C, et al. Video generation models as world simulators[EB/OL]. (2024-02-15)[2024-02-26]. .
[2] Agostinelli A, Denk T I, Borsos Z, et al. MusicLM: generating music from text[EB/OL]. arXiv 2023:2301.11325.(2023-01-26)[2024-02-26]. .
[3] Huang Q Q, Park D S, Wang T, et al. Noise2Music: text-conditioned music generation with diffusion models[EB/OL]. arXiv 2023:2302.03917. (2023-02-08)[2024-02-26]. .
[4] Copet J, Kreuk F, Gat I, et al. Simple and controllable music generation[EB/OL]. arXiv 2023:2306.05284. (2023-06-08)[2024-02-26]. .
[5] OpenAI. Introducing ChatGPT[EB/OL]. (2022-11-30)[2024-02-26]. .
[6] Kaliakatsos-Papakostas M, Floros A, Vrahatis M N. Artificial intelligence methods for music generation: a review and future perspectives[M]//Nature-Inspired Computation and Swarm Intelligence. Amsterdam: Elsevier, 2020: 217-245. DOI: 10.1016/b978-0-12-819714-1.00024-5 .
[7] Tesoriero M, Rickard N S. Music-enhanced recall: an effect of mood congruence, emotion arousal or emotion function? [J]. Musicae Scientiae201216(3): 340-356. DOI: 10.1177/1029864912459046 .
[8] Millet B, Chattah J, Ahn S. Soundtrack design: the impact of music on visual attention and affective responses[J]. Applied Ergonomics202193: 103301. DOI: 10.1016/j.apergo.2020.103301 .
[9] V?stfj?ll D. Emotion induction through music: a review of the musical mood induction procedure[J]. Musicae Scientiae20015(): 173-211. DOI: 10.1177/10298649020050S107 .
[10] Bachorik J P, Bangert M, Loui P, et al. Emotion in motion: investigating the time-course of emotional judgments of musical stimuli[J]. Music Perception200926(4): 355-364. DOI: 10.1525/mp.2009.26.4.355 .
[11] Di S Z, Jiang Z R, Liu S, et al. Video background music generation with controllable music transformer[C]//Proceedings of the 29th ACM International Conference on Multimedia. October 20 - 24, 2021, Virtual Event, China. ACM, 2021: 2037-2045. DOI: 10.1145/3474085.3475195 .
[12] Ferreira L N, Mou L L, Whitehead J, et al. Controlling perceived emotion in symbolic music generation with Monte Carlo Tree search[C]//Proceedings of the Eighteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment. 202218(1): 163-170. ACM, 2022: 163-170. DOI: 10.1609/aiide.v18i1.21960 .
[13] Gómez-Ca?ón J S, Cano E, Eerola T, et al. Music emotion recognition: toward new, robust standards in personalized and context-sensitive applications[J]. IEEE Signal Processing Magazine202138(6): 106-114. DOI: 10.1109/MSP.2021.3106232 .
[14] Wu X P. An analysis of the origin, integration and development of contemporary music composition and artificial intelligence and human-computer interaction[C]//International Conference on Human-Computer Interaction. Cham: Springer Nature Switzerland, 2023: 259-268. DOI: 10.1007/978-3-031-34609-5_19 .
[15] Briot J P. From artificial neural networks to deep learning for music generation: history, concepts and trends[J]. Neural Computing and Applications202133(1): 39-65. DOI: 10.1007/s00521-020-05399-0 .
[16] Wang L, Zhao Z Y, Liu H W, et al. A review of intelligent music generation systems[J]. Neural Computing and Applications202436(12): 6381-6401. DOI: 10.1007/s00521-024-09418-2 .
[17] Yang X T, Yu Y, Wu X Y. Double linear transformer for background music generation from videos[J]. Applied Sciences202212(10): 5050. DOI: 10.3390/app12105050 .
[18] Karlin F, Wright R. On the track: a guide to contemporary film scoring[M]. New York: Schirmer Books, 1990. DOI: 10.4324/9780203643907 .
[19] Wu S L, Yang Y H. The Jazz Transformer on the front line: exploring the shortcomings of AI-composed music through quantitative measures[EB/OL]. arXiv 2020:2008.01307. (2020-08-04)[2024-02-26]. .
[20] Hsiao W Y, Liu J Y, Yeh Y C, et al. Compound word transformer: learning to compose full-song music over dynamic directed hypergraphs[J]. Proceedings of the AAAI Conference on Artificial Intelligence202135(1): 178-186. DOI: 10.1609/aaai.v35i1.16091 .
[21] Davis A, Agrawala M. Visual rhythm and beat[J]. ACM Transactions on Graphics201837(4): 122. DOI: 10.1145/3197517.3201371 .
[22] Hernandez-Olivan C, Beltr??n J R. Music composition with deep learning: a review [M]//. Advances in Speech and Music Technology . Cham: Springer, 2023: 25-50. DOI: 10.1007/978-3-031-18444-4_2 .
[23] Huang Y S, Yang Y H. Pop music transformer: beat-based modeling and generation of expressive pop piano compositions[C]//Proceedings of the 28th ACM International Conference on Multimedia. Seattle WA USA. ACM, 2020: 1180-1188. DOI: 10.1145/3394171.3413671 .
[24] Eerola T, Toiviainen P. MIDI toolbox: MATLAB tools for music research [EB/OL]. University of Jyv?skyl?Finland.(2017-05-17)[2024-02-26]. .
[25] Hung H T, Ching J, Doh S, et al. EMOPIA: a multi-modal pop piano dataset for emotion recognition and emotion-based music generation [EB/OL]. arXiv 2021:2108.01374. (2021-08-03)[2024-02-26]. .
[26] Sulun S, Davies M E P, Viana P. Symbolic music generation conditioned on continuous-valued emotions[J]. IEEE Access202210: 44617-44626. DOI: 10.1109/ACCESS.2022.3169744 .
[27] Pesek M, Strle G, Kav?i? A, et al. The Moodo dataset: integrating user context with emotional and color perception of music for affective music information retrieval[J]. Journal of New Music Research201746(3): 246-260. DOI: 10.1080/09298215.2017.1333518 .
[28] Russell J A. A circumplex model of affect[J]. Journal of Personality and Social Psychology198039(6): 1161-1178. DOI: 10.1037/H0077714 .
[29] Katharopoulos A, Vyas A, Pappas N, et al. Transformers are RNNs: fast autoregressive transformers with linear attention[EB/OL]. arXiv 2020: 2006.16236. (2020-06-29) [2024-02-26]. .
[30] Rae J W, Potapenko A, Jayakumar S M, et al. Compressive transformers for long-range sequence modelling[EB/OL]. arXiv 2019:1911.05507. (2019-11-13)[2024-02-26]. .
[31] Dai Z H, Yang Z L, Yang Y M, et al. Transformer-XL: attentive language models beyond a fixed-length context[C]//proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Florence, Italy. Stroudsburg, PA, USA: Association for Computational Linguistics, 2019: 2978-2988. DOI: 10.18653/v1/p19-1285 .
[32] Vaswani A, Shazeer N, Parmar N, et al. Attention is all you need [EB/OL]. arXiv 2017:1706.03762.(2017-06-12) [2024-02-26]. .
Outlines

/