Welcome to Journal of University of Chinese Academy of Sciences,Today is

Journal of University of Chinese Academy of Sciences ›› 2026, Vol. 43 ›› Issue (4): 566-575.DOI: 10.7523/j.ucas.2024.040

• Electronics and Computer Science • Previous Articles    

BM-Transformer: a generative model for film and television soundtracks enriched with melodic elements

Bingshuang ZHAO, Tiejian LUO(), Chengjie WANG   

  1. School of Computer Science and Technology,University of Chinese Academy of Sciences,Beijing 101408,China
  • Received:2024-03-04 Revised:2024-05-06 Online:2026-07-15
  • Contact: Tiejian LUO

Abstract:

Multimodal models show great potential in tasks such as generating language, video, and musical scores. However, they still face problems such as emotional consistency and professional guidance for background music generation tasks oriented toward rich melodic elements. In this paper, we propose multidimensional interaction guidance and temporal scaling encoding to generate music vector representations with multidimensional interaction properties by embedding emotion labels, rhythmic density, and rhythmic intensity into music expression compound words. We design the background music transformer (BMT), a model for generating film and television soundtracks with rhythmic and harmonic correspondence between video and music, and provide a corresponding unpaired data-driven network training method for the model. We propose a “creator-audience” dual-view evaluation strategy to evaluate the interactivity and soundtrack effect of the model in a more comprehensive way. The experimental results show that the BMT model improves the objective evaluation index by 20.6%, and the model inference efficiency increases from 0.125 to 14.370, and also has obvious advantages in subjective evaluation index.

Key words: deep neural network, emotional coherence, video soundtrack generation model, multidimensional interaction guidance

CLC Number: