Multimodal Video-to-Music Recommendation
via Semantic Retrieval and Temporal Reranking

Seungheon Doh1,*, Minhee Lee1, Sangmoon Lee2, Ben Sangbae Chon2, Juhan Nam1

1KAIST ยท 2Gaudio Lab, Inc.

* Work completed while Seungheon was visiting Gaudio Lab.

Abstract

We present VTMR, a two-stage framework for Video-To-Music Recommendation designed to address two fundamental limitations of existing methods: reliance on unimodal video representations and the loss of temporal context from average-pooled global embedding search. To overcome these challenges, VTMR first encodes comprehensive multimodal video and music signals into a tightly aligned joint audio-visual-text representation space. We then deploy a retrieve-and-rerank pipeline: Stage~1 retrieves semantically compatible candidates at scale via global embedding search, and Stage~2 reranks them by attending over the full temporal sequences of both video and music. Evaluated on the video-to-music recommendation task, the multimodal retrieval stage improves R@10 from 14.2 to 15.9 and Median Rank from 75 to 58 over the strongest baseline; the temporal reranker further boosts R@10 to 18.3 and Median Rank to 46, demonstrating complementary gains from richer query encoding and temporal alignment. A human preference study confirms that VTMR is on par with a commercial baseline in overall preference, while outperforming a generative baseline in music quality.

Overview diagram of the VTMR video-to-music retrieval framework
Demo
Reference VTMR (Ours) VidMuse [1] Firefly [2]
VTMR (Ours) VidMuse [1] Firefly [2]
References

[1] Tian, Zeyue, et al. "Vidmuse: A simple video-to-music generation framework with long-short-term modeling." Proceedings of the Computer Vision and Pattern Recognition Conference. 2025.

[2] Adobe. "Adobe Firefly." Adobe, Accessed 26 May 2026.