Local Multimodal Music Alignment from Global Supervision

Irmak Bukey1, Zachary Novack2, Jongmin Jung3, Dasaem Jeong4, Chris Donahue1
1 Carnegie Mellon University    2 University of California San Diego    3 Neutune    4 Sogang University

Accepted at ISMIR 2026

Abstract

Introducing FuSiLi: Fused Sinkhorn-Localized Similarity

Understanding music requires understanding localized relationships across data modalities, e.g., how time in performance audio maps onto position in a score image. Yet supervision for such local correspondences is difficult to obtain—in practice, we often only have access to coarser global supervision like paired segments of audio and images. To address this gap, we propose FuSiLi (Fused Sinkhorn-Localized Similarity), a similarity score for multimodal contrastive learning operating directly on local image patch and audio frame features via Sinkhorn-based soft alignment. We show that FuSiLi (i) effectively learns local relationships, (ii) requires only global supervision, and (iii) retains the global alignment capabilities of conventional contrastive approaches. We fine-tune pretrained CLIP and CLAP encoders on pairs of raw sheet music images and audio using a hybrid contrastive objective combining FuSiLi with conventional global similarity. We evaluate on cross-modal retrieval and frame-level alignment tasks against a range of global and local baselines, showing that our approach outperforms them on local alignment while remaining competitive on retrieval.

Method Overview

Pipeline of standard contrastive learning vs. FuSiLi. Standard approaches compute batch-wise similarity using pre-pooled representations, removing any capacity for explicit cross-modal locality. In contrast, our fuse-then-pool FuSiLi computes frame-wise similarity matrices (frame similarity Sij) for each batch pairing. We then apply Sinkhorn iteration to enforce soft one-to-one correspondences at the frame level. Only after this alignment step do we pool representations into global embeddings, which now encode localized cross-modal structure in addition to global semantics.

Qualitative Results

We evaluate our best-performing configuration (FuSiLi, Same Piece + Pos. Mutations) qualitatively on frame-level cross-modal alignment between score images and audio performances. We visualize retrieved image patches for each audio frame to assess local correspondence quality under both in-domain and out-of-domain settings, and compare it against a standard global contrastive baseline trained with the same configuration (Baseline, Same Piece + Pos. Mutations) to highlight localized alignments learned by FuSiLi. We also include the alignment matrices produced for each example.

In-Domain Alignment (Top-1 Image Patch Retrieval)

On the MSMD dataset [1], we visualize top-1 image patch retrieval for each audio frame.

Example 1

FuSiLi
Global Baseline
Ground Truth

Example 2

FuSiLi
Global Baseline
Ground Truth

Example 3

FuSiLi
Global Baseline
Ground Truth

Out-of-Domain Alignment (Top-3 Image Patch Retrieval)

We further evaluate robustness under domain shift using the YTSV dataset [2]. In this setting, we visualize the top-3 most similar image patches per audio frame, reflecting increased ambiguity in cross-domain alignment.

Example 1

FuSiLi
Global Baseline

Example 2

FuSiLi
Global Baseline

Quantitative Summary

Local Eval Point & Retrieve Global (PDMX[3]) Global (YTSV[2]) Global (MSMD[1])
Method Top-1PPL I2AA2I I2A R@1I2A MRR A2I R@1A2I MRR I2A R@1I2A MRR A2I R@1A2I MRR I2A R@1I2A MRR A2I R@1A2I MRR
Baseline 0.2436.55 0.010.05 0.730.83 0.730.83 0.070.12 0.030.07 0.270.38 0.280.41
FuSiLi 0.3034.24 0.140.14 0.720.83 0.730.83 0.050.10 0.040.08 0.250.38 0.250.39

Bolded values correspond to the task visualized in the videos above.

[1] Matthias Dorfer, Jan Hajič jr., Andreas Arzt, Harald Frostel, Gerhard Widmer. Learning Audio-Sheet Music Correspondences for Cross-Modal Retrieval and Piece Identification. Transactions of the International Society for Music Information Retrieval, issue 1, 2018.

[3] Jongmin Jung, Dongmin Kim, Sihun Lee, Seola Cho, Hyungjoon So, Irmak Bukey, Chris Donahue, Dasaem Jeong.
U-MusT: A Unified Framework for Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio.
IEEE Transactions on Audio, Speech and Language Processing, 2025. IEEE.

[2] Phillip Long, Zachary Novack, Taylor Berg-Kirkpatrick, Julian McAuley. PDMX: A Large-Scale Public Domain MusicXML Dataset for Symbolic Music Processing. ICASSP 2025 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP), 2025. IEEE.