DistantNews

Sogang University Team's Research on Musical Modality Translation Published in Prestigious IEEE Journal

From Hankyoreh · (43m ago) Korean Positive tone

Translated from Korean, summarized and contextualized by DistantNews.

TLDR

  • A research team led by Professor Jeong Da-saem from Sogang University's Department of Art & Technology has published a paper in the prestigious international journal IEEE TASLPRO.
  • The paper proposes a unified framework, 'U-MusT,' capable of simultaneously learning translation tasks across various musical modalities like score images, symbolic music, and performance audio.
  • The research achieved the lowest symbol error rate in piano score recognition and can generate expressive performance audio directly from score images, with the team also releasing a large dataset to contribute to the music information retrieval field.

Sogang University's Department of Art & Technology, under the leadership of Professor Jeong Da-saem, has achieved a significant milestone in the field of signal processing. Their collaborative research, involving scholars from Seoul National University and Carnegie Mellon University, has resulted in a groundbreaking paper published in the esteemed IEEE Transactions on Audio, Speech and Language Processing (TASLPRO).

The paper, titled 'U-MusT: A Unified Framework for Cross-modal Translation of Score Images, Symbolic Music, and Performance Audio,' introduces a novel, universal model. This model is designed to learn translation tasks between diverse musical modalities, such as score images, symbolic music notation, and performance audio, all at once. This represents a significant advancement over previous studies that focused on specialized models for individual translation tasks.

This research is particularly noteworthy for its practical achievements. The proposed model has attained the lowest symbol error rate to date in piano score recognition. Furthermore, it is the first model globally to generate expressive performance audio directly from score images without intermediate steps. To foster further research in Music Information Retrieval (MIR), the team has generously released a dataset comprising over 1300 hours of paired score image-performance audio data.

This work will be presented at ICASSP 2026, the world's largest signal processing conference, highlighting its international recognition and potential impact. The contribution from Sogang University underscores South Korea's growing prowess in cutting-edge technological research and its commitment to advancing global academic fields.

DistantNews Editorial

Originally published by Hankyoreh in Korean. Translated, summarized, and contextualized by our editorial team with added local perspective. Read our editorial standards.