Audio-to-symbolic Arrangement via Cross-modal Music

Audio-to-symbolic Arrangement via Cross-modal Music Representation Learning

Could we automatically derive the score of a piano accommodation based on the audio of a pop song? This is the audio-to-symbolic arrangement problem we tackle in this paper. A good arrangement model should not only consider the audio content but also have prior knowledge of piano composition (so that the generation “sounds like” the audio and meanwhile maintains musicality). To this end, we contribute a cross-modal representation-learning model, which 1) extracts chord and melodic information from the audio, and 2 ) learns texture representation from both audio and a corrupted ground truth arrangement. We further introduce a tailored training strategy that gradually shifts the source of texture information from corrupted score to audio. In the end, the score-based posterior texture is reduced to a standard normal distribution, and only audio is needed for inference. Experiments show that our model captures major audio information and outperforms baselines in generation quality.

SaveSavedRemoved 0

Audio-to-symbolic Arrangement via Cross-modal Music

Hot-Refresh Model Upgrades with Regression-Free Compatible Training

Uncertainty Modeling for Out-of-Distribution Generalization

What is Post-Quantum Cryptography?

How to Choose the Best Crypto NFT Wallet

How To Send Money Abroad And Make International Payments

How To Accept Cryptocurrency As A Business?

To Get Daily Health Newsletter