
Automatic Music Transcription (AMT) is a central task within MIR, enabling various subsequent applications. Despite advancements thanks to deep learning, improving AMT remains challenging due to the scarcity of large, high-quality annotated datasets. Recognizing pitches in multi-instrument settings beyond solo piano is particularly difficult, as models struggle to generalize across domains due to dataset biases and overfitting. AMT research appears to have hit a glass ceiling, where further progress is difficult to achieve and to measure. To address this, we propose cross-version consistency---an annotation-free evaluation framework that assesses a model's transcription consistency across different recordings of the same musical work. We formalize this concept and systematically analyze its relationship with standard evaluation metrics on the AMT subtask of multi-pitch estimation. Our results show that cross-version consistency enables model assessment using only unlabeled multi-version datasets, making it particularly valuable in domains where annotated data is scarce but multi-version recordings are easy to obtain, such as orchestral music. Beyond this, our results indicate that cross-version consistency can also provide insights into a model's robustness, i. e., its ability to generalize to out-of-domain data.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
