Automated Music Transcription
We are developing an approach for the transcription of MIDI recordings of performances into digital music scores, based on an abstract intermediate representation of music notation as trees, and on advanced quantitative parsing algorithm. It is implemented into the qparse library, and experiments have shown very good results on challenging datasets.
Music transcription is the conversion of a performance into a score. When the performance is given in discrete, sequential form (a series of notes with exact pitches but approximate dates and durations called a piano roll), it is referred to as MIDI-to-score (M2S) transcription. This case is particularly interesting in the context of score editing, as the piano roll representation corresponds to the MIDI format output by electronic musical keyboards. It is also the target format for music transcription tools that process audio signals (known as Audio-to-MIDI transcription A2M), with which M2S approaches are therefore complementary, with a view to global transcription (Audio-to-Score).
We are developing an approach based on formal languages for M2S transcription, relying on a hierarchical representation of musical notation similar to rhythm trees (see these articles presented at MCM’15, MEC’15, TENOR’17 or in the IFIP WG 1.6 on term rewriting), following the principle of proportional definition of durations in music (1 half note = 2 quarter notes, 1 quarter note = 2 eighth notes, etc.). A weighted tree automaton describes an a priori notation model, associating each rhythm tree with a notational complexity value in a commutative semi-ring. This automaton can be learned on a corpus of scores using grammatical inference techniques (SMC 2019). Given a sequential, unstructured performance as input, a so-called 1-best-parsing algorithm searches for a parsing tree that minimizes both its complexity value in the a priori language and the deviation from the corresponding rhythm performance. From this parse tree, an intermediate representation (IR) is constructed in the TSM model that we are developing. After a number of post-processing steps, in particular by rewriting, the IR is converted into a musical score in XML encoding.
Development
The development in C++ of the transcription procedure briefly described above (library qparse)
began at Inria Paris in 2017 and continued in collaboration with Philippe Rigaux’s team at Cedric/CNAM and Masahiko Sakai’s team at Nagoya University (Yamaha grant and JSPS project with JAIST).
A previous approach, preliminary to qparse in terms of efficiency and functionality, was developed with Adrien Ycart and Jean Bresson at Ircam in 2015-16 (LISP library RQ for OpenMusic). RQ is oriented towards composition assistance, offering an interactive approach where the user is presented with transcription solutions listed in an orderly manner, thanks to a graphical interface integrated into OpenMusic. The development of RQ began during Adrien Ycart’s Master’s research in 2015 at Ircam, with Ycart doing most of the development, co-supervised by Jean Bresson (lead developer of OpenMusic), and continued with an internal project (UPI) at Ircam on the subject.
Originality and difficulty
Transcription is one of the oldest and most difficult challenges in computer music, and has been the subject of much research in Music Information Retrieval (MIR). To date, no completely satisfactory solution has been found.
The main originality of our approach to M2S transcription is the direct production of structured musical scores (in XML encodings). Most other approaches (particularly signal-based approaches that take audio files as input) produce sequences of MIDI events with quantized timestamps (or non-quantized timestamps in the case of audio-to-MIDI), with the structuring of the output score being delegated to a mainstream score editor (Musecore or Finale), which often requires significant manual corrections.
Our approach uses parsing (converting an unstructured sequence into a tree structure), processing the two steps above—quantifying durations and structuring notation—in tandem, avoiding error propagation, for more relevant results, particularly in the case of complex rhythms. Another advantage of parsing is that it establishes global relationships between arbitrarily distant events, typically linked to the metric structure. This long-term view is lacking in sequential statistical models.
This approach relies in particular on a tree-structured abstract intermediate representation of music notation, and on tool from formal language theory (weighted tree automata) to grasp the deep structure of music notation. To summarize, roughly, qparse transcribes by parsing the sequential input against an a priori hierarchical language model, describing a preferred style of music notation.
Validation
Evaluations of qparse were performed on a database of scores consisting of monophonic excerpts (one note at a time) from the classical repertoire, with varying rhythmic complexity. They demonstrated fast processing times (around 100 ms for a one-page score) and good results, particularly in the case of complex rhythms, grace notes, and rests, characteristics that tend to cause difficulties for other systems, especially commercial ones.
It was also used for the transcription into music scores in the MEI encoding of human performances on electronic drum kits, captured in MIDI files, of Magenta’s Groove MIDI Dataset. All the 440 beats files of this dataset were successfully transcribed, without pre-training or manual adjustments, covering various genres, time signatures, tempi and of length up to 260 measures.
Dissemination
RQ has been evaluated by composers, taught in composition courses at Ircam, and used for writing pieces, such as this creation by Alessandro Ratoci for the 2016 Ravenna Festival.
qparse is currently a set of command-line prototypes. The ultimate goal is to integrate it as a plugin into music notation software (Sibelius, Dorico, or the free software MuseScore, etc.) or a DAW.
