OMR Datasets
Datasets are used for training model and testing. It is the food for model building. Good datasets with paird notation with high faithfulness is expensive. Correct and annotated datasets are resource heavy and expensive. Inspired by OMR datasets by Apache.
Maintained by: Alina Liu
Dataset Name | Description | Sample Image | Number of Samples | Digitalformat |
|---|---|---|---|---|
Josquin project Stanford | The Josquin Research Project (JRP) changes what it means to engage with Renaissance music. | kern, MEI, musicxml, json piano roll | ||
Tasso | Digital Edition of the Settings of Torquato Tasso's Poetry | |||
RISM works | RISM Online, a service from the RISM Digital Center provides access to many musical sources through extensive search and browse functions. Multiple entry points, with extensive Source, Person, and Institution search interfaces, provide users with the ability to quickly narrow down relevant results using several different approaches, including keyword searches, filters, and relationship information. | |||
Byrd dataset | A small dataset of 34 high quality images with individual music score pages of increasing difficulty. | |||
IMSLP Petrucci music library | Petrucci music library. Contains 280571 works, 27853 composes. Share the worlds public domain music. | |||
Apacha music score classification dataset | A python script that trains a model that can learn to distinguish between music scores and arbitrary content. A dataset of 2000 images, containing 1000 images of music scores and 1000 images of other objects including text documents. The images were taken with a smartphone camera from various angles and different lighting conditions. | |||
Apacha printed music symbols dataset | This repository contains a set of about 200 printed music symbols of 36 different classes with and without context. | |||
Accidental detection dataset | This work takes place in the context of the thesis of of Kwon-Young Choi on the subject of resolving segmentation problems in dense and damaged printed piano scores from the 1750 to 1950 period. This data set contains 2955 images with annotations for the task of detecting or rejecting the correct accidental for a given note head. | |||
International audio lab Erlangen | Tools for semi automatic bounding box annotation of musical measures | |||
Polish digital scores | ||||
Polish heritage open access | The Open Access project will expand the informational infrastructure in place at The Chopin Institute, which was designed as part of the POPC project ‘Chopin Heritage in Open Access’ (DCOD). The main aim is to broaden the portal to the most important and representative works of Polish music from the 16th to the 19th centuries, allowing searches according to musical elements and metadata. | |||
CMME datasets | Mensural music database and study project | |||
Sheet Music Benchmark | SMB (Sheet Music Benchmark) is a dataset of printed Common Western Modern Notation scores developed at the University of Alicante at the Pattern Recognition and Artificial Intelligence Group. | |||
Musicalion | Membership music sheet download. | pdf, capella, midi, musicxml | ||
OpenOMR dataset | A dataset of 706 symbols (g-clef, f-clef) and symbol primitives (note-heads, stems with flags, beams) of 16 classes created by Arnaud F. Desaedeleer as part of his master thesis to train artificial neural networks.D | |||
Rebelo | Three datasets of perfect and scanned music symbols including an extensive set of synthetically modified images for staff-line detection and removal. Contains approximately 15000 music symbols. | |||
SEILS dataset | The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, MEI, **mens, agnostic, semantic and pdf formats, in white mensural and modern notation. The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, **mens, MEI, agnostic, semantic, and pdf formats, in white mensural and modern notation. | XML, MIDI, Finale, kern, MEL, | ||
Capitan collection | A corpus collected by an electronic pen while tracing isolated music symbols from Early manuscripts. | |||
MK free sheet music | A large collection of piano, voilin, vocal, trumpet, etc. | |||
Olympic dataset | Create a dataset of synthetic and scanned pianoform music for end-to-end OMR, called OLiMPiC. The dataset is built upon the OpenScore Lieder Corpus. | LMX format (linearized MusicXML) | ||
Quartets sheet transformer | Generated together with GrandStaff by Alicante group. Quartets is a well-known collection employed in the Audio to Score field. | |||
Mozart scores Mozarteum by DIME | Printed score and MEI download of 30+ works. | |||
Humdrum Polish digital scores | Songs and music from Poland, including Chopin. | |||
RISM | Reportoire International de Sources Musicales contains catalogged music sources funded by barious agencies. It includes libraries in Berlin and Bavaria. | 1617184 | ||
PDMX | Public domain muxicxml dataset for symbolic music processing. | |||
MUSCIMA++ | A subset of MUSCIMA that is fully annotated. It used MuNG music notation graph format. | mung | ||
Musicxml collection | A collection of music library with musicxml | musicxml | ||
MusicXML Sample | Small sample of music in musicXML format | musicxml | ||
MSMD Multimodal | MSMD is a synthetic dataset of 497 pieces of (classical) music that contains both audio and score representations of the pieces aligned at a fine-grained level. Music of Bach, Chopin, Froberger, and others in full page (typically 1) in pdf format. | midi, ly, yml, npy, xml, mung | ||
Primus | Printed Images of Music Staves. Two packages of synthetic music with one line segments printed. No musicxml and kern. | mid, mei, pae, agnostic | ||
MUSCIMA | Handwritten music score images. | |||
DoReMe | Part of Dorico project, most pictures are one line long and one staff long. File in png format | 6432 | MEI, MIDI, XML | |
Open Score Lieder | Collection of 1200 10th century songs annotated by a team of volunteers. | 1200 | mxl, mscz, mscx, | |
Humdrum Polish scores | This repository contains transcriptions from the Heritage of Polish Music in Open Access project at the Chopin Institute in the Humdrum digital score format. Files are organized by the RISM siglum ID of the source archive. Here is a promotional booklet about the project in Polish and English. | 79 | kern | |
DeepScore | Page long or less png files with annotation in json. A tool set accompanies. | 150 | ||
GrandStaff | Music segments from Beethoven, Chopin, Hummel, Juplin, Mozart, and Scarlatti. From the paper Sheet Music Transformer. | 70000 | kern, bekrn, | |
ClassicComposers | Lots of classic music sheet and kern/MEI files. | 6497 | kern, MEI | |
OMRT | OMR datasets for testing OMR tools comprehensively. | 330000 | ||
OMR datasets | A collection of collections. A collection of images from various sources of datasets. Some no longer work. | 891 |

Josquin project Stanford
The Josquin Research Project (JRP) changes what it means to engage with Renaissance music.

Tasso
Digital Edition of the Settings of Torquato Tasso's Poetry

RISM works
RISM Online, a service from the RISM Digital Center provides access to many musical sources through extensive search and browse functions. Multiple entry points, with extensive Source, Person, and Institution search interfaces, provide users with the ability to quickly narrow down relevant results using several different approaches, including keyword searches, filters, and relationship information.

Byrd dataset
A small dataset of 34 high quality images with individual music score pages of increasing difficulty.

IMSLP Petrucci music library
Petrucci music library. Contains 280571 works, 27853 composes. Share the worlds public domain music.

Apacha music score classification dataset
A python script that trains a model that can learn to distinguish between music scores and arbitrary content. A dataset of 2000 images, containing 1000 images of music scores and 1000 images of other objects including text documents. The images were taken with a smartphone camera from various angles and different lighting conditions.

Apacha printed music symbols dataset
This repository contains a set of about 200 printed music symbols of 36 different classes with and without context.

Accidental detection dataset
This work takes place in the context of the thesis of of Kwon-Young Choi on the subject of resolving segmentation problems in dense and damaged printed piano scores from the 1750 to 1950 period. This data set contains 2955 images with annotations for the task of detecting or rejecting the correct accidental for a given note head.

International audio lab Erlangen
Tools for semi automatic bounding box annotation of musical measures

Polish digital scores

Polish heritage open access
The Open Access project will expand the informational infrastructure in place at The Chopin Institute, which was designed as part of the POPC project ‘Chopin Heritage in Open Access’ (DCOD). The main aim is to broaden the portal to the most important and representative works of Polish music from the 16th to the 19th centuries, allowing searches according to musical elements and metadata.

CMME datasets
Mensural music database and study project

Sheet Music Benchmark
SMB (Sheet Music Benchmark) is a dataset of printed Common Western Modern Notation scores developed at the University of Alicante at the Pattern Recognition and Artificial Intelligence Group.

Musicalion
Membership music sheet download.

OpenOMR dataset
A dataset of 706 symbols (g-clef, f-clef) and symbol primitives (note-heads, stems with flags, beams) of 16 classes created by Arnaud F. Desaedeleer as part of his master thesis to train artificial neural networks.D

Rebelo
Three datasets of perfect and scanned music symbols including an extensive set of synthetically modified images for staff-line detection and removal. Contains approximately 15000 music symbols.

SEILS dataset
The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, MEI, **mens, agnostic, semantic and pdf formats, in white mensural and modern notation. The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, **mens, MEI, agnostic, semantic, and pdf formats, in white mensural and modern notation.

Capitan collection
A corpus collected by an electronic pen while tracing isolated music symbols from Early manuscripts.

MK free sheet music
A large collection of piano, voilin, vocal, trumpet, etc.

Olympic dataset
Create a dataset of synthetic and scanned pianoform music for end-to-end OMR, called OLiMPiC. The dataset is built upon the OpenScore Lieder Corpus.

Quartets sheet transformer
Generated together with GrandStaff by Alicante group. Quartets is a well-known collection employed in the Audio to Score field.

Mozart scores Mozarteum by DIME
Printed score and MEI download of 30+ works.

Humdrum Polish digital scores
Songs and music from Poland, including Chopin.

RISM
Reportoire International de Sources Musicales contains catalogged music sources funded by barious agencies. It includes libraries in Berlin and Bavaria.

PDMX
Public domain muxicxml dataset for symbolic music processing.

MUSCIMA++
A subset of MUSCIMA that is fully annotated. It used MuNG music notation graph format.

Musicxml collection
A collection of music library with musicxml

MusicXML Sample
Small sample of music in musicXML format

MSMD Multimodal
MSMD is a synthetic dataset of 497 pieces of (classical) music that contains both audio and score representations of the pieces aligned at a fine-grained level. Music of Bach, Chopin, Froberger, and others in full page (typically 1) in pdf format.

Primus
Printed Images of Music Staves. Two packages of synthetic music with one line segments printed. No musicxml and kern.

MUSCIMA
Handwritten music score images.

DoReMe
Part of Dorico project, most pictures are one line long and one staff long. File in png format

Open Score Lieder
Collection of 1200 10th century songs annotated by a team of volunteers.

Humdrum Polish scores
This repository contains transcriptions from the Heritage of Polish Music in Open Access project at the Chopin Institute in the Humdrum digital score format. Files are organized by the RISM siglum ID of the source archive. Here is a promotional booklet about the project in Polish and English.

DeepScore
Page long or less png files with annotation in json. A tool set accompanies.

GrandStaff
Music segments from Beethoven, Chopin, Hummel, Juplin, Mozart, and Scarlatti. From the paper Sheet Music Transformer.

ClassicComposers
Lots of classic music sheet and kern/MEI files.

OMRT
OMR datasets for testing OMR tools comprehensively.

OMR datasets
A collection of collections. A collection of images from various sources of datasets. Some no longer work.