top of page

OMR Datasets

Datasets are used for training model and testing.  It is the food for model building.  Good datasets with paird notation with high faithfulness is expensive.  Correct and annotated datasets are resource heavy and expensive.  Inspired by OMR datasets by Apache.

Maintained by: Alina Liu

Dataset Name
Description
Sample Image
Number of Samples
Digitalformat
Josquin project Stanford
The Josquin Research Project (JRP) changes what it means to engage with Renaissance music.
kern, MEI, musicxml, json piano roll
Tasso
Digital Edition of the Settings of Torquato Tasso's Poetry
RISM works
RISM Online, a service from the RISM Digital Center provides access to many musical sources through extensive search and browse functions. Multiple entry points, with extensive Source, Person, and Institution search interfaces, provide users with the ability to quickly narrow down relevant results using several different approaches, including keyword searches, filters, and relationship information.
Byrd dataset
A small dataset of 34 high quality images with individual music score pages of increasing difficulty.
IMSLP Petrucci music library
Petrucci music library. Contains 280571 works, 27853 composes. Share the worlds public domain music.
Apacha music score classification dataset
A python script that trains a model that can learn to distinguish between music scores and arbitrary content. A dataset of 2000 images, containing 1000 images of music scores and 1000 images of other objects including text documents. The images were taken with a smartphone camera from various angles and different lighting conditions.
Apacha printed music symbols dataset
This repository contains a set of about 200 printed music symbols of 36 different classes with and without context.
Accidental detection dataset
This work takes place in the context of the thesis of of Kwon-Young Choi on the subject of resolving segmentation problems in dense and damaged printed piano scores from the 1750 to 1950 period. This data set contains 2955 images with annotations for the task of detecting or rejecting the correct accidental for a given note head.
International audio lab Erlangen
Tools for semi automatic bounding box annotation of musical measures
Polish digital scores
Polish heritage open access
The Open Access project will expand the informational infrastructure in place at The Chopin Institute, which was designed as part of the POPC project ‘Chopin Heritage in Open Access’ (DCOD). The main aim is to broaden the portal to the most important and representative works of Polish music from the 16th to the 19th centuries, allowing searches according to musical elements and metadata.
CMME datasets
Mensural music database and study project
Sheet Music Benchmark
SMB (Sheet Music Benchmark) is a dataset of printed Common Western Modern Notation scores developed at the University of Alicante at the Pattern Recognition and Artificial Intelligence Group.
Musicalion
Membership music sheet download.
pdf, capella, midi, musicxml
OpenOMR dataset
A dataset of 706 symbols (g-clef, f-clef) and symbol primitives (note-heads, stems with flags, beams) of 16 classes created by Arnaud F. Desaedeleer as part of his master thesis to train artificial neural networks.D
Rebelo
Three datasets of perfect and scanned music symbols including an extensive set of synthetically modified images for staff-line detection and removal. Contains approximately 15000 music symbols.
SEILS dataset
The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, MEI, **mens, agnostic, semantic and pdf formats, in white mensural and modern notation. The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, **mens, MEI, agnostic, semantic, and pdf formats, in white mensural and modern notation.
XML, MIDI, Finale, kern, MEL,
Capitan collection
A corpus collected by an electronic pen while tracing isolated music symbols from Early manuscripts.
MK free sheet music
A large collection of piano, voilin, vocal, trumpet, etc.
Olympic dataset
Create a dataset of synthetic and scanned pianoform music for end-to-end OMR, called OLiMPiC. The dataset is built upon the OpenScore Lieder Corpus.
LMX format (linearized MusicXML)
Quartets sheet transformer
Generated together with GrandStaff by Alicante group. Quartets is a well-known collection employed in the Audio to Score field.
Mozart scores Mozarteum by DIME
Printed score and MEI download of 30+ works.
Humdrum Polish digital scores
Songs and music from Poland, including Chopin.
RISM
Reportoire International de Sources Musicales contains catalogged music sources funded by barious agencies. It includes libraries in Berlin and Bavaria.
1617184
PDMX
Public domain muxicxml dataset for symbolic music processing.
MUSCIMA++
A subset of MUSCIMA that is fully annotated. It used MuNG music notation graph format.
mung
Musicxml collection
A collection of music library with musicxml
musicxml
MusicXML Sample
Small sample of music in musicXML format
musicxml
MSMD Multimodal
MSMD is a synthetic dataset of 497 pieces of (classical) music that contains both audio and score representations of the pieces aligned at a fine-grained level. Music of Bach, Chopin, Froberger, and others in full page (typically 1) in pdf format.
midi, ly, yml, npy, xml, mung
Primus
Printed Images of Music Staves. Two packages of synthetic music with one line segments printed. No musicxml and kern.
mid, mei, pae, agnostic
MUSCIMA
Handwritten music score images.
DoReMe
Part of Dorico project, most pictures are one line long and one staff long. File in png format
6432
MEI, MIDI, XML
Open Score Lieder
Collection of 1200 10th century songs annotated by a team of volunteers.
1200
mxl, mscz, mscx,
Humdrum Polish scores
This repository contains transcriptions from the Heritage of Polish Music in Open Access project at the Chopin Institute in the Humdrum digital score format. Files are organized by the RISM siglum ID of the source archive. Here is a promotional booklet about the project in Polish and English.
79
kern
DeepScore
Page long or less png files with annotation in json. A tool set accompanies.
150
GrandStaff
Music segments from Beethoven, Chopin, Hummel, Juplin, Mozart, and Scarlatti. From the paper Sheet Music Transformer.
70000
kern, bekrn,
ClassicComposers
Lots of classic music sheet and kern/MEI files.
6497
kern, MEI
OMRT
OMR datasets for testing OMR tools comprehensively.
330000
OMR datasets
A collection of collections. A collection of images from various sources of datasets. Some no longer work.
891
primus dataset title.png
Josquin project Stanford

The Josquin Research Project (JRP) changes what it means to engage with Renaissance music.

primus dataset title.png
Tasso

Digital Edition of the Settings of Torquato Tasso's Poetry

primus dataset title.png
RISM works

RISM Online, a service from the RISM Digital Center provides access to many musical sources through extensive search and browse functions. Multiple entry points, with extensive Source, Person, and Institution search interfaces, provide users with the ability to quickly narrow down relevant results using several different approaches, including keyword searches, filters, and relationship information.

primus dataset title.png
Byrd dataset

A small dataset of 34 high quality images with individual music score pages of increasing difficulty.

primus dataset title.png
IMSLP Petrucci music library

Petrucci music library. Contains 280571 works, 27853 composes. Share the worlds public domain music.

primus dataset title.png
Apacha music score classification dataset

A python script that trains a model that can learn to distinguish between music scores and arbitrary content. A dataset of 2000 images, containing 1000 images of music scores and 1000 images of other objects including text documents. The images were taken with a smartphone camera from various angles and different lighting conditions.

primus dataset title.png
Apacha printed music symbols dataset

This repository contains a set of about 200 printed music symbols of 36 different classes with and without context.

primus dataset title.png
Accidental detection dataset

This work takes place in the context of the thesis of of Kwon-Young Choi on the subject of resolving segmentation problems in dense and damaged printed piano scores from the 1750 to 1950 period. This data set contains 2955 images with annotations for the task of detecting or rejecting the correct accidental for a given note head.

primus dataset title.png
International audio lab Erlangen

Tools for semi automatic bounding box annotation of musical measures

primus dataset title.png
Polish digital scores

primus dataset title.png
Polish heritage open access

The Open Access project will expand the informational infrastructure in place at The Chopin Institute, which was designed as part of the POPC project ‘Chopin Heritage in Open Access’ (DCOD). The main aim is to broaden the portal to the most important and representative works of Polish music from the 16th to the 19th centuries, allowing searches according to musical elements and metadata.

primus dataset title.png
CMME datasets

Mensural music database and study project

primus dataset title.png
Sheet Music Benchmark

SMB (Sheet Music Benchmark) is a dataset of printed Common Western Modern Notation scores developed at the University of Alicante at the Pattern Recognition and Artificial Intelligence Group.

primus dataset title.png
Musicalion

Membership music sheet download.

primus dataset title.png
OpenOMR dataset

A dataset of 706 symbols (g-clef, f-clef) and symbol primitives (note-heads, stems with flags, beams) of 16 classes created by Arnaud F. Desaedeleer as part of his master thesis to train artificial neural networks.D

primus dataset title.png
Rebelo

Three datasets of perfect and scanned music symbols including an extensive set of synthetically modified images for staff-line detection and removal. Contains approximately 15000 music symbols.

primus dataset title.png
SEILS dataset

The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, MEI, **mens, agnostic, semantic and pdf formats, in white mensural and modern notation. The SEILS dataset is a corpus of scores in lilypond, music XML, MIDI, Finale, **kern, **mens, MEI, agnostic, semantic, and pdf formats, in white mensural and modern notation.

primus dataset title.png
Capitan collection

A corpus collected by an electronic pen while tracing isolated music symbols from Early manuscripts.

primus dataset title.png
MK free sheet music

A large collection of piano, voilin, vocal, trumpet, etc.

primus dataset title.png
Olympic dataset

Create a dataset of synthetic and scanned pianoform music for end-to-end OMR, called OLiMPiC. The dataset is built upon the OpenScore Lieder Corpus.

primus dataset title.png
Quartets sheet transformer

Generated together with GrandStaff by Alicante group. Quartets is a well-known collection employed in the Audio to Score field.

primus dataset title.png
Mozart scores Mozarteum by DIME

Printed score and MEI download of 30+ works.

primus dataset title.png
Humdrum Polish digital scores

Songs and music from Poland, including Chopin.

primus dataset title.png
RISM

Reportoire International de Sources Musicales contains catalogged music sources funded by barious agencies. It includes libraries in Berlin and Bavaria.

primus dataset title.png
PDMX

Public domain muxicxml dataset for symbolic music processing.

primus dataset title.png
MUSCIMA++

A subset of MUSCIMA that is fully annotated. It used MuNG music notation graph format.

primus dataset title.png
Musicxml collection

A collection of music library with musicxml

primus dataset title.png
MusicXML Sample

Small sample of music in musicXML format

primus dataset title.png
MSMD Multimodal

MSMD is a synthetic dataset of 497 pieces of (classical) music that contains both audio and score representations of the pieces aligned at a fine-grained level. Music of Bach, Chopin, Froberger, and others in full page (typically 1) in pdf format.

primus dataset title.png
Primus

Printed Images of Music Staves. Two packages of synthetic music with one line segments printed. No musicxml and kern.

primus dataset title.png
MUSCIMA

Handwritten music score images.

primus dataset title.png
DoReMe

Part of Dorico project, most pictures are one line long and one staff long. File in png format

primus dataset title.png
Open Score Lieder

Collection of 1200 10th century songs annotated by a team of volunteers.

primus dataset title.png
Humdrum Polish scores

This repository contains transcriptions from the Heritage of Polish Music in Open Access project at the Chopin Institute in the Humdrum digital score format. Files are organized by the RISM siglum ID of the source archive. Here is a promotional booklet about the project in Polish and English.

primus dataset title.png
DeepScore

Page long or less png files with annotation in json. A tool set accompanies.

primus dataset title.png
GrandStaff

Music segments from Beethoven, Chopin, Hummel, Juplin, Mozart, and Scarlatti. From the paper Sheet Music Transformer.

primus dataset title.png
ClassicComposers

Lots of classic music sheet and kern/MEI files.

primus dataset title.png
OMRT

OMR datasets for testing OMR tools comprehensively.

primus dataset title.png
OMR datasets

A collection of collections. A collection of images from various sources of datasets. Some no longer work.

bottom of page