Explainer

what is optical
music recognition?

The technology that turns a photograph of a page back into notes: why it is much harder than reading text, and where it still falls over.

Updated · August 22, 2026 8 min read Plain English

The short definition

Optical music recognition, or OMR, is the process of reading musical notation from an image and turning it back into music that software can understand. Point a camera at a printed page, and OMR is what works out that the shapes on it mean a C sharp, a dotted quaver, and a change to three-four time.

SheetlyAudiveris, the best free scanner
96%
88%
Clean scans
95%
50%
Phone photos
90%
58%
Real scans
Notes read correctly on the same pages, August 2026. How we measured
A phone photograph of a printed trombone partThe same photograph with each detected staff coloured
Step one of optical music recognition: finding the staves in a photograph that is neither flat nor evenly lit.

It is often described as “OCR for music”, which is a useful starting point and a slightly misleading one. The goal is the same: recover the meaning a printer put on a page. The difficulty is not remotely the same.

Why it is much harder than reading text

Text is a line of symbols. Each letter means the same thing regardless of where it sits, meaning runs in one direction, and when a shape is ambiguous a dictionary usually settles it. Music has none of those conveniences.

Position carries the meaning

The same oval notehead is a G on one line and an A on the space above it. Getting the shape right and the position wrong is a complete failure. In text, a letter recognised correctly is correct wherever it sits.

Several things happen at once

A staff can carry two independent voices with different rhythms, written with stems pointing in opposite directions. A piano brace carries two staves that are one instrument. An orchestral page carries a dozen staves that are not. Working out what belongs with what is a structural problem before it is a recognition problem.

Context propagates

A key signature at the start of a system silently alters every matching note until it changes. An accidental alters the rest of its bar. A clef alters everything after it. One misread symbol therefore does not produce one wrong note. It produces a wrong page. This is the property that makes music recognition unforgiving in a way text recognition is not.

Symbols overlap and touch

Engravers cram beams, slurs, ties, ledger lines and dynamics into a small space, and in dense writing they collide. Separating a slur from the stem it crosses is a segmentation problem that has no equivalent in a line of printed prose.

There is no dictionary

Text recognition leans hard on language models: “teh” is almost certainly “the”. Music has conventions: bars should contain the right number of beats, harmony tends to behave. But there is no lookup that says this passage must be those notes. The strongest available check is arithmetic: does each bar add up?

What a decoded page looks like

The Sheetly player on an iPhone. Every page you scan opens like this, ready to play, loop and slow down.

How a page is actually read

Systems differ, but nearly all of them pass through these stages in roughly this order.

  1. PreprocessingStraighten the image, correct perspective, even out the lighting, and separate ink from paper. With a photograph this stage is doing a great deal of work before any music is involved.
  2. Staff detectionFind the five-line staves. Everything downstream is measured relative to them, so an error here is fatal rather than merely damaging.
  3. Staff groupingWork out which staves are braced into one instrument. This is what decides whether a page is a piano or two separate players, and it is one of the quietest, most consequential failure points in the whole pipeline.
  4. Symbol recognitionIdentify the noteheads, stems, beams, flags, rests, accidentals, clefs and time signatures, and where each one sits.
  5. Musical interpretationTurn positions into pitches using the clef and key, turn shapes and beams into durations, assign notes to voices, and resolve the accidentals that carry through a bar.
  6. ValidationCheck the result against the rules music has to obey, most usefully whether each bar contains the beats its time signature demands. A bar that does not add up is a reliable flag that something upstream went wrong.
  7. EncodingWrite the result out as MusicXML or MIDI.

Scans versus photographs

Almost all published OMR research assumes a flat scan: the page pressed against glass, lit evenly, square to the sensor. That assumption is doing more work than it appears to.

A phone photograph breaks it in four ways at once. Perspective makes the far end of the page smaller, so staff spacing is no longer constant. Curvature from a bound book bends the staff lines that everything is measured against. Uneven light means one threshold cannot separate ink from paper across the whole page. Glare simply deletes whatever is underneath it.

The size of that gap is measurable. Our own 31-piece benchmark is scored by strict F1, where a note counts only if pitch, onset and duration are all correct. On it, Audiveris 5.11.0 scored 88% on clean renders and 50% on the same pieces degraded to a moderate photograph. Sheetly scored 96% and 95%. The clean-page difference is modest; the photograph difference is the whole ballgame, and it is the case that matters when the scanner in question is a phone.

Sheetly’s own measurements, on identical inputs through an identical scorer. The method is written up here.

What it still cannot do

  • Handwritten manuscript. Engraving is standardised; handwriting is not. This remains an open research problem, and any product claiming to have solved it deserves careful testing.
  • Dense orchestral scores. Many staves, small print and heavy overlap compound every difficulty above.
  • Tablature, chord diagrams and drum notation. Different symbol systems, generally needing purpose-built recognition.
  • Interpretation. OMR recovers what is written. Rubato, voicing and pedalling are decisions a performer makes, and no amount of reading the page will supply them.

Knowing the boundary is what makes the tool useful. A system that reads a printed page well and says so plainly is worth considerably more than one that claims to read everything and quietly fails on a third of it.

Ask Sheetly: a coach that has read your score

Not a chatbot bolted onto a music app. The coach has the actual notes of the page you just photographed, so you can ask about your bar 12 and get an answer about what is written in it.

bar 12 keeps falling apart, what am I doing wrong?
Bar 12 is where the left hand crosses over. You are almost certainly re-taking the thumb. Try 2–1–3 and keep the wrist moving right through the bar rather than resetting on the downbeat.
what should I drill tonight?
Bars 12–15, hands separately, at 60. That is the only place the texture changes; the rest of the page is a variation on the opening you already have.

We know of no other sheet-music scanner that does this. Every other tool on this page hands you a transcription and stops.

Download on the App StoreFree to download · iPhone and iPad

Questions

What does OMR stand for?

Optical music recognition. It is the process of reading printed or handwritten musical notation from an image and converting it into a machine-readable form such as MusicXML or MIDI, and it is the musical equivalent of OCR for text.

How is OMR different from OCR?

OCR reads text, which is one-dimensional: symbols follow one another left to right, and each one means the same thing wherever it sits on the line. Music is two-dimensional. A notehead’s vertical position is its pitch, several voices can sound at the same instant, and symbols such as key signatures and clefs change the meaning of every note that follows them. Music also has no equivalent of a dictionary to fall back on when a symbol is ambiguous.

Why is optical music recognition so hard?

Because meaning is carried by position, because context propagates so that one misread clef transposes an entire system, because symbols overlap and touch each other in dense writing, and because a single error can cascade through the rest of the page. Photographs make all of it harder by adding perspective, shadow, curl and glare on top.

Is OMR a solved problem?

For clean, printed, single-line music it is close to solved. For dense multi-voice piano writing, for photographs rather than flat scans, and for handwritten manuscript, it is not. Handwriting in particular remains an active research problem rather than a shipped product feature.

What formats does OMR usually output?

MusicXML, which preserves the written score with its clefs, spelling, beaming and layout, and MIDI, which preserves the performance as a list of note events. MusicXML carries more information; MIDI is what instruments and DAWs speak.

Read next