Research

how accurate is a sheet music scanner,
measured properly.

Sheetly, Audiveris 5.11.0 and oemer on 31 clean renders, the same 31 as simulated phone photographs, and 60 systems cut from real scanned pages we did not author. Strict F1, one scorer, one scoring window. Every input, every raw transcription and the scorer itself is published, so you can recompute the table rather than believe it.

Run · August 2026 Data published Updated · August 23, 2026 12 min read

Summary

Sheetly scored 96.44% strict F1 on 31 clean scores, 95.28% on the same 31 degraded to simulated phone photographs, and 90.3% mean / 99.2% median on 60 systems cut from real scanned pages in the public OLiMPiC suite. Audiveris 5.11.0 scored 88.43%, 50.43% and 57.6% / 64.4% on the same pages, or 69.2% / 74.4% if its ten failures are excluded rather than scored zero. oemer scored 38.45% and 9.81% on the 31-piece suite.

None of that has to be taken on trust. The ground truth, every system’s raw MusicXML output, the scorer and the photograph degradation are published as a 2 MB download; one command recomputes the two 31-piece tables in well under a minute, on a laptop, with no dependencies. The OLiMPiC table needs one extra step, because that dataset’s ground truth is not ours to redistribute. If a figure here were flattering, the files would say so. Check it yourself →

A trombone part photographed in a ring binder with every staff the scanner found highlighted in its own colour
A photographed page with every staff the scanner found highlighted. That is the first step behind every number on this page.
In short
  • Metric: strict F1. A note counts only when pitch, onset and duration are all correct.
  • Second metric (OMR-NED, whole notation, lower is better): clean renders Audiveris 0.216 · Sheetly 0.236 · oemer 0.734, so Audiveris wins; photographs Sheetly 0.260 · Audiveris 0.538 · oemer 0.944. Why both metrics →
  • Clean renders (31): Sheetly 96.44% · Audiveris 5.11.0 88.43% · oemer 38.45%.
  • Simulated phone photographs (same 31): Sheetly 95.28% · Audiveris 50.43% · oemer 9.81%.
  • Real scans (OLiMPiC, 60 systems, not ours): Sheetly 90.3% mean / 99.2% median, 60 of 60 usable · Audiveris 57.6% / 64.4% counting its ten failures as zero, or 69.2% / 74.4% over the 50 it transcribed.
  • Checkable: every per-piece score is on this page, and the inputs, the raw transcriptions and the scorer are downloadable. One command reproduces the headline figures.
  • Not measured: PlayScore 2, Halbestunde, ScanScore, PhotoScore, Sheet Music Scanner. None can be driven over a batch of files, so this page claims nothing about them, but the test pages and the protocol are published, and we will publish the result either way.

Every system was measured on the same pinned input files with the same scorer: Sheetly’s August 2026 production pipeline, Audiveris 5.11.0 and oemer 0.1.x. Both are the current releases of those projects, and unchanged since they were run. This is a dated snapshot rather than a permanent property of any of the three, and it is written by the people who make one of them. That is exactly why the inputs, the outputs and the scorer are published instead of described.

Why a benchmark

Every sheet music scanner describes itself as accurate, and almost none says accurate at what, measured how, on which kind of page. A benchmark replaces that sentence with three things that can be argued with: a metric, a set of inputs, and a procedure that treats every system identically. Without those, a comparison is an opinion with a percentage sign on it.

The particular question we wanted answered is not “can OMR read a flatbed scan”, which it can, and has for years, but “what happens when the page is a phone photograph, with perspective, shadow and curl”. That is the input a player in a practice room actually has, and it is where published accuracy figures, which are almost always measured on clean input, stop describing the experience.

The second question is whether a system's own suite can be trusted. It cannot, entirely, which is why the third suite on this page is one we did not author.

What is measured

Every figure on this page is strict F1 at the note level. A transcribed note is counted as correct only when its pitch, its onset within its bar and its duration all match the ground truth. F1 is the harmonic mean of precision, the fraction of transcribed notes that were right, and recall, the fraction of true notes that were found. A score of 100% is a perfect transcription; 0% is no usable notes at all.

How notes are matched

Ground truth and transcription are each flattened to a sequence of notes ordered by onset, and the two sequences are aligned by longest common subsequence under one predicate: same pitch, same onset within the bar to a tolerance of a sixty-fourth, same duration to the same tolerance. Extra notes in the transcription cost precision; missing notes cost recall.

Onset is measured inside its bar rather than from the start of the piece, deliberately: a system that drops a rest in bar 3 should lose that bar, not every bar after it. The bar a note lands in is not part of the predicate either, only the order the alignment must respect. This is a real relaxation and it is worth saying who it helps. It is not us. Re-scored with the bar number in the predicate, Sheetly falls 0.0 and 2.1 points on the two suites, Audiveris 1.9 and 13.6, oemer 16.6 and 3.8. A stricter metric than the one published would widen every gap on this page.

Why pitch-only accuracy flatters

Most quoted OMR accuracy figures count pitches. A transcription that gets every pitch right and every rhythm wrong scores close to 100% on a pitch measure and is, to a player, unusable: it will play back as a different piece. Strict F1 scores that transcription as the failure it is. It also punishes the common partial failures that pitch-only measures either ignore or count once: a dropped dot, a tuplet read as straight quavers, a missed clef change that shifts a passage by an octave.

The consequence is that strict F1 numbers look lower than the figures scanners usually quote. That is the point. A 98% here means roughly one note in fifty is wrong in some respect; a 98% on a pitch-only measure can hide a rhythmically broken page.

What the metric does not measure

“Strict” is strict about three things and silent about the rest, and the silences should be named. Scored: pitch, onset and duration, for every pitched note. Not scored: rests, which staff or hand a note was assigned to, enharmonic spelling (G♯ against A♭), clefs, key signatures, time signatures, repeats and voltas, slurs, articulations, dynamics and lyrics. A transcription that puts the whole left hand on the right staff, or spells every accidental the wrong way, loses nothing here. The scorer computes staff accuracy, spelling accuracy and an exact-measure rate as well, published per piece in the results JSON as staff_acc, spell_acc and measure_exact, but the headline figure is note-level only, and it applies identically to every system in the table.

The established metric: OMR-NED

After this page was first published, the maker of a competing engine (Soundslice) made two public criticisms of it: strict F1 scores only pitched notes and ignores the rest of the notation, and the scorer is our own code when established metrics exist. Both points are fair, so as of August 31 the page reports the established whole-notation metric alongside strict F1: OMR-NED, the OMR normalized edit distance introduced with the Sheet Music Benchmark at ISMIR 2025 (arXiv:2506.10488) as the successor to the older symbol error rate. It counts edit operations over every notation symbol, normalized by the symbol count of both scores: note heads, beams, dots, accidentals, ties, slurs, articulations, ornaments, clefs, key and time signatures, dynamics, directions and lyrics. 0.0 is a perfect transcription and lower is better. None of the scoring code is ours this time: every comparison is musicdiff 5.2, the metric authors’ own tool, at default settings, over the same published page-one files strict F1 is scored on.

OMR-NED on the 31-piece suites, page one of each piece, computed with musicdiff 5.2 at default settings. Lower is better; 0.0 is a perfect transcription. Mean rows count a piece an engine produced nothing for as 1.0; corpus rows are total edit operations over total symbols across the pieces the engine transcribed (Audiveris missed 3 photographs; oemer missed 8 clean and 14 photographs).
Suite · OMR-NED Sheetly Audiveris 5.11.0 oemer 0.1.x
Clean renders, mean (n = 31, missing = 1.0) 0.2360.2160.734
Simulated photographs, mean (n = 31, missing = 1.0) 0.2600.5380.944
Clean renders, corpus-level 0.2650.2600.690
Simulated photographs, corpus-level 0.2980.5130.892

The result is worth stating against our own interest: on clean renders, Audiveris wins this metric, 0.216 to Sheetly’s 0.236 mean and 0.260 to 0.265 corpus-level, precisely because it reads slurs, articulations, dynamics and ornaments that Sheetly does not read at all. Which marks those are is published one by one on the symbols page, and this table is what those gaps cost when the whole notation is scored. On phone photographs, the input the product exists for, Sheetly leads under both metrics, 0.260 against 0.538. Strict F1 stays as the headline because it answers the player’s question of whether they will hear the right notes, while OMR-NED answers the researcher’s: how much of the notation survived. Which of the two describes your experience depends on the job. The marks Sheetly does not read are dynamics, slurs and ornaments. They say how a page should be interpreted, not which notes sound, so they cost little when the job is photographing a part and practising against it, and a lot when the job is re-engraving or archiving a score with every marking intact. For that second job, on a clean flatbed scan, this table says to use Audiveris, and tools built for notation work, which is a different product category from a practice app. A fair comparison needs both numbers, which is why both are now published, per piece, with the code: omrned-results.json and omr_ned.py.

The three suites

Three sets of inputs were used, each answering a different question. The first two are ours; the third is not, and is the one to trust most.

1. Thirty-one clean renders

Thirty-one pieces chosen to span the cases that break scanners: single-line parts, grand-staff piano writing with independent voices and cross-staff beams, multi-part vocal and chamber scores, repeats and voltas, tuplets, dense ledger lines, and unusual metres and clefs. Each has ground-truth notation and was rendered clean, as a flat, evenly lit, square-on page. This is the easiest case, and the one most published figures describe.

2. The same thirty-one as simulated phone photographs

Each clean render was passed through a deterministic degradation that imitates a phone photograph of a page on a stand: perspective skew, uneven lighting and shadow, page curl, mild blur and sensor noise, at a fixed “moderate” severity. Deterministic means the degradation is reproducible from the same source file rather than sampled anew per system. Because we designed this degradation, this suite is the weakest evidence on the page: it would be possible, in principle, to tune a system to survive exactly this and nothing else.

Moderate is the middle of three settings

The degradation model, published as photo_sim.py in the bundle, defines three severities. Light is a careful capture in good light. Moderate, which every figure on this page uses, is a typical handheld shot: visible perspective, uneven light, soft focus. Heavy is described in the file itself as “the kill zone: strong angle, shadow, page curl, dim light, JPEG mush”.

One property of the model matters more than its severity label and is easy to miss: it does not resample. The degraded image is the original render at full scale on a larger canvas, so staff lines keep their ideal pixel width and the blur is a 1.1-pixel Gaussian across a 2,000-pixel page. A real phone capture varies its resolution, its distance and its focus, and resolution is the single variable OMR is most sensitive to. So the photograph column measures perspective, shadow, curl and mild blur, not the loss of detail that a real handheld shot also brings. Read the −1.16 point cost against that.

We do not publish a current heavy figure, and it is fair to ask why. The only heavy run we have is from July 2026, on a six-piece subset and an older pipeline, where Sheetly scored 76.2%. That number is too stale and too small to sit in a table next to the rest, and quoting it as though it were current would be worse than the gap. It is here rather than nowhere, and the next run will measure all three severities across the full suite.

3. Sixty systems cut from real scanned pages (OLiMPiC)

OLiMPiC 1.0-Scanned is a public dataset from ICDAR 2024, released under CC BY-SA, which pairs crops of original IMSLP library scans with MusicXML ground truth. It is the one suite here we did not author, and the artefacts in it are genuine rather than simulated: bleed-through, skew, faded print, binding shadow. This is the check on suite 2: if a system only looked good on our own degradation, it would show here.

Two things about it should be stated plainly, because they make it both stronger and narrower evidence than the other suites.

  • Each sample is one system, not a page. OLiMPiC crops per staff system, so a sample is a grand-staff system lifted out of a scanned page: 16 to 122 notes, 52 on average. That removes the page-level problems a phone photograph has: finding the staves across a whole page, grouping systems, following the music from one system to the next. A figure measured here is a figure for reading real engraved ink, not for reading a whole page.
  • The 60 were not chosen by us, but the rule has a bias worth naming. The selector sorts the dataset’s own sample ids and takes them round-robin, one system per score; the first 60 are therefore always the same 60, and picking a different set would mean changing the rule rather than the result. Because 60 is one full round, every sample is the first system of page one of its score. That is the strip that carries the clef, key and time signature, and the one least likely to sit in a page interior. The list is in samples.txt; a run with --n=200 would reach further into each score.

Systems compared and how each was run

Three systems appear: Sheetly, Audiveris 5.11.0 and oemer 0.1.x. The rule for inclusion was simple. A system is in the table only if it can be run head-to-head, with identical input files in, a machine-readable score out and the same scorer applied, so that the comparison is checkable rather than reported.

  • Sheetly. The production recognition pipeline, run on each input exactly as an uploaded photograph would be, with one deliberate difference: the LLM and vision passes are switched off, so every run is deterministic. Those passes read the title block and guess the instrument; they do not read notes, and every figure here is note-level. You can see the switch in our own output files, each of which carries the line LLM/vision passes are disabled in the offline pipeline. It costs us nothing on this metric and it means a re-run cannot drift.
  • Audiveris 5.11.0. The open-source desktop engine, run from its command-line interface in batch mode on each input with default settings, exporting MusicXML.
  • oemer 0.1.x. The open-source end-to-end neural OMR project, run from its command-line interface on each input, exporting MusicXML.

Every output was scored by the same strict F1 scorer against the same ground truth, and no system was given a friendlier page or a friendlier metric. Inputs were byte-identical on the clean suite. On the photographed suite they were not: the competitor runner was staged with page one of each piece only, so Sheetly saw second pages that Audiveris and oemer never received. That is why the published figures score page one for everybody. The scoring is like-for-like, the staging was not, and both halves of that sentence should be read together.

Page one of each piece, for everybody

The 31-piece figures below score page one of each piece only. Eighteen of the pieces run to a second page, and when the photographed suite was staged for the competitor runner only the first page of each piece was staged, so Audiveris and oemer were never shown those second pages. Scoring them against ground truth that included the missing page charged them for music they had not been given, and that is how this benchmark was first published: it made Audiveris’s photograph figure 38.55% when the pages it actually saw score 50.43%.

Restricting every engine to page one removes the asymmetry, and it costs 44% of the notes: 594 bars and 6,610 notes are scored, out of 935 bars and 11,861 notes across all pages. So “31 pieces” here means 31 opening pages of between 7 and 36 bars, median 18, not 31 complete works. On that basis Sheetly scores 96.44% and 95.28% rather than 97.03% and 96.00%. Both sets of files are in the bundle, and python3 score.py . --all-pages prints the multi-page version. The correction is recorded in the changelog.

Pinned versions, pinned files

This is a fixed-input comparison, not a timed race. What makes it a comparison is that every system was scored by the same code over the same window: page one of each piece, which is the part every system was given. Sheetly’s figures are its August 2026 pipeline; Audiveris 5.11.0 and oemer 0.1.x were run in July 2026 and are still the current releases of both projects, so running them again on the files they were given produces the same MusicXML, which is published in the bundle, page by page, for anyone who would rather check than take our word for it. Sheetly’s own earlier figures are kept in the changelog rather than deleted.

Why PlayScore 2 and Halbestunde are not in the table

They are the two products people actually weigh Sheetly against, and their absence is the biggest hole in this page. The reason is mechanical rather than strategic: both are phone apps with no batch interface, and in both the MusicXML export needed to score anything sits behind a paid tier. There is no way to push 62 files through either one automatically. A fair comparison needs a person with a phone, importing 62 images by hand, exporting 62 files, and correcting none of them.

So we have published everything that person would need.

  • The test pages (29 MB). Page one of all 31 pieces, as clean renders and as the photographed versions: the pages every figure on this page is scored over. The second pages, which only some engines received and which nothing here is scored on, are not in it.
  • The ground truth and the scorer (2 MB), so an export from any app can be scored by the same code that produced every figure here.
  • The protocol: import each image, first recognition attempt only, no manual correction, export MusicXML, record the app version and the date. An app that refuses a page scores zero for that page rather than being excused from it, the same rule Audiveris and oemer were held to.

If you run it, send the exports to hello@trysheetly.com and we will publish the table with your name on the method and a link to your files, whatever it says. If Sheetly loses a column, the column goes up losing, dated, next to the rest. Until someone does that, this page makes no claim about either product, and neither should anyone quoting it.

Results

On the 31-piece suite Sheetly scored 96.44% clean and 95.28% photographed; Audiveris 88.43% and 50.43%; oemer 38.45% and 9.81%. On the 60 OLiMPiC systems Sheetly scored 90.3% mean and 99.2% median with usable output on all 60 scores; Audiveris 57.6% mean and 64.4% median with usable output on 50 of 60, or 69.2% and 74.4% over only the 50 it transcribed.

31-piece suite: clean renders and simulated photographs

Mean strict F1 across 31 pieces, page one of each piece, from the transcriptions published in the data bundle. Page one is the window every engine received. A piece an engine produced nothing for scores zero and stays in the mean. Clean = flat rendered page. Photographed = the same page through a fixed moderate phone-photo degradation. Identical scorer, identical scoring window. Higher is better; 100% is a perfect transcription.
Suite Sheetly Audiveris 5.11.0 oemer 0.1.x
Clean renders (n = 31) 96.44%88.43%38.45%
Simulated phone photographs (n = 31) 95.28%50.43%9.81%
Drop from clean to photograph −1.16 pts−38.00 pts−28.64 pts
Pieces with no usable output 0 and 00 and 38 and 14

OLiMPiC: 60 systems cut from real scanned pages

Strict F1 on 60 systems cut from real scanned pages, from the public OLiMPiC suite, which we did not author. “Usable output” counts samples for which the system produced a parseable transcription at all; a sample with no output contributes 0% rather than being dropped, which is the difference between the two mean rows and the two median rows. Sheetly transcribed all 60, so both of its conventions coincide; quoting Audiveris’s 64.4% without its 74.4% would be quoting the harsher of two defensible rules. Both engines’ transcriptions for all 60 are in the bundle, so these four rows can be recomputed. oemer was not run on this suite.
Measure Sheetly Audiveris 5.11.0
Mean strict F1 90.3%57.6%
Median strict F1, over all 60 99.2%64.4%
Median over only what it transcribed 99.2%74.4%
Usable output 60 of 6050 of 60
Mean over only what it transcribed 90.3%69.2%

The ordering is the same on all three suites, and the shape of the result is that the photograph column has almost stopped costing anything: a page photographed on a stand scores just over a point below the same page rendered flat, where the best free engine loses thirty-eight. The gap to that engine is eight points on clean pages, forty-five on photographs, and thirty-three on real scans.

What changed between July and August

Fourteen changes to the recognition pipeline, not one of them a retrained model. Each is a deterministic, evidence-gated correction that only fires when it can prove its case on the page in front of it, and each had to leave a set of six reference pieces at a perfect score before it was allowed to ship. Most fix one structural misreading: a pair of staves welded into a single row, a tenor clef whose small printed 8 was being read as a lyric syllable, a time signature inferred from bar lengths instead of read off the page, a system of unequal size re-shredded rather than aligned to its neighbours. Eight of the thirty-one clean pieces improved, ten of the same pieces photographed, and five of the sixty real scans; one photographed piece slipped by three thousandths, and nothing else went backwards. Those per-piece counts come from the July run, whose per-piece figures we have not published, so they are the one set of numbers on this page you cannot check against a file. The suites, the inputs and the scorer did not change, which is what makes July and August comparable at all.

Every piece, every score

Benchmarks are usually published as three numbers, which is exactly enough to hide behind. Here is the whole thing: all 31 pieces, every system, clean and photographed, sorted by size. The pieces where Sheetly does worst are in it too. A dash means the system produced no usable transcription for that page, which is counted as zero in the means above rather than quietly dropped.

Strict F1 per piece. Sheetly’s two columns are shaded. A dash is no usable output. Every one of these figures can be recomputed from the published bundle.
Piece Notes Sheetlyclean Sheetlyphoto Audiverisclean Audiverisphoto oemerclean oemerphoto
Joplin: Maple Leaf Ragtwo-hand ragtime piano (advanced)518100.0%99.4%92.0%69.5%72.3%6.9%
Clara Schumann: Polonaise op. 1 no. 1romantic piano polonaise w/ key change50891.1%91.1%72.6%59.6%43.6%6.3%
Verdi: La donna è mobilearia in 3/8 with lyrics37099.6%95.1%81.4%79.3%––
Chopin: Mazurka op. 6 no. 2mazurka, ornamented romantic piano35991.1%94.4%96.6%87.2%49.7%53.8%
Schubert: Der Lindenbaumlied, triplet accompaniment, key change, lyrics342100.0%100.0%98.0%51.4%48.6%9.9%
Weber: Clarinet Concertinoclarinet concertino, 3 meters, 2 keys, runs33799.4%67.6%93.6%0.0%52.1%–
C. P. E. Bach: H. 186galant keyboard piece30084.0%83.7%89.4%60.0%39.7%24.7%
Clara Schumann: Trio op. 17, IIIpiano trio mvt in 6/827692.6%92.0%73.1%66.1%24.1%–
Beach: Prayer of a Tired Childchoral + piano, lyrics everywhere26390.2%90.0%94.1%88.3%12.7%10.8%
Mozart: Quartet K. 155, IIquartet slow mvt, triplets + graces258100.0%100.0%92.2%78.1%––
Mozart: Quartet K. 458 ‘Hunt’, IHunt quartet opener, 6/8 with pickup257100.0%100.0%99.6%6.5%––
Schumann: Dichterliebe no. 2lied with pickup + lyrics25495.8%96.2%75.9%16.4%44.3%29.8%
Luca: Gloriamedieval gloria, 3 meter changes23198.3%98.3%96.1%6.2%35.7%–
Beethoven: Quartet op. 18 no. 1, Ilong sonata-form quartet mvt21595.3%95.3%94.3%–––
Haydn: Quartet op. 74 no. 1, Iclassical quartet allegro198100.0%100.0%96.4%–––
42d Highland Regiment Strathspeyfiddle tune with repeat barlines19292.7%100.0%82.2%27.8%61.1%0.5%
Handel: Lascia ch’io piangaaria, meter + key change, lyrics17597.4%97.4%95.9%41.9%52.9%–
Beethoven: Große Fuge op. 133Grosse Fuge string quartet, dense 4-part (extreme)174100.0%98.3%94.5%92.5%––
Bach: Chorale BWV 66.64-voice chorale (4 staves/system)165100.0%100.0%100.0%–––
Schumann: Quartet op. 41 no. 1, IIIquartet mvt in cut time w/ meter change15198.0%100.0%99.0%66.7%––
O’Neill’s Collection: tuneO'Neill's collection tune (pinned GT)13392.5%89.5%77.3%44.1%33.0%0.0%
Ryan’s Mammoth: Acrobat’s Hornpipedotted hornpipe with repeats12399.2%95.9%99.2%45.2%52.6%0.0%
Ryan’s Mammoth: Annie Hughes’ Jig6/8 jig with repeats11987.4%87.4%93.0%36.3%56.9%65.8%
Lili‘uokalani: Aloha ‘Oesong with pickup, many staves115100.0%100.0%100.0%100.0%47.0%53.0%
Synthetic: repeats and endingsmonophonic melody with ||: :|| and 1st/2nd endings103100.0%100.0%97.5%97.5%74.8%–
Schoenberg: op. 19 no. 2atonal piano, accidentals everywhere (extreme)102100.0%97.1%75.2%41.9%52.9%5.9%
Palestrina: Agnus Deirenaissance 5-voice mass mvt97100.0%100.0%90.1%43.8%29.0%0.0%
Foster: Jeanie with the Light Brown Hairlead sheet, melody + lyrics95100.0%100.0%28.8%27.3%82.1%27.4%
Aird’s Airs: highland tuneAird's Airs highland tune (pinned GT)8498.8%98.8%81.2%77.7%85.2%6.0%
Essen folksong: old GermanEssen folksong, old German (pinned GT)6091.7%91.7%87.8%82.3%65.0%3.3%
Essen folksong: balladEssen folksong, ballad (pinned GT)3694.4%94.4%94.4%69.8%76.7%–

Twelve of the 31 come back note-perfect on clean renders and eleven survive the photograph untouched. Two clean pieces and four photographed ones fall below 90%, and those are where the next run’s work is. On the other side of the table, oemer produced nothing at all for eight clean pages and fourteen photographed ones, and Audiveris for three photographed ones.

The aggregate also hides five clean pieces where Audiveris beats us: the Chopin mazurka (96.6% against our 91.1%), Beach’s Prayer of a Tired Child (94.1% against 90.2%), Ryan’s jig (93.0% against 87.4%), the C. P. E. Bach (89.4% against 84.0%) and the third Schumann quartet movement (99.0% against 98.0%). The eight-point gap on clean renders is not uniform superiority; it is Sheetly being steadier, with Audiveris collapsing on a handful of pages rather than losing a little everywhere. On photographs it is behind on thirty of the thirty-one, and ties the remaining one, Aloha ‘Oe, where both are perfect.

The 60 real scans, one by one

The same again for the OLiMPiC suite, with Audiveris alongside. The distribution matters more than the mean here, so it is worth saying exactly what it is: 27 of the 60 are perfect, 47 are at 90% or better, 11 fall below 80%, and 4 fall below 50%, the worst of them at 8.9%. That tail is the whole distance between the 99.2% median and the 90.3% mean. In practical terms, most real scans come back essentially right and roughly one in fifteen comes back badly enough to be worth rescanning or fixing by hand. Anyone quoting the 90.3% without the tail, us included, would be quoting it wrong.

Strict F1 for each of the 60 OLiMPiC systems, best first by Sheetly’s score. Sample identifiers are the dataset’s own. A dash is no usable output, counted as zero in the means above.
OLiMPiC sampleNotesSheetlyAudiveris 5.11.0
4919798/p1-s185100.0%37.8%
4978492/p1-s154100.0%59.3%
5000459/p1-s167100.0%98.5%
5000467/p1-s196100.0%63.4%
5026309/p1-s141100.0%90.2%
5026316/p1-s145100.0%87.4%
5062159/p1-s158100.0%92.0%
5069022/p1-s139100.0%100.0%
5069028/p1-s131100.0%68.9%
5071644/p1-s138100.0%21.3%
5659823/p1-s143100.0%46.5%
5667944/p1-s133100.0%47.3%
5701612/p1-s144100.0%90.9%
6114679/p1-s171100.0%99.3%
6209594/p1-s146100.0%50.0%
6209605/p1-s126100.0%96.0%
6232117/p1-s165100.0%18.5%
6243836/p1-s116100.0%86.7%
6320420/p1-s120100.0%–
6377942/p1-s144100.0%62.8%
6378337/p1-s124100.0%95.7%
6378373/p1-s125100.0%95.8%
6379527/p1-s136100.0%0.0%
6379750/p1-s118100.0%83.9%
6482032/p1-s148100.0%94.9%
6567603/p1-s141100.0%–
6568021/p1-s136100.0%80.0%
5820797/p1-s18299.4%75.5%
6243839/p1-s16099.2%79.7%
6351349/p1-s16099.2%94.9%
5062158/p1-s15899.1%87.9%
5700433/p1-s15499.1%69.8%
5062156/p1-s14898.9%82.6%
6447286/p1-s16998.5%–
6569855/p1-s16598.5%86.9%
4982505/p1-s16498.4%60.0%
5656522/p1-s15698.2%–
5834392/p1-s15798.2%–
5654342/p1-s12297.7%31.8%
5000464/p1-s14097.5%95.0%
6565727/p1-s17097.1%–
6209602/p1-s13096.7%78.0%
6568070/p1-s15694.6%46.8%
6209608/p1-s14794.5%90.7%
6481218/p1-s13694.4%25.4%
6209607/p1-s14493.2%40.9%
5840228/p1-s18890.0%–
6447342/p1-s15488.9%65.4%
6548435/p1-s18181.5%–
6569865/p1-s16078.3%98.3%
5026306/p1-s15378.0%52.7%
6115317/p1-s16377.8%63.1%
6112984/p1-s18073.0%41.7%
6243838/p1-s13671.4%89.6%
6447372/p1-s13159.0%71.0%
5837811/p1-s110257.1%–
4982465/p1-s112246.5%59.4%
6481941/p1-s13234.4%73.3%
5098713/p1-s13620.3%30.8%
6302395/p1-s1498.9%–

Multi-staff and dense scores

The most common thing said about a young scanner, usually without anyone having measured it, is that it copes with a melody line but falls apart on real polyphony: a quartet, a chorale, a piano texture with independent voices. It is a fair thing to want checked, so the suite was built with those cases in it from the start and the per-piece results are published rather than summarised.

Strict F1 on the multi-staff and multi-voice pieces in the 31-piece suite, page one of each. Sheetly’s columns are shaded; Audiveris 5.11.0 is shown for scale. Every figure here is recomputable from the published bundle.
Piece Notes Sheetlyclean Sheetlyphoto Audiverisclean
Bach: Chorale BWV 66.64 voices on 4 staves165100.0%100.0%100.0%
Palestrina: Agnus Dei5-voice Renaissance mass movement97100.0%100.0%90.1%
Beethoven: Große Fugedense 4-part string quartet174100.0%98.3%94.5%
Haydn: Quartet op. 74 no. 1classical quartet allegro198100.0%100.0%96.4%
Mozart: Quartet K. 4586/8 with a pickup257100.0%100.0%99.6%
Mozart: Quartet K. 155triplets and grace notes258100.0%100.0%92.2%
Lili‘uokalani: Aloha ‘Oesong across many staves115100.0%100.0%100.0%
Luca: Gloriathree meter changes23198.3%98.3%96.1%
Schumann: Quartet op. 41 no. 1cut time with a meter change15198.0%100.0%99.0%
Beethoven: Quartet op. 18 no. 1long sonata-form movement21595.3%95.3%94.3%
Beach: Prayer of a Tired Childchoir plus piano, lyrics throughout26390.2%90.0%94.1%

Seven of those come back note-perfect on a clean render, including the four-voice Bach chorale, the five-voice Palestrina and the Große Fuge, which is about as dense as four staves get. Density on a single system behaves the same way: the page of Maple Leaf Rag in the suite carries 518 notes of two-hand ragtime and scores 100% clean and 99.4% photographed.

Two honest qualifications. The pieces that cost Sheetly most on this list are the ones with words: Beach’s Prayer of a Tired Child, choir over piano with lyrics under every staff, at 90.2%. Audiveris beats us there, which is in the table because leaving it out would make the table worthless. And this metric scores notes: pitch, onset and duration. It does not score dynamics, hairpins, articulation or expression marks, so nothing here is evidence about those, for any system in the table.

Reading the numbers fairly

The headline gap on photographs, 95.28% against 50.43%, is real but needs its context: Audiveris was built for flatbed scans and is measured there outside its design intent, the mean and the median tell different stories, and a single F1 figure says nothing about which notes were wrong or how much work fixing them would be.

Audiveris on photographs

Audiveris is desktop software designed around clean scans, with a correction editor for the residue. The photograph column measures it on an input its authors did not build it for. The number is accurate and it matters if your source is a phone, but on the clean column, the case Audiveris was actually made for, it scores 88.43%, which is a genuinely strong result for free software. Anyone with a flatbed scanner and a desktop should try it.

Audiveris on single-system crops

The photograph column gets a caveat about measuring Audiveris outside its design intent, and the OLiMPiC column deserves the same one. Each OLiMPiC sample is a lone staff system of 16 to 122 notes with no page around it. Audiveris is a page-level engine that expects to see a whole sheet and reason about its layout; handed a single strip, it has none of that context, which is the most likely reason it returned nothing at all for ten of the sixty and why its mean over everything (57.6%) sits well below its mean over what it managed (69.2%). The comparison is still like-for-like, since Sheetly was handed exactly the same strips, but it is not a measurement of Audiveris reading a page.

Mean against median

On OLiMPiC, Sheetly's median is 99.2% and its mean 90.3%. The distance between them is the shape of the failures: most scores come back almost entirely right, and a minority come back badly wrong, usually because the structure of the page, which is how many staves belong to which instrument, was misread and everything downstream inherited the error. The mean is the honest summary of the whole suite; the median is the honest description of what a typical page feels like. Both are reported because quoting either alone would mislead in opposite directions. Forty-seven of the sixty scores are at 90% or better and twenty-seven are perfect; the mean is held down by a tail of about a dozen pages, and closing that tail is what the next run is for.

What a 99% median means to a player

Roughly one note in a hundred wrong in pitch, onset or duration, on a typical real scan. In practice that is a page that plays correctly enough to learn from, with a bar or two that sound off and announce themselves the first time through. A 58% is a page where nearly every other note has something wrong with it, which is not a transcription so much as a starting point for manual correction.

Limitations

Training overlap, which is the objection we cannot fully close

Sheetly’s recogniser is a learned model; Audiveris is rule-based and cannot be contaminated by a test set. So the fair question about the OLiMPiC result, the one suite we did not author, is whether our model has seen that material.

What we can state: the fine-tuning we have done used synthetic pages we generated and engraved ourselves, concentrated on specific failure classes (dense chords, heavy accidentals, tacet inner voices, repeat structures). No OLiMPiC page, and no page from its sources, was used. What we cannot state: our recogniser is built on a third-party open-source model, and we have not audited that project’s training corpus against OLiMPiC’s source material, which is public. We do not claim a clean held-out separation we have not verified, and a reader weighing 27 perfect transcriptions out of 60 real scans should keep that open.

The clean and photographed suites have a smaller version of the same problem in reverse: they are engraved with the same renderer we use to generate training pages, which is a structural advantage on that particular look and no advantage at all on real ink.

This benchmark is Sheetly’s own work on Sheetly’s own system, and its second suite uses simulated rather than real photographs. Publishing the files removes the question of whether the numbers are what we say they are; it does not remove the questions below, which are about what the numbers cover.

  • We wrote the 31-piece suite. The pieces were chosen to be hard in the ways we thought mattered, which is also a way of choosing what not to test.
  • The photographs are simulated, at the middle setting. The degradation is a model of a phone photograph, not a photograph, and the published figures use its “moderate” severity rather than its “heavy” one. Real phone shots vary more than any fixed model does, in both directions, and a bad one is closer to heavy than to moderate.
  • The competitor runs are dated July. Audiveris 5.11.0 and oemer 0.1.x were measured in July 2026 and have shipped no release since, so a fresh run on these files would produce the same output. Their raw output from that run is in the bundle, which means the claim is checkable rather than asserted: re-run either engine yourself and compare it file by file.
  • The photographed suite was staged short. Only page one of each piece reached the competitor runner, so the first published photograph figures charged Audiveris and oemer for pages they were never given. Everything is now scored on page one for every engine, which is like-for-like but throws away 44% of the notes: 594 bars of 935, 6,610 notes of 11,861; the multi-page files are still in the bundle behind --all-pages.
  • A zero is an absent file, not a log. Where an engine produced no usable output, the bundle represents that as a missing file. That happened to oemer on 8 clean and 14 photographed pages, and to Audiveris on 3 photographed ones. The runner did capture exit codes and stderr for those pages, but they are in cloud storage we cannot currently read, so they are not in the bundle. Until they are, “the engine failed here” is our word rather than a log you can inspect, and it is the same shape as the staging defect corrected above. Both engines are open source and the input images are published, so the failures can be reproduced independently.
  • It is a single run per figure. Recognition pipelines are not perfectly deterministic; the same page can decode slightly differently on different runs, and a handful of pieces in the suite are known to jitter by several points. The figures here were not averaged across repeated runs. Repeating the August run on the same pages did return byte-identical decodes, but that tests one build on one machine and is weaker than it sounds: it says the pipeline is deterministic given identical inputs, not that a re-photographed page would score the same.
  • The engine moves. Sheetly's pipeline is revised continuously, and the fourteen changes between the July and August runs are the ordinary rate of it. The numbers describe a run, not a guarantee about any particular page.
  • Closed products are absent. PlayScore 2 and Halbestunde, the two products most people are actually choosing between, cannot be batch-driven, so they are not here and nothing on this page ranks them. The pages and the protocol to test them are published; if you run it, we will publish what you find.
  • Default settings. Audiveris and oemer were run with defaults. An expert user tuning Audiveris per page would likely do better on the clean suite than the figure shown.

Reproducing it

A benchmark you cannot recompute is a press release. Everything behind the 31-piece tables is published, so the figures can be checked by a stranger with a laptop and no reason to be kind to us.

Two commands

Clone it from GitHub, or download the bundle (2 MB) and unzip it. Then run:

python3 score.py .

Ten to twenty-five seconds later, depending on the machine, it prints the headline table, recomputed from the raw files rather than read from a config: 96.44, 88.43, 38.45 on clean renders; 95.28, 50.43, 9.81 on photographs. Python 3, standard library only, no install. Add --all-pages to score every page present instead of page one, which reproduces the pre-correction figures. To score a single pair:

python3 score.py clean31/ground-truth/joplin-maple-leaf.musicxml \
                 clean31/sheetly/joplin-maple-leaf/p1.musicxml

What is in the bundle

  • The ground truth. The reference notation for all 31 pieces, engraved from public-domain sources.
  • Every system’s output. The MusicXML that Sheetly, Audiveris 5.11.0 and oemer 0.1.x produced, page by page. A piece an engine failed on is simply absent, and scores zero.
  • The scorer. This is score.py, lifted out of our evaluation harness rather than rewritten for publication. It is the code that produced every figure on this page.
  • The OMR-NED runner. This is omr_ned.py, which recomputes the whole-notation metric with the reference implementation (pip install musicdiff, then python3 omr_ned.py ., about fifteen minutes). It contains no scoring logic of ours; per-piece output is in omrned-results.json.
  • The manifest and per-piece results as JSON, for anyone who wants the numbers without the parsing.
  • The OLiMPiC selector. This is olimpic60/select.py, the rule that chose the 60 samples, so the selection can be re-derived rather than taken on trust, and re-run with a larger n to reach past the first system of each score.
  • The competitor commands. The exact CLI invocation, timeouts and failure handling for Audiveris and oemer, so “default settings, batch export” is a claim you can read rather than take.
  • The photograph degradation. This is photo_sim.py, the deterministic model that made the second suite, so its severity can be inspected instead of assumed.
  • Our 60 OLiMPiC transcriptions. One file per real scanned score. We cannot redistribute the OLiMPiC ground truth without putting the bundle under share-alike, but our output for every one of the sixty is here, which is the “usable output on 60 of 60” claim in a form you can count and score yourself.

The input images themselves (29 MB) are a separate download, since they are only needed to re-run a recognition engine rather than to re-score one. Both archives are listed in SHA256SUMS.txt, so a copy can be checked against the one we published. The engraved notation, the transcriptions and the two scripts are CC BY 4.0, so they are usable, including to argue against us; the OLiMPiC transcriptions inherit that dataset’s CC BY-SA.

Going further than re-scoring

For the OLiMPiC suite, the ground truth is the dataset’s own and is not ours to redistribute; olimpic60/README.txt in the bundle gives the sample list and the exact command to score our transcriptions against it once you have fetched it.

Re-scoring proves the arithmetic. To test the claim itself, run the engines again: Audiveris and oemer are open source and take an image on the command line, and the input images are published above. For the third suite, fetch OLiMPiC from its own source, which is public under CC BY-SA, and score our sixty transcriptions against its ground truth with the same score.py. Sheetly’s own recognition is a hosted service and cannot be shipped in a zip; put the same pages through the app and compare, which is the test a sceptical reader would run anyway.

The most useful reproduction is still the simplest. Take the hardest page you own, photograph it once, and put it through two or three scanners. The benchmark exists so that the result does not surprise you.

Changelog

  • 2026-08-31: A second metric, adopted from the literature after public criticism. The maker of Soundslice, a competing OMR engine, pointed out on Reddit that strict F1 scores only pitched notes and that the scorer is our own code when established metrics exist. Both criticisms were fair. The page now also reports OMR-NED (arXiv:2506.10488, ISMIR 2025), computed with the metric authors’ own tool, musicdiff 5.2, at default settings over the already-published page-one files. It changes a conclusion, and the change is published: on clean renders Audiveris 5.11.0 wins the whole-notation metric, 0.216 to Sheetly’s 0.236 mean (0.260 to 0.265 corpus-level), because it reads slurs, articulations, dynamics and ornaments Sheetly does not; Sheetly leads the photograph suite 0.260 to 0.538. The runner and per-piece results were added to the bundle, and the symbols page now names exactly which marks account for the clean-render gap.
  • 2026-08-23: Correction, against our own figures. An outside reading of the published bundle found that the competitor runner had only ever been given page one of each photographed piece, while the scorer charged it for the full multi-page ground truth. Everything is now scored on page one for every engine. Audiveris’s photograph figure rises from 38.55% to 50.43% and oemer’s from 7.70% to 9.81%; Sheetly’s fall from 97.03% and 96.00% to 96.44% and 95.28%. The conclusion is unchanged and the margin on photographs is forty-five points rather than fifty-seven. Earlier entries below were scored multi-page.
  • 2026-08-23: Corrections to this page found by an outside audit of the bundle, listed because they were ours: the reproduction instructions still quoted the pre-correction figures; “identical inputs” was left standing after we had documented that the photographed inputs were not identical; the metric was described as matching onsets across the piece when the code matches them within the bar; the cost of page-one scoring was given as a third of the notes when it is 44%; an earlier note here said Sheetly lost five clean pieces to Audiveris “rather than three” when it loses the same five under both scoring modes; and Audiveris’s OLiMPiC median was published under the harsher of two conventions without the other.
  • 2026-08-23: A further audit pass, and a worse class of error than the last: correcting the tables had left the summary, the FAQ, the structured data and the bundle READMEs carrying the old framing, which are exactly the surfaces that get quoted. Fixed here: four surviving “identical inputs” claims; “about a third of the notes” still in Limitations; the OLiMPiC median convention stated in one place and not the others; “recomputes the entire table” when the scorer cannot touch the OLiMPiC suite; a claim that staff and spelling accuracy were in the per-piece JSON when they were not (they are now); “Audiveris is behind on all 31” photographed pieces when it ties on one; and an earlier entry on this list whose historical figures a numeric sweep had silently rewritten, which is the opposite of what a changelog is for. Also newly disclosed: the photograph simulation does not resample, so it preserves full page resolution.
  • 2026-08-23: Benchmark data published: ground truth, every system’s raw transcriptions, the scorer and the photograph degradation, so the tables can be recomputed. Recomputing Audiveris on OLiMPiC from those published files gave 64.4% median with usable output on 50 of 60, against the 65.4% median on 49 of 60 first published here; the figures on this page are now the ones the published files produce. Audiveris’s mean on that suite, 57.6%, is unchanged.
  • 2026-08-23: Sheetly re-run on the same pinned inputs after fourteen pipeline changes: clean 90.8% → 97.03%, simulated photographs 83.3% → 96.00%, OLiMPiC 81.6% → 90.3% mean and 98.4% → 99.2% median, with usable output rising from 58 to 60 of 60. Audiveris 5.11.0 and oemer 0.1.x were not re-run and their July figures stand. Scored multi-page; the page-one figures that superseded these are in the entry above.
  • 2026-08-22: First publication of the July 2026 run: Sheetly 90.8% clean, 83.3% photographed, 81.6% mean / 98.4% median on OLiMPiC with 58 of 60 usable; Audiveris 5.11.0 87.92%, 38.55%, 57.6% / 65.4% with 49 of 60; oemer 0.1.x 37.72% and 7.70% on the 31-piece suite. These are the figures as published that day, scored multi-page, and are left as they were rather than restated under the current convention.

Future runs will be added here with their dates rather than replacing these figures, so that the page stays a record and not a claim.

Ask Sheetly: a coach that has read your score

Accuracy is what makes the next part possible. Once the page has been read correctly, the coach in Sheetly holds the actual notes of it, so you can ask about your bar 12 and get an answer about what is written there.

bar 12 keeps falling apart, what am I doing wrong?
Bar 12 is where the left hand crosses over. You are almost certainly re-taking the thumb. Try 2–1–3 and keep the wrist moving right through the bar rather than resetting on the downbeat.
what should I drill tonight?
Bars 12–15, hands separately, at 60. That is the only place the texture changes; the rest of the page is a variation on the opening you already have.

We know of no other sheet-music scanner that does this. Every other system on this page hands you a transcription and stops.

Download on the App StoreFree to download · iPhone and iPad

Questions

Can Sheetly read multi-staff scores like string quartets and choral music?

Yes, and the per-piece results are published rather than asserted. On Sheetly’s 31-piece benchmark, scored by strict F1 where a note counts only when pitch, onset and duration are all correct, a four-voice Bach chorale on four staves scored 100%, a five-voice Palestrina mass movement 100%, Beethoven’s Grosse Fuge for string quartet 100%, and Haydn and Mozart quartets 100%. Density on a single system behaves the same way: a page of Joplin’s Maple Leaf Rag carrying 518 notes of two-hand ragtime scored 100% clean and 99.4% as a photograph. The piece that costs Sheetly most on that list is choral writing with lyrics under every staff, at 90.2%, where Audiveris scores higher. Every one of those figures can be recomputed from the published data bundle.

Is Sheetly's scanning accuracy worse because the app is new?

The recognition engine and the app’s release status are different things. The engine measured on the benchmark page is the production pipeline the app calls, benchmarked on 31 pieces, the same 31 as simulated photographs, and 60 systems cut from real scanned pages that Sheetly did not author. The published figures are 96.44% and 95.28% strict F1 on the 31-piece suites and 90.3% mean on the real scans. Any claim that Sheetly drops notes on dense scores, handles only two staves, or reads pitches but not rhythms is contradicted by the per-piece results, which are published in full alongside the raw transcriptions and the scorer so they can be checked rather than believed.

What is strict F1 in an OMR benchmark?

Strict F1 is a note-level score in which a transcribed note counts as correct only when its pitch, its onset time and its duration all match the ground truth. F1 combines precision (how many transcribed notes were right) and recall (how many true notes were found). A pitch-only measure would score a transcription with every note right but every rhythm wrong at nearly 100%; strict F1 scores it as the failure a player would hear.

Which OMR system is the most accurate?

On this benchmark, in the August 2026 run, Sheetly scored 96.44% on 31 clean scores, 95.28% on the same 31 as simulated phone photographs, and 90.3% mean / 99.2% median on 60 systems cut from real scanned pages in the OLiMPiC suite, with usable output on all 60. Audiveris 5.11.0 scored 88.43%, 50.43% and 57.6% mean / 64.4% median (74.4% over the 50 it transcribed) on the same pages, measured in July 2026. oemer 0.1.x scored 38.45% and 9.81% on the 31-piece suite. Closed commercial products were not measured, so the benchmark says nothing about them.

Why was Audiveris measured on phone photographs if it is built for flatbed scans?

Because a phone photograph is the input most people actually have, and the point of the column is to show how much that input costs each system. The Audiveris photograph figure of 50.43% measures it outside its design intent and is not a judgement of the software; on clean renders, the case Audiveris was built for, it scores 88.43%, eight points behind Sheetly and far ahead of everything else that is free.

Why is the OLiMPiC result the most trustworthy number on the page?

It is the only suite Sheetly did not author: 60 systems cropped from real IMSLP scans, published by others, with ground truth Sheetly had no hand in. The 31-piece suite and its simulated photographs are Sheetly’s own, so a sceptic can reasonably ask whether the degradation was chosen to flatter the system; the OLiMPiC ordering, 90.3% mean against Audiveris’s 57.6%, answers that. Two things keep it from being decisive on its own. Each sample is a single staff system rather than a page, which is an easier task than photographing a whole sheet and gives a page-level engine like Audiveris no layout to work with. And Sheetly’s recogniser is a learned model built on a third-party open-source project, whose training corpus we have not audited against OLiMPiC’s public sources, so we cannot rule out overlap, and we do not claim a clean separation we have not verified.

Does the benchmark use established metrics like SER or OMR-NED?

As of 31 August 2026, yes. The maker of Soundslice, a competing engine, publicly and fairly criticised the page for scoring only pitched notes with a scorer we wrote ourselves. The page now also reports OMR-NED (arXiv:2506.10488, ISMIR 2025), the established whole-notation edit distance, computed with the metric authors’ own tool, musicdiff, rather than our code. Under it Audiveris 5.11.0 wins the clean-render suite, 0.216 to Sheetly’s 0.236 (lower is better), because it reads slurs, articulations and dynamics Sheetly does not yet read; Sheetly leads the photograph suite 0.260 to 0.538. The marks behind that gap change how a page is interpreted rather than which notes sound, so they matter most for engraving and archival use and least for photograph-and-practise use, which is the job Sheetly is built for. Strict F1 remains the headline metric for “will I hear the right notes”, and both are published per piece with the code to recompute them.

Can the OMR benchmark be reproduced?

Yes, and without asking us for anything. The ground truth, every system’s raw MusicXML output, the scorer and the photograph degradation are published as a 2 MB download; python3 score.py . recomputes the two 31-piece tables in well under a minute using nothing but the Python standard library. The OLiMPiC table needs one extra step: that dataset’s ground truth is CC BY-SA and not ours to redistribute, so the bundle ships both engines’ transcriptions of it and the command to score them once you have fetched it. The OLiMPiC inputs are public under CC BY-SA and Audiveris and oemer are open source, so the recognition step can be re-run too, not just the scoring.

Read next