Sheetly OMR benchmark — inputs, outputs and scorer
https://www.trysheetly.com/compare/omr-benchmark

This archive contains everything needed to recompute the headline figures on
that page without trusting us: the ground-truth notation, the raw MusicXML
that each recognition system produced, and the scorer that turns the two into
a number.

    python3 score.py .

takes ten to twenty-five seconds on a laptop, needs no dependencies beyond the Python
standard library, and prints:

    == clean31 ==
      sheetly    strict F1  96.44%   median(all) 98.8%  median(transcribed) 98.8%  (31 of 31)
      audiveris  strict F1  88.43%   median(all) 93.6%  median(transcribed) 93.6%  (31 of 31)
      oemer      strict F1  38.45%   median(all) 44.3%  median(transcribed) 52.1%  (23 of 31)
    == photo31 ==
      sheetly    strict F1  95.28%   median(all) 97.4%  median(transcribed) 97.4%  (31 of 31)
      audiveris  strict F1  50.43%   median(all) 51.4%  median(transcribed) 59.8%  (28 of 31)
      oemer      strict F1   9.81%   median(all)  0.0%  median(transcribed)  6.9%  (17 of 31)

which are the numbers published on the page. To score a single pair:

Only page one of each piece is scored, because that is the only page every
engine received: the photographed suite was staged to the competitor runner as
page one per piece. Scoring their later pages would charge Audiveris and oemer
for ground truth they were never shown. `python3 score.py . --all-pages` scores
every page present and prints the earlier, multi-page figures instead.

What strict F1 covers: pitch, onset and duration of pitched notes. What it does
not cover: rests, staff/hand assignment, enharmonic spelling, clefs, key and
time signatures, repeats, slurs, articulations, dynamics, lyrics. Those
omissions apply identically to every engine here.

    python3 score.py clean31/ground-truth/joplin-maple-leaf.musicxml \
                     clean31/sheetly/joplin-maple-leaf/p1.musicxml

Layout
    manifest.json                       pieces, pages and measures per page
    score.py                            the scorer, lifted from our harness
    photo_sim.py                        the phone-photograph degradation, for
                                        inspection: it is the exact model used,
                                        at severity "moderate", but re-making the
                                        suite from it needs numpy and OpenCV and
                                        the clean page renders, which are in the
                                        separate 29 MB image download
    COMPETITOR-RUNS.txt                 the exact commands the competitors were run with
    benchmark-results.json              per-piece scores for every engine and suite
    clean31/ground-truth/<piece>.musicxml
    clean31/<engine>/<piece>/p<N>.musicxml
    photo31/…                           same pieces, phone-photograph degradation

OMR-NED (added 2026-08-31)
    After public criticism from the maker of a competing engine — that strict
    F1 ignores everything but pitched notes, and that the scorer is our own
    code — the bundle also carries the established whole-notation metric,
    OMR-NED (arXiv:2506.10488, ISMIR 2025; lower is better, 0.0 = perfect):

        pip install musicdiff        # the reference implementation
        python3 omr_ned.py .         # ~15 min; OMR_NED_WORKERS=8 to speed up

        clean31  mean(all, missing=1.0)  sheetly 0.236  audiveris 0.216  oemer 0.734
        photo31  mean(all, missing=1.0)  sheetly 0.260  audiveris 0.538  oemer 0.944

    On clean renders Audiveris wins this metric — it reads slurs,
    articulations, dynamics and ornaments that Sheetly does not (the exact
    list: trysheetly.com/guides/supported-music-symbols). On photographs
    Sheetly leads under both metrics. omr_ned.py contains no scoring logic
    of ours: every comparison is musicdiff at default settings
    (github.com/gregchapman-dev/musicdiff); per-piece numbers are in
    omrned-results.json.

Engines and dates
    sheetly     production recognition pipeline, run 2026-08-22
    audiveris   5.11.0, default settings, batch export, run 2026-07-15
    oemer       0.1.x,  default settings, run 2026-07-15

Both competitor versions are the current releases and are unchanged since they
were run.

Inputs were byte-identical on the clean suite. On the photographed suite they
were not: the competitor runner was staged with page one of each piece only, so
Sheetly saw second pages the competitors never received. Scoring page one for
every engine is what makes the comparison like-for-like; see COMPETITOR-RUNS.txt.

Sheetly's own runs have the LLM and vision passes switched off, which is why
every file it produced carries "LLM/vision passes are disabled in the offline
pipeline". Those passes read the title block and guess the instrument; they do
not read notes, and every figure here is note-level.

A piece an engine could not transcribe at all is scored zero rather than
dropped, which is why the means differ from the average over the files present.

The third suite on the page is 60 systems cut from real scanned pages, from the
public OLiMPiC dataset (CC BY-SA). Each sample is one staff system lifted out of
a scanned page, not a whole page. We do not redistribute the dataset itself,
because that would put this bundle under share-alike. What is here, in
olimpic60/, is both engines' output — ours for all 60, Audiveris 5.11.0's for
the 50 it managed — plus samples.txt listing the 60 sample ids in selection
order. Fetch OLiMPiC from its own source and score these against its ground
truth with the same score.py; see olimpic60/README.txt.

LICENCE
    The notation in clean31/photo31 is engraved from public-domain sources, and
    those files, the transcriptions and the two scripts are released CC BY 4.0 —
    use them, including to argue against us, with attribution.

    olimpic60/ is different. Those are our transcriptions of pages from OLiMPiC,
    which is CC BY-SA 4.0, so treat anything derived from them as share-alike;
    that is also why the dataset's own ground truth is not redistributed here.
