Sheetly and Audiveris on 60 OLiMPiC systems

OLiMPiC 1.0-Scanned (ICDAR 2024, CC BY-SA) pairs per-system crops of original
IMSLP library scans with MusicXML ground truth. Each sample is one staff system
cut out of a real scanned page — not a whole page. We report 60 of them: the
first system of page one from each of 60 different scores.

The selection is deterministic and not ours to tune. sample ids are sorted, then
picked round-robin, one system per score; taking the first 60 always yields
these 60. Because 60 is exactly one full round, every sample is the FIRST system
of its score — the strip that carries the clef, key and time signature, and never
a page interior. That is a real bias and it is worth knowing before quoting the
figure. The list is in samples.txt and the selector itself is select.py, so both
the rule and its output can be checked.

WHAT IS HERE
    sheetly/<score>-<system>.musicxml     our transcription, all 60
    audiveris/<score>-<system>.musicxml   Audiveris 5.11.0's, the 50 it produced
    samples.txt                           the 60 sample ids, in selection order

We do not redistribute OLiMPiC's ground truth, because that would put this
bundle under share-alike. Fetch the dataset from its own source, then:

    python3 ../score.py <olimpic>/samples/<score>/<system>.musicxml \
                        sheetly/<score>-<system>.musicxml

and read strict_f1. Doing that for all 60 gives the published figures:

    Sheetly     90.3% mean · 99.2% median · usable output on 60 of 60
    Audiveris   57.6% mean · 64.4% median · usable output on 50 of 60
                69.2% mean · 74.4% median over only the 50 it transcribed

A system an engine produced nothing for scores zero and stays in the mean; that
is why Audiveris's mean over all 60 (57.6%) is below its mean over the 50 it
managed (69.2%). Per-sample figures for both engines are published as JSON at
https://www.trysheetly.com/benchmark/benchmark-results.json
