Choose material neither tool trained on
A benchmark starts with a declared test set. If a system has seen the song, stems, alternate mix or a near duplicate during training, its score no longer answers the question of generalisation. The 2018 Signal Separation Evaluation Campaign introduced MUSDB18 specifically as a common music-separation dataset and released evaluation tooling. That shared setup makes comparisons more meaningful than a vendor's hand-picked before-and-after clip.
For a small editorial comparison, do not claim hidden training knowledge you cannot verify. Instead say what you know: the tracks are from a public held-out set, the paper reports its train/test split, or the tool's training data is undisclosed. Add a second, rights-cleared practical set that resembles the intended job, but keep it separate from the benchmark result.
Report the metric with its definition
Signal-to-distortion ratio (SDR) is commonly used to quantify how closely an estimated stem resembles a reference stem; higher is not a universal definition of ‘better.’ Report the version, aggregation method, source type and whether numbers are medians or means. The SiSEC paper describes MUSDB18 and its evaluation setting, while later work shows that listener judgments can line up differently across instruments and metrics.
Never collapse several measurements into a scorecard without showing the ingredients. A vocal extraction may score well overall yet be awkward for dialogue replacement because breaths, consonants or reverb tails matter to that task. Conversely, a noisier isolated stem may be acceptable as a quiet practice aid. The metric answers a signal question; the job defines which defects matter.
Add a blind listening pass
Prepare matched excerpts, normalise playback level, randomise tool labels and ask listeners to rate a stated purpose: vocal intelligibility, drum-transient preservation, karaoke usefulness or remix editability. Let listeners select ‘cannot judge’ rather than forcing a preference. Write the same questions before listening to avoid quietly changing the standard after hearing a favourite result.
Log difficult material separately: dense backing vocals, distorted guitars, live ambience, abrupt edits and very short clips. These examples are not proof that one system universally wins; they explain where the aggregate score stops being enough. Keep the original mix and the exact generated stems so another listener can retrace the result.
Publish limits with the result
A fair conclusion sounds like: ‘On this declared set, using this metric and this listening question, Tool A was preferred for vocals; the finding does not establish legal permission to use the separated material.’ Separation changes access to audio; it does not transfer rights in the underlying recording. This keeps performance evaluation distinct from the archive's wider rights guidance.
If you cannot run the test, do not invent a ranking from product claims. Read papers for their dataset, split, metric and listening design, then use that information to decide what evidence would be needed for your own task.
Continue in Listening & research, or Open source made stem separation common but not a rights grant.
Evidence & further reading
Go to the source.
Primary references checked 19 September 2026. Practical exercises are editorial guidance, not reported product tests.
- The 2018 Signal Separation Evaluation Campaign ↗
Primary campaign paper describes MUSDB18 and common evaluation tooling for source separation.
arXiv · Source date 2018-04-17 · Accessed 2026-09-19 - Musical Source Separation Bake-Off: Comparing Objective Metrics with Human Perception ↗
Primary study compares objective separation metrics with listener ratings across source types.
arXiv · Source date 2025-07-09 · Accessed 2026-09-19 - 2019 EUSIPCO proceedings paper on source separation evaluation ↗
Primary proceedings source documents MUSDB18 and SDR reporting in a separation comparison.
EURASIP · 2019 proceedings; exact publication day not established · Accessed 2026-09-19
