RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The archive · 100 retrospective records ↗
← The archive

A shared dataset made separation scores comparable, not perceptual

MUSDB18 and the SDR metric let separation systems be ranked by decibels, which the campaign's own paper says diverges from listening preference.

Historical event
December 17, 2017
First source published
December 17, 2017
Site publication
September 18, 2026
Visual for this record: A shared dataset made separation scores comparable, not perceptual
Visual published by fnzhan.com, shown for identification of the record. Credit: fnzhan.com · source page ↗ Rights: owner-review-pending. Source

What happened

On 17 December 2017, Zafar Rafii, Antoine Liutkus, Fabian-Robert Stöter, Stylianos Ioannis Mimilakis and Rachel Bittner published MUSDB18, a dataset of 150 full songs with isolated drum, bass, vocal and 'other' stems, on Zenodo. The dataset, split into 100 training and 50 test tracks, became the shared benchmark for music source separation research described on the SigSep project page. It underpinned the 2018 Signal Separation Evaluation Campaign, SiSEC, whose paper states the campaign itself dates to 2008 and this was its first release built for training data-hungry systems, not only for testing them.

What the documents say

The SiSEC paper explains the standard scoring tools, the BSS Eval metrics, assess separation through 'Source to Distortion, to Artefact, to Interference ratios', SDR, SAR and SIR, computed after optimally aligning an estimated stem to the true one through a matching filter. A higher SDR indicates less distortion energy relative to the target signal, not a direct perceptual rating. The paper states this explicitly: 'BSS Eval scores are in fine relative to squared-error criteria', and reports that earlier 'perceptual studies showed' a method preferred by human listeners could still score lower on SAR than a method they liked less, because the metric optimises for a mathematical fit, not a listening preference.

Why it matters for makers

The mechanism is signal-level error measured in decibels against a known ground truth, which is exactly why it is useful for comparing separation systems on the same held-out songs: no listening panel is required, and the number is reproducible. But a separation plug-in advertising a leaderboard SDR score is quoting a distortion-energy measurement on MUSDB18's specific genres and mix styles, not a promise about how clean the vocal stem will sound on your particular mix, especially one outside the pop, rock and electronic material the dataset weights toward.

What to check before you use it

When a separation tool cites an SDR figure, check whether it was measured on MUSDB18 or a different, possibly easier, test set, since scores are not comparable across datasets. Listen to the stem on your own material rather than trusting a decibel figure alone; this is an editorial recommendation the papers themselves support, given their own finding that SDR and human preference can diverge. Note MUSDB18's licence permits non-commercial, educational use, which matters if a tool's benchmark claim rests on training against it.

MUSDB18 and SDR gave source separation a shared, reproducible yardstick, which most audio tasks still lack. The campaign's own paper is the clearest source for the limits of that yardstick: a decibel gain is not the same claim as a listener's preference.

Sources & reading trail

Publishes the dataset with its release date, authors, licence terms and track/stem composition.

Source published: 17 December 2017 · Retrieved: 16 September 2026

Describes the dataset's train/test split, stem categories and intended use for supervised separation research.

Source published: Not established · Retrieved: 16 September 2026

Defines SDR/SAR/SIR via BSS Eval, dates SiSEC to 2008, and states BSS Eval scores can diverge from perceptual studies.

Source published: 6 July 2018 · Retrieved: 16 September 2026

Papers, terms and official documents establish the record; the maker reading and the checks are Signal to Song editorial analysis. This retrospective draft does not imply the site published on the event date.

Continue reading

Sources & reading trail

The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.

Published September 18, 2026, not on the date of the event described.