A broadcast standard specifies how to blind-test audio
ITU-R’s MUSHRA recommendation gives a small studio a tested method for comparing generated and human tracks without fooling itself.
- First source published
- February 1, 2015
- Site publication
- September 18, 2026

What happened
The International Telecommunication Union’s broadcasting standards arm maintains two recommendations for subjective listening tests, both periodically revised: BS.1116, for detecting small impairments, most recently approved in February 2015, and BS.1534, the “MUlti Stimulus test with Hidden Reference and Anchor” method – MUSHRA – for intermediate audio quality, most recently approved in October 2015. MUSHRA exists specifically because BS.1116’s method, the recommendation itself notes, “is not suitable for assessing systems with intermediate audio quality,” which is closer to where a comparison between a generated track and a human recording usually sits than the near-transparent differences BS.1116 targets.
What the documents say
The MUSHRA recommendation specifies a structure a home listening test typically skips: every trial includes the unprocessed original as a hidden, unlabelled reference alongside the test items, plus at least one mandatory “hidden anchor” of known, deliberately degraded quality, so a listener’s ratings can be checked against a known floor and ceiling rather than trusted at face value. It also sets a post-screening rule: an assessor should be excluded from the results if they rate the hidden reference below 90 out of 100 more than 15 percent of the time, since that suggests they cannot reliably identify the original when it is disguised among the test items. The companion BS.1116 recommendation confirms the two methods are deliberately scoped for different jobs, not competing versions of the same test.
Why it matters for makers
The mechanism the standard protects against is expectation bias: a listener who knows which track is “the AI one” rates it differently than the same audio heard blind, and a listener never shown the real original has no anchor for how far “close to the original” actually stretches. A producer comparing a generated stem or a mastering assistant’s output against a human reference is running exactly this kind of test, informally, usually without hiding which is which. The music information retrieval research community publishes its own reproducibility resources separately from these broadcast standards.
What to check before you use it
This is an editorial adaptation of a broadcast-engineering standard for a small studio, not a substitute for a properly resourced listening test. Include the unprocessed original in the comparison set, unlabelled, alongside whatever is being evaluated. Add at least one deliberately degraded version as a low anchor, so “good” and “bad” have fixed points on the scale rather than only relative ones. Ask listeners to identify the hidden original among the set before trusting their other ratings, since MUSHRA’s own screening rule treats failure to do so as grounds to discard that listener’s data.
- Is the unprocessed original included in the comparison, hidden and unlabelled?
- Is there a known, deliberately degraded anchor to calibrate what “poor” means on the scale?
- Can each listener correctly pick out the hidden original before their other ratings are trusted?
None of this requires broadcast-lab equipment, only the discipline of the method: hide identities, include a known reference and a known anchor, and screen out listeners who cannot tell the difference before drawing a conclusion from the ones who can.
Sources & reading trail
Specifies the MUSHRA method's hidden reference, mandatory hidden anchor, and post-screening rule for excluding unreliable listeners.
Source published: 1 October 2015 · Retrieved: 16 September 2026
Confirms BS.1116 and BS.1534 are scoped for different quality ranges, small impairments versus intermediate quality.
Source published: 1 February 2015 · Retrieved: 16 September 2026
Establishes ISMIR as the research body for music information retrieval and its published resources on reproducible research.
Source published: Not established · Retrieved: 16 September 2026
Papers, terms and official documents establish the record; the maker reading and the checks are Signal to Song editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- Two decades of shared MIR evaluation still exclude generation
- A shared dataset made separation scores comparable, not perceptual
- A generation metric adapted from images admits a moderate correlation
- Browse the complete the archive
Sources & reading trail
- BS.1534: Method for the subjective assessment of intermediate quality level of audio systems
Source published: October 1, 2015 · Retrieved: September 16, 2026 - BS.1116: Methods for the subjective assessment of small impairments in audio systems
Source published: February 1, 2015 · Retrieved: September 16, 2026 - International Society for Music Information Retrieval
Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.