Open models show two ways to disclose training data
Stable Audio Open and MusicGen name different amounts about their training data, against what the EU’s new code asks providers to describe.
- First source published
- July 22, 2024
- Site publication
- September 18, 2026

What happened
Two openly released music-generation models have published different levels of detail about what trained them. Stability AI’s announcement of Stable Audio Open, dated 22 July 2024, states the model was trained on 472,618 recordings from Freesound and 13,874 tracks from the Free Music Archive, all under Creative Commons licences, with samples screened using “the PANNs audio tagger” and sent to “Audible Magic’s content detection company” to remove potential copyrighted music. Meta’s documentation for MusicGen states training on “20K hours of licensed music,” combining an internal 10K-track dataset with data from “ShutterStock and Pond5,” while stating plainly the underlying datasets are not released.
What the documents say
Read together, the two disclosures name different things. Stability’s announcement names specific source platforms, exact recording counts and a copyright-screening method – a reader could, in principle, go and inspect Freesound’s own licence terms. Meta’s MusicGen documentation names commercial licensors and an hours figure but not the tracks themselves. The EU’s Code of Practice for General-Purpose AI Models, Transparency Chapter, published 10 July 2025, sets out what its Model Documentation Form asks a signatory to describe: “the scope and main characteristics of the training, testing and validation data, such as domain..., geography..., language, modality coverage,” plus any “measures to detect unsuitability of data sources.” Several of those fields are marked as owed to the EU AI Office and national authorities rather than to the public, meaning the code’s detail does not automatically become a reader-facing disclosure.
Why it matters for makers
The mechanism here is that “trained on licensed data” is a claim with a range of possible evidence behind it, from a named platform and screening method to a vague assurance. A producer weighing rights risk in a generated track is asking how traceable the claim is: can it be checked against an independent record, or does it rest on the provider’s word about an unreleased internal set? The EU code’s category list is useful even where it stays private, because it names exactly the questions any public disclosure should answer.
What to check before you use it
This is an editorial reading of the cited documents, not a rights clearance. Look for named sources rather than adjectives: a dataset name and a link a person could independently verify carries more weight than the word “licensed” alone. Check whether a screening method for unsuitable content is described, and whether it was applied before or after collection. Where a model card gives no source detail at all, treat any downstream rights claim as unverified rather than as cleared.
- Does this model name specific datasets or platforms, or only describe them in general terms?
- Is there a stated method for screening out copyrighted or unsuitable material?
- If a rights question arose later, is there anything here a maker could actually go and check?
Neither disclosure proves the data was properly licensed beyond what each company states; both are self-reported. But the gap between naming a platform and citing an hours figure is informative, and it is the gap the EU’s categories are designed to close once, and if, they reach the public.
Sources & reading trail
States the named training sources (Freesound, Free Music Archive), recording counts, licence types and the Audible Magic screening method.
Source published: 22 July 2024 · Retrieved: 16 September 2026
States MusicGen's training data comprised 20K hours of licensed music including an internal dataset and ShutterStock/Pond5 data, none of which is released.
Source published: Not established · Retrieved: 16 September 2026
Sets out the categories of training-data description (domain, geography, unsuitability screening) the Model Documentation Form asks signatories to complete, and which recipients receive them.
Source published: 10 July 2025 · Retrieved: 16 September 2026
Papers, terms and official documents establish the record; the maker reading and the checks are Signal to Song editorial analysis. This retrospective draft does not imply the site published on the event date.
Continue reading
- EU law now makes AI model makers disclose training summaries
- Stable Audio switched from a licensed library to Creative Commons
- Meta open-sourced MusicGen with a licensed dataset and an NC licence
- Browse the complete the archive
Sources & reading trail
- Stable Audio Open: Research Paper
Source published: July 22, 2024 · Retrieved: September 16, 2026 - MusicGen
Retrieved: September 16, 2026 - Code of Practice for General-Purpose AI Models: Transparency Chapter
Source published: July 10, 2025 · Retrieved: September 16, 2026
The documents above establish the record. The reading and the questions are this publication’s editorial analysis, written after the fact.
Published September 18, 2026, not on the date of the event described.