All Articles
Sean Power

Why Music Publishers Need a Data Strategy for AI Licensing Before It Is Too Late

music publishers
Why Music Publishers Need a Data Strategy for AI Licensing Before It Is Too Late

AI developers are approaching music publishers for training licenses. Some of these conversations are happening now, quietly, without public announcement. The terms being discussed vary widely because the market has no established rate card, no precedent license structure, and no settled legal framework governing what the license is actually for.

In that environment, negotiating position depends almost entirely on information asymmetry. A publisher that knows exactly which recordings in their catalog appear in which training datasets, with documentation to support that knowledge, is negotiating from a fundamentally different position than one that cannot answer those questions. The difference is not a legal advantage; it is a data advantage. And data advantages are built before negotiations begin, not during them.

The Catalog Transparency Problem

Most music publishers have a clearer picture of their catalog than they did a decade ago, but "clearer" is not the same as "complete." Catalog acquisitions compound the problem: each acquisition brings its own metadata conventions, ISRC registration completeness, and ownership chain documentation quality. A publisher that has grown through acquisition often has a catalog that is, in practice, a federation of different catalogs with different levels of internal coherence.

For streaming royalties, the gaps in catalog data are a known problem with established remediation processes. Missing ISRCs can be registered retroactively. Unmatched works can be resolved through the MLC's matching process. The cost of gaps is delayed or unclaimed royalties on a per-play basis, which accumulates but does not create a single large exposure event.

For AI training licensing, the exposure structure is different. Training is a point-in-time event, not a continuous use. If your catalog was used in a training run that happened eighteen months ago and you do not have documentation of that use, the window for establishing a license and seeking compensation for that specific training run is closing. It is not infinite. The ability to identify what was used depends on tools that need to operate while the training event is still analytically traceable, not years after the fact when competing interests have had time to establish different narratives about what happened.

What a Data Strategy Actually Means in Practice

A data strategy for AI licensing is not a new initiative requiring a separate team and a multi-year project. It is a set of decisions about catalog data quality that have immediate operational value and that also happen to be the prerequisites for AI attribution analysis.

The three core components are: ISRC completeness, ownership chain clarity, and training exposure documentation.

ISRC completeness means every recording in the catalog has a registered ISRC that is linked to rights holder identifiers in a format that attribution analysis systems can query. Gaps in ISRC registration are attribution analysis gaps; a recording without an ISRC cannot be returned as a match result because there is nothing to match against in the database.

Ownership chain clarity means the rights split for every work in the catalog is documented and current, with co-publisher shares, songwriter splits, and any controlled composition clauses that affect royalty calculations. For AI training licensing, where the royalty basis is still being defined, clear ownership chains allow any eventual payment to flow correctly rather than sitting in suspense while splits are resolved.

Training exposure documentation means having a systematic approach to identifying which of your catalog recordings appear in publicly disclosed or discoverable training datasets. This is where attribution analysis tools become relevant: they provide the mechanism for building that documentation rather than relying on AI developers to disclose it voluntarily.

The Negotiation Asymmetry in Detail

Consider two hypothetical scenarios for how a training license negotiation might proceed. In the first scenario, an AI developer approaches a publisher for a retroactive training license. The developer offers a flat fee for use of the publisher's catalog. The publisher has no independent data on which recordings were actually in the training set, how heavily they contributed, or what the market rate for that contribution should be. The negotiation is essentially about accepting or declining the developer's framing of the situation.

In the second scenario, the same developer approaches a publisher that has run attribution analysis on the developer's model outputs and has a documented list of catalog recordings that appear as training contributors, with confidence scores and relative contribution weights. The publisher knows the composition of the contribution from their catalog. They can quantify it against the total training corpus size and construct an argument about fair compensation based on that quantification. The conversation is different because the information basis is different.

We are not saying either scenario produces a legally certain outcome. The legal framework for what training licensing should look like is unresolved. But the publisher in the second scenario has a factual foundation for the negotiation that the publisher in the first scenario does not. That foundation is worth building regardless of which legal theory ultimately prevails, because it provides optionality: the data can support a licensing negotiation, a fair use challenge, a contribution to an industry collective negotiation, or documentation for eventual litigation if the legal framework evolves in that direction.

The Window Is Not Permanent

The technical capacity to trace training attribution is strongest closest to the training event. Attribution signals in model outputs are measurable now. The question is whether publishers will build the documentation while the window is open or wait until external pressure forces action when the signals have weakened or the legal landscape has been defined by others.

There is a reasonable objection to acting before the legal framework is clear: why invest in documentation when you do not know what you will be able to do with it? The answer is that the documentation is an option, not a commitment. Building a catalog of training exposure data does not commit a publisher to any particular legal strategy. It preserves choices. Not building it closes those choices.

The publishers who will be in the strongest position when AI training licensing settles into a defined market are the ones who started building their data infrastructure when the market was forming. That process is happening now. The data strategy work is not difficult; it is a matter of priority. Treating catalog accuracy as a prerequisite for AI licensing negotiation, rather than a background administrative task, is the change that matters.

Stay Informed

Get perspectives on AI music attribution and rights infrastructure.

Request API Access