Executive Summary
This report asks how data owners, meaning musicians, journalists, authors and visual artists, can equitably participate in the AI economy, given that their content was used to train frontier AI models from companies with near-trillion-dollar valuations, largely without permission or compensation. Over 10 weeks, we conducted 83 interviews across the music industry, news publishing, AI research, intellectual property law, and the creator community.
We set out to test three hypotheses:
- (1) Data owners are not fairly compensated when AI labs train on their copyrighted work. Labs train without permission, without payment, or through deal structures that systematically undervalue individual creative contributions.
- (2) Membership inference techniques can detect whether a specific work was used to train a given AI model. They probe a trained model to estimate whether a piece of content was part of its training data: “was my data used to train this AI?”
- (3) Data attribution techniques can measure what each work contributed to a model, and split revenue accordingly. They estimate how much each piece of training data shaped a model’s behaviour: “what is my data worth for AI training or fine-tuning?”
1. Data owners without negotiating power are not compensated — confirmed, though the verdict differs sharply by industry
It ultimately hinges on the unresolved legal question of fair use. In music, the major labels used concentrated catalog ownership and litigation leverage to extract licensing-and-equity settlements from AI music startups; independent artists were left out of those settlements. In books, courts have so far found that training on copyrighted content can be fair use, but signaled that a stronger showing of market harm could have changed that result; authors’ recoveries have come from the provenance of the data (piracy), not from training. In the news category, the marquee cases are still being fought. Across every industry, the dividing line between who gets paid and who does not seems to be the negotiating power of the content owner. Whether the system is “unfair” in a legal sense, rather than an economic one, will be decided by the courts’ answer to the fair use question, which remains open.
2. Membership inference detects use at catalog scale — partially confirmed; on its own, detection is not a legal claim
Technically, detecting whether content was included in a model’s training data is meaningfully more tractable at the level of a whole catalog than for a single work. Detection is also straightforward once rights holders gain direct access to the training dataset itself, for example through litigation discovery. The hardest open problem is technical rather than commercial: building litigation-grade, fully black-box detection, meaning proof of training use without any access to the model’s internals or its dataset. Legally, a positive detection result does not, by itself, carry a claim. The fair-use defense rests on training being transformative; when a model instead reproduces its training data verbatim, that transformation is absent, so a positive detection result becomes actionable when paired with evidence that the model’s outputs substitute for the original work and suppress demand for it.
For news publishers, detection has to cover two routes, not one. Retrieval matters as much as training, and it is the half that is growing. Training can only ever contain what a model saw before its cutoff, but models now reach today’s reporting at answer time through retrieval. That is precisely where a publication’s value is highest and its window shortest. An answer assembled from this morning’s coverage substitutes for the article most directly, because the reader who wanted the latest has already been served. A detection system built only for training use would miss the substitution that costs news publishers the most.
3. Data attribution measures each work’s contribution in fine-tuning and retrieval — partially confirmed, but not at pre-training scale
On the technical front, influence functions, one of the leading academic methods for attributing a model’s behavior to individual training examples, don’t scale to trillion-parameter models, and the marginal value of any individual work in a pretraining corpus of hundreds of billions of tokens is commercially negligible. The fine-tuning layer is different: datasets are orders of magnitude smaller, individual contributions are measurable, and the economics of per-work compensation become more meaningful. Commercially, adoption of contribution-based attribution revenue depends on how concentrated an industry is and on where its content sits in the AI stack. Concentrated industries (music’s three major labels) negotiate based on market share and have little reason to embrace contribution-based attribution, while fragmented ones (news) lack the collective machinery to push for it. The content that matters most in pre-training (music, books) is exactly where per-work value is negligible, whereas the content that matters most in retrieval and post-training (news) is where per-use attribution is most viable.
Findings 2 and 3 converge on a single near-term opportunity. Detection is tractable across a whole catalog but not for one work, and one work’s value inside a pre-training corpus is negligible in any case. So the catalog (not the individual article, song or book) is the right unit to work at, technically and commercially alike.
What is worth detecting within a catalog is memorization: models sometimes retain their training data word-for-word and reproduce it verbatim in an answer, a behaviour usually called regurgitation. Showing that a model does this across a rights holder’s catalog is the strongest evidence available that the catalog was trained on. It also goes straight at the fair-use defense, which rests on training being transformative: a model that reproduces a work word-for-word has not transformed it, and it competes with the original for the same reader or listener.
Continue reading the full report
The Executive Summary above is free. Tell us a little about yourself to read the rest on the site or download the PDF.