US National WireUS NATIONAL WIRE
TechOpinion

The AI Royalty War: Who Actually Owns the Training Data?

Portrait of Alicia Ferro
Alicia Ferrofintech & paymentsSep 7AI
The AI Royalty War: Who Actually Owns the Training Data?

AI-generated image · US National Wire

A $1.5 billion settlement with Anthropic is revealing a systemic struggle over intellectual property carve-outs between authors, publishers, and agents.

The headline of the Anthropic copyright settlement is the $1.5 billion figure. But for those of us tracking the plumbing of the generative AI economy, the real story is the fight over where that money actually lands.

As TechCrunch first reported, Anthropic settled a copyright class action suit after a judicial ruling determined that while training AI models on copyrighted material is legal under fair use, the act of pirating that material is not. The resulting deal, finalized in July, targets nearly 500,000 titles, with a payment of $3,000 allocated for each pirated work.

On paper, the distribution mechanism is straightforward: if a book is still in-print with a traditional publisher, the payment is split 50-50 between the author and the publisher. If the work was self-published or the rights reverted to the author, the author receives 100%.

However, the execution of this payout is exposing a messy reality regarding IP ownership. According to TechCrunch, authors are reporting that publishers are claiming payments for books where rights have already reverted. Mystery and thriller author April Henry highlighted a specific dispute with HarperCollins, claiming the publisher sought payment for a book that reverted at least 17 years ago.

Victoria Strauss, writing for the blog Writers Beware, notes that complaints generally fall into two buckets: publishers claiming works they no longer own, and publishers seeking 100% of a payment when they are only entitled to 50%. While Strauss suggests these may be systemic errors rather than routine glitches, Mary Rasenberger, CEO of the Authors Guild, told The New York Times that she does not view this as a deliberate "grab" by publishers, attributing the chaos to poor record-keeping and a confusing settlement process.

Adding another layer of complexity is the role of intermediaries. Strauss reports that some literary agencies are also making claims on the settlement, despite the fact that agents are not rightsholders in the books they sell. This prompted a blunt reaction from author Courtney Milan (pen name of Heidi Bond) on Bluesky, who argued that agents should not be claiming percentages of the fund.

From a market perspective, the friction points here are critical. The settlement includes a specific "download date" of August 10, 2022; for an author to claim 100% of the payment, the rights reversion must have occurred before that date.

This dispute is a canary in the coal mine for the AI era. As model trainers move toward licensing deals and settlements to sanitize their datasets, the industry is discovering that the ledger of who owns what—and who is entitled to the resulting royalties—is dangerously outdated. If the infrastructure for tracking rights is this broken, the transition to a compensated AI training economy will be a logistical nightmare.

Sources

More from Alicia Ferro