US National WireUS NATIONAL WIRE
TechOpinion

The 'Nobody Reads It' Defense: Microsoft's Attempt to Sanitize AI Theft

Portrait of Tobias Lund
Tobias Lundtelecom & connectivitySep 4AI
The 'Nobody Reads It' Defense: Microsoft's Attempt to Sanitize AI Theft

AI-generated image · US National Wire

By claiming Copilot rarely regurgitates full sentences, Microsoft is trying to convince a judge that stealing journalistic labor is a victimless crime.

Microsoft is attempting a classic Big Tech pivot: arguing that because the theft of intellectual property is often invisible to the end user, it isn't actually theft.

In new legal filings as first reported by The Verge, Microsoft is fighting copyright claims from book authors and publishers, including The New York Times. The company's strategy is to use the sheer volume of its data to hide the crime. Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers—logs specifically selected because they contained keywords linked to the plaintiffs' websites.

Microsoft's conclusion? Virtually nobody is actually reading the stolen goods. The company claims that fewer than 1 percent of those 8.2 million logs matched at least 16 words of the news content used to ground the AI model. Specifically, they point to a finding that only 59,545 logs hit that 16-word threshold. They further claim that an expert for the Center for Investigative Reporting found only 51 instances of "substantial overlap," and an expert in the authors' suit found only 24 responses with at least 30 matching words. Out of 212 books evaluated, Microsoft claims only 10 had any matches.

This is a cynical play. Microsoft is arguing that because Copilot rarely reproduces substantive chunks of text that could substitute for an original article, the use of copyrighted material for training datasets constitutes "fair use." They claim the resulting AI systems serve a "significantly different purpose" than the original works and that occasional reproduction of text does not undermine the "transformative purpose" of training Large Language Models (LLMs).

But let's be clear: the value of the LLM isn't in the verbatim regurgitation of a paragraph; it's in the synthesis of the knowledge contained within those millions of stolen words. The AI doesn't need to quote a full sentence to profit from the labor of the journalist who spent weeks reporting the story.

The New York Times is not buying the math. Ian Crosby, lead counsel for The Times, stated that the discovery process leads to only one conclusion: Microsoft and OpenAI stole from The New York Times to create commercial products that undermine the industry and substitute for its journalism.

Microsoft is now pushing for a summary judgment to end the case early. They want the court to believe that if the theft is fragmented enough, it ceases to be theft. It is a bold attempt to frame the open web as a free buffet for Big Tech, where the labor of creators is harvested for profit, provided the AI doesn't accidentally repeat too many words at once.

Sources

More from Tobias Lund