Internal Documents Reveal Big Tech's 'Astonishing Theft' of Creative Work

AI-generated image · US National Wire
Unsealed court filings suggest Microsoft and OpenAI viewed the mass ingestion of internet data as a disruption of the very supply chain their models rely on.
Unsealed court documents from a copyright battle between The New York Times, OpenAI, and Microsoft reveal internal admissions regarding the ingestion of creative intellectual property, as first reported by 404 Media and detailed by The Register.
In a 92-page combined brief, Microsoft's Director of Applied Science, Dr. Brent Hecht, is quoted describing the case as "an astonishing theft of unprecedented proportions" and potentially the "largest theft of labor in human history." Further internal Microsoft policy documents cited in the papers admit that generative AI could "significantly disrupt the employment" of the people who created the training data, noting that LLMs are products that destroy their own supply chain—a cycle described as a "doom loop."
Additional filings cited by The Register reference sworn deposition testimony from an OpenAI corporate representative, who stated they were unaware of any efforts to detect or remove paywall-protected content from training datasets. The report further notes that when OpenAI cofounder Greg Brockman was informed the company could bypass firewalls, he responded, "ah, nice."
Microsoft and OpenAI are currently disputing these allegations in the US District Court for the Southern District of New York, where Judge Sidney H. Stein has yet to issue a ruling. The companies are utilizing a "fair use" defense, arguing that training LLMs on public books and articles advances public knowledge rather than serving as an "unlawful economic market substitute."

