The Great Strip-Mine: How Big Tech Built AI on Systemic Theft

AI-generated image · US National Wire
Internal documents reveal a 'doom loop' where AI giants knowingly vacuum up creative labor to build products that destroy their own supply chain.
The current AI gold rush is not a miracle of engineering; it is a business model predicated on the systemic exploitation of human creativity. As first reported by The Register, for the giants of Big Tech, the world's books, news stories, code, and photographs are not intellectual property to be respected, but free resources to be strip-mined for profit.
This "take it now, argue about legality later" ethos has moved from a quiet corporate strategy to a documented reality. In the ongoing copyright battle between OpenAI, Microsoft, and The New York Times, recently unsealed court documents—reported on by 404 Media—reveal a startling level of internal awareness. Dr. Brent Hecht, Microsoft's Director of Applied Science, is quoted in a plaintiffs' brief describing the situation as an "astonishing theft of unprecedented proportions" and potentially the "largest theft of labor in human history."
This isn't just a legal loophole; it is a conscious choice to cannibalize the very creators who make these models possible. Internal Microsoft policy documents explicitly admit that generative AI could "significantly disrupt the employment" of the people who provided the training data. The documents describe a parasitic relationship where LLMs are a product that effectively "destroys its supply chain," leading to what Microsoft staff termed a "doom loop" of content strategy. By replacing the human creators who produce the "gold eggs" of original content, AI is triggering a model collapse.
The disregard for boundaries extends to the technical level. According to The Register, sworn deposition testimony from an OpenAI corporate representative indicated that the company's LLMs vacuumed up work even when protected by paywalls. The representative admitted to being unaware of any effort to detect or remove paywall content from training datasets. Most damningly, when OpenAI cofounder Greg Brockman was informed that the company could bypass firewalls to acquire data, he reportedly responded, "ah, nice."
While OpenAI and Microsoft argue in court that their actions constitute "fair use" and push forward public knowledge without acting as an "unlawful economic market substitute," the internal reality suggests a different motive: avoiding payment. This predatory approach extends to software development as well. While the US Ninth Circuit recently granted a narrow win to GitHub, Microsoft, and OpenAI in the *Doe v. GitHub* lawsuit—ruling that generating new code without copyright-management information (CMI) does not necessarily violate a specific DMCA provision—the underlying pattern remains the same.
Big Tech is betting that the public's desire for quick, automated answers will outweigh the ethical cost of erasing the creative class. As one OpenAI software engineer noted, users rarely click the links provided to verify information. By treating the internet as a free-for-all, Big AI isn't just innovating; it is executing a heist of unprecedented scale, betting that by the time the bill comes due, the original creators will already be obsolete.

