-
3 minutes, 35 seconds
The complaint, filed in a U.S. federal court, accuses the AI companies of engaging in what it calls “one of the largest and most blatant ongoing thefts of intellectual property in history.” At the heart of the allegation is the claim that the defendants systematically scraped and reproduced copyrighted works—including books, articles, and other written material—without authorization to train their large language models. The plaintiffs argue this was not an incidental or accidental use, but a deliberate and sustained business strategy. They assert that the companies built commercially valuable AI systems by ingesting protected expression, thereby directly competing with the original creators in the marketplace. The lawsuit seeks to hold the firms accountable for what it describes as willful infringement, demanding both monetary damages and injunctive relief to halt the alleged unauthorized use. The core legal question, as framed by the plaintiffs, is whether the massive extraction of copyrighted text for AI training constitutes fair use or, as they contend, a clear violation of existing copyright law. The case sets the stage for a landmark legal battle.
The lawsuit centers on a specific corpus of copyrighted works—namely, a vast collection of news articles and books—that were allegedly used without authorization. The plaintiffs claim this material was scraped from the open web and incorporated into a training dataset. The scale of the alleged theft is significant: the complaint asserts that the dataset includes over 100,000 books and millions of news articles, all reproduced or summarized without permission.
Critically, the plaintiffs argue that this goes beyond mere ingestion of facts. They contend that the AI system’s outputs can generate near-verbatim excerpts from the protected works, effectively reproducing the expressive elements that copyright law protects. The alleged theft is not just of individual pieces but of the entire corpus, used to build a commercial product that competes with the original sources. The sheer volume and the systematic nature of the copying form the core of the legal claim, distinguishing it from isolated instances of infringement.
The lawsuit names OpenAI, Microsoft, and GitHub as defendants, alleging that their AI coding tool Copilot was trained on open-source code without proper attribution or licensing compliance. In response, GitHub publicly stated that it “has always been committed to responsible innovation” and emphasized that Copilot was designed to respect developers’ workflows, though it did not directly address the core theft allegation. OpenAI issued a brief statement asserting that its models are trained on “publicly available data” and that it believes “fair use” protects its practices, while declining to comment on the specific code at issue. Microsoft, the largest backer of OpenAI and owner of GitHub, has remained largely silent, issuing only a corporate boilerplate response that it “takes intellectual property rights seriously” and will review the claims. As of the filing date, none of the companies had publicly acknowledged any wrongdoing, and all three have declined to provide the plaintiffs with detailed training data logs requested in the complaint. This collective refusal to engage with the specifics has intensified scrutiny of their legal position.
The outcome of this case could set a critical precedent for how generative AI models are trained. If the court rules against the defendants, it may force AI developers to fundamentally alter their data acquisition methods, potentially requiring explicit licensing agreements for all copyrighted material used in training datasets. This would significantly increase development costs and slow the pace of innovation, particularly for startups that lack the resources of major tech firms.
Conversely, a ruling for the companies could be seen as a broad endorsement of "fair use" for AI training, potentially freezing the current legal landscape. However, the case also highlights the inadequacy of existing IP law, which was not designed for machine learning. As this Reuters report notes, the lawsuit underscores a growing tension between creators’ rights and technological progress. The industry now faces a period of uncertainty, with the need for new legislative frameworks becoming increasingly urgent to clarify the boundaries of ownership and fair use in the age of AI.
Comment