Legal docs in authors' lawsuit against Microsoft/OpenAI just got
By AI Update World · 2026-09-27

The tension between copyright holders and AI model developers has roots in decades of earlier tech disruption. When search engines first emerged, they indexed the entire web without explicit permission from every page author. When music streaming arrived, it raised questions about how artists should be compensated when their work is delivered algorithmically rather than purchased individually. The copyright system, shaped largely before the internet, assumes human readers or listeners consuming finished works. Large language models operate differently: they train by processing enormous quantities of text to recognize statistical patterns. This involves reading copyrighted books, articles, and other written works in ways that differ fundamentally from how a human reader engages with them.
At the core of current legal disputes is a question of copyright doctrine. Existing copyright law includes an exception called fair use, which permits limited use of copyrighted material without permission in certain contexts like criticism, education, and research. The law also distinguishes between copying a work and learning from it. When you read a book and retain what you learned, you have not violated copyright. But when a company copies millions of books into a database to train a system, the scale and commercial application shift the calculation. Courts have not yet established clear precedent for whether training an AI system on copyrighted text falls within fair use or constitutes infringement that requires licensing.
The economics driving this ambiguity matter greatly. Training large language models requires immense computational resources and substantial data. Companies could theoretically license every text used in training, but this would be expensive and logistically complex at the scale needed. Authors and publishers, meanwhile, were generally not asked for consent or offered compensation when their works were included in training datasets. This created a business model where copyright holders bore the original creative cost while AI companies derived commercial value. The question of whether this arrangement is legal remains unsettled.
Unsealed litigation documents offer the public a rare window into how companies actually acquired and used training data. Depositions and internal communications can reveal what company executives knew, what decisions they made about sourcing copyrighted material, and whether they considered licensing alternatives. Such evidence helps courts understand intent and reasonableness, factors that influence fair use analysis. The documents also often illuminate the practical constraints companies faced and the reasoning behind their choices, providing context beyond what either side argues in legal briefs.
These disputes carry implications beyond the immediate parties. If courts determine that training on copyrighted works without permission is infringement, it could reshape how AI systems are built. Companies might