Is it legal to train AI models on copyrighted books? It’s complicated
If you have been following the rapid evolution of generative AI, you are likely aware that the engines behind tools like ChatGPT, Claude, and Gemini are fueled by massive, seemingly bottomless repositories of human knowledge. These datasets encompass hundreds of millions of books, academic journals, news articles, and virtually every scrap of data accessible across the open web.
For many authors, this reality is unsettling. Their life’s work has been ingested into the development of technologies that many fear will ultimately erode their professional livelihoods—and all of this has occurred without their explicit consent or compensation. It feels like a clear-cut case of intellectual property theft, yet the legal landscape is far from straightforward.
The Complexity of Copyright in the Age of AI
The intersection of generative AI and copyright law is a legal minefield, characterized by shifting precedents and deep-seated industry tensions.
“I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on. It’s very complex and there are a lot of raw feelings about what is happening, both for and against.” — Cathy Gellis, intellectual property and technology attorney.
Last year, a landmark ruling involving Anthropic brought these tensions to the forefront. A federal judge ordered the company to pay a $1.5 billion settlement to a group of authors whose works were utilized in model training. While the headlines painted this as a massive win for creators, the legal reality was more nuanced. Judge William Alsup actually affirmed that the act of training an AI on copyrighted works is lawful. The massive fine was not for the training process itself, but for the company’s reliance on pirated content sourced from illegal "shadow libraries."
In his ruling, Judge Alsup drew a fascinating parallel: he likened the way a Large Language Model (LLM) processes trillions of words to a human writer studying literature to hone their craft. He noted that the AI was not designed to replicate or replace the original works, but rather to "turn a hard corner and create something different."
Is Reading the Same as Copying?
For AI companies, the distinction between "reading" and "copying" is the cornerstone of their legal defense. Cathy Gellis suggests that the Anthropic ruling is a significant victory for the industry. When viewed against the backdrop of an AI company projected to hit $200 billion in annual revenue by 2028, a $1.5 billion fine is a manageable cost of doing business.
The core of the argument is that copyright law is fundamentally designed to regulate the copying of works, not the consumption or experience of them. Because the law has not seen a major overhaul since 1976, judges are currently forced to apply half-century-old guidelines to technologies that were unimaginable at the time.
The "Fair Use" Dilemma
The central debate often revolves around the doctrine of fair use, which allows for the use of copyrighted material without permission for purposes like criticism, education, or parody. To determine if an AI’s use of data is "fair," courts look at several factors:
- Purpose and nature of the use: Is it transformative?
- Amount of the work used: How much of the original is consumed?
- Market impact: Does the AI output directly compete with the original work?
Jason Henderson, founder of the IP & Media Practice at JWL International, notes that the courts are struggling to find a consistent standard. “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question,” Henderson explained.
When Competition Becomes the Deciding Factor
While the "transformative" nature of AI training is a strong defense, it isn't a blanket pass. A notable case involving Thomson Reuters and the research firm Ross Intelligence highlights the limits of fair use. In that instance, the court ruled that Ross Intelligence’s use of Reuters' content was not transformative because it was used to build a platform that directly competed with the original source.
Judge Stephanos Bibas was clear in his assessment: if the purpose of the training is to create a product that competes directly with the source material, the courts are far more likely to intervene. However, for most authors, proving that a chatbot is "competing" with their specific book remains a difficult hurdle to clear in court.
The Future of AI-Generated Content
Beyond the training phase, there is the separate, equally thorny issue of whether AI-generated output can be copyrighted. In the case of Thaler v. Perlmutter, the court established that works created entirely by AI are not eligible for copyright protection. This creates a new set of questions: 1. How do we verify if a work was AI-assisted? 2. What percentage of human input is required to claim authorship?
Gellis draws a comparison to modern word processors: “If you write your novel in Microsoft Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel. [AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.”
A Long Road Ahead
We are currently in the "opening volleys" of a legal battle that will likely span years. While early rulings provide some guidance, they are subject to being overturned or contradicted as more cases move through the judicial system.
For now, AI companies must navigate a landscape where every legal decision acts as a potential precedent for the future of the industry. As Gellis aptly put it, these decisions are shaping the trajectory of the entire field, and ignoring them would be a significant strategic error for any company operating in the AI space. Until the law catches up to the technology, the status of our creative works remains in a state of high-stakes limbo.