
Your favourite chatbot has almost certainly read your favourite novel, without asking you, the author, or the publisher for permission. So is that legal? The short answer, according to AI copyright law as it stands in 2026, is: it depends on what the AI company did with the book, not just whether it used it. Courts have started drawing a sharp line between training on a copyrighted book (often allowed) and pirating one (not allowed), and that distinction is quietly deciding the future of the entire AI industry.
If that sounds confusing, you’re not alone. Even IP attorneys admit this area of law is a mess right now. “I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.”
That messiness is precisely why AI copyright law matters so much right now, not just to authors and publishers, but to every student, developer, and content creator who uses AI tools built on top of copyrighted material. Let’s unpack what’s actually been decided, what’s still up in the air, and what it means for young professionals in India building careers around AI.
What Is Fair Use, and Why Does It Decide AI Copyright Law?
Fair use is a legal carve-out in copyright law that allows someone to use a copyrighted work without asking permission, as long as the use is limited and “transformative”, think criticism, parody, education, or commentary. It’s the exception that stops copyright from blocking every book review, remix, or classroom handout ever made.
Fair use isn’t a free pass, though. US judges weigh a few specific factors before deciding: the purpose and nature of the new use, how much of the original was used, and, critically, whether the new work damages the market for the original. This is exactly the test courts have been applying to generative AI training, and it’s why almost every major AI copyright law case right now hinges on the word “transformative.”
Does fair use automatically mean AI companies can use any book they want? No. Fair use is decided case by case, and the outcome depends heavily on why the AI was trained and what it competes with. As IP and media attorney Jason Henderson, founder of the IP & Media Practice at JWL International, put it, AI training tends to survive legal scrutiny when the resulting product doesn’t directly compete with the original work, and tends to lose when it does.
The Anthropic Case: A $1.5 Billion Fine That Actually Favoured AI Companies
The biggest AI copyright law ruling to date came from Judge William Alsup, who ordered Anthropic, the company behind Claude, to pay a $1.5 billion settlement to a group of authors whose books were used to train its models. On the surface, that looks like a huge win for writers. It wasn’t, exactly.
Judge Alsup actually ruled that training Anthropic’s AI models on copyrighted books was lawful. What he penalised the company for wasn’t the training itself, it was pirating the books from illegal “shadow libraries” (unauthorised online repositories of copyrighted texts) instead of acquiring them legitimately. In his ruling, Alsup compared how a large language model ingests text to how a writer studies literature, noting that the goal was to “turn a hard corner and create something different,” not to replicate or replace the original works.
Was Anthropic actually punished for training AI on books? Not for the training itself, for how it obtained the books. The fine was about piracy, not about the legality of AI training on copyrighted material, which the judge found to be fair use.
Attorney Cathy Gellis, who specialises in IP and technology law, argues the ruling is actually more favourable to AI companies than it first appears. Anthropic is reportedly projecting close to $200 billion in annual revenue by 2028, a number that makes a $1.5 billion settlement look more like a cost of doing business than a deterrent. Gellis frames the court’s reasoning as treating AI training more like reading a copyrighted work than copying it, since “copyright law hinges on copying… it doesn’t hinge on using the work or experiencing the work.”
Training vs. Piracy: The Line Courts Keep Drawing
This is the single most important distinction in AI copyright law right now, and it’s easy to miss if you only skim the headlines.
Definition + Expansion, “Transformative use”: A transformative use is one that adds new meaning, purpose, or character to the original work rather than simply substituting for it. In the Anthropic case, the court treated model training as transformative because the AI wasn’t reproducing books for readers, it was learning patterns from them to generate something new. That’s different from, say, uploading a scanned copy of a novel for people to read for free, which offers no transformation at all.
So training on legally acquired copyrighted books can be fair use. Training on pirated copies, or building a product that directly substitutes for the original market, is where companies get into trouble.
Copyright law in the US, notably, hasn’t been substantially updated since 1976, decades before anyone imagined a chatbot trained on hundreds of millions of books. Judges are essentially applying 50-year-old guidelines to a technology nobody anticipated, which is a big part of why rulings feel inconsistent from one court to the next.
Why AI Copyright Law Feels So Inconsistent Right Now
If you’ve noticed that one AI copyright law ruling seems to contradict the next, you’re reading the situation correctly, it genuinely is inconsistent, and legal experts say that’s normal for a young, unsettled area of law. “Everybody is very worried right now because the law is all over the place,” Jason Henderson told TechCrunch, noting that courts are still working out how a 1976-era copyright framework applies to models trained on internet-scale data.
Most major AI companies are currently sitting inside pending litigation over exactly these questions, which means there isn’t yet a single, settled rulebook for AI copyright law, there are early, influential decisions that could still be reshaped by appeals and future cases. As Gellis put it, the initial rulings are shaping how the entire industry behaves right now, even though “that influence itself could be undone if other courts decide different things.” In other words: what counts as legal AI training today could look different in two or three years, once higher courts weigh in on the same questions.
This matters practically. AI companies are making product and data-sourcing decisions based on rulings that are still, technically, provisional. For anyone building a career around AI, whether that’s prompt engineering, AI-assisted content, or building your own tools, it’s worth remembering that the legal ground here is still moving.
When Does AI Training Cross the Line? Thomson Reuters vs. Ross Intelligence
Not every AI training case has gone the way Anthropic’s did. In Thomson Reuters v. Ross Intelligence, the legal research firm Ross Intelligence was sued for using Thomson Reuters’ content to train a competing AI-based legal research platform.
Did the court side with the AI company here too? No, this time, the court ruled against the AI company. Judge Stephanos Bibas found that Ross’s use of Thomson Reuters’ material was not transformative, because the resulting product didn’t have a “further purpose or different character” than the original, it was built to directly compete with Thomson Reuters in the same market.
That’s the key differentiator across nearly every AI copyright law case so far: does the AI-trained product compete directly with the copyrighted source, or does it do something meaningfully different? Authors have tried to argue that AI chatbots compete with them by generating new “synthetic” books, but so far, that specific argument hasn’t won in court.
Comparing the Two Landmark AI Copyright Rulings
| Factor | Anthropic (Authors v. Anthropic) | Thomson Reuters v. Ross Intelligence |
| What was trained on | Books, legally and illegally sourced | Thomson Reuters’ proprietary legal content |
| Core legal question | Was AI training on books fair use? | Was training a competing product fair use? |
| Court’s finding on training itself | Training on books = lawful, transformative | Not transformative, no “different character” |
| What triggered the penalty | Acquiring books from pirate shadow libraries | Building a direct market competitor |
| Financial outcome | $1.5 billion settlement (for piracy, not training) | Ruled against Ross Intelligence |
| What it signals for AI copyright law | Training ≠ automatically illegal | Competing directly with the source ≠ fair use |
Can AI-Generated Content Even Be Copyrighted?
Here’s the flip side of the same debate: if AI training on copyrighted books is complicated, so is the question of who owns what an AI creates. In Thaler v. Perlmutter, the court ruled that a work generated entirely by AI, with no meaningful human authorship, is not copyrightable at all.
That ruling immediately raises a harder, unresolved question: how do you prove a work was AI-generated, and if a human edited or guided it, what percentage of AI involvement disqualifies it? Attorney Cathy Gellis offers a useful comparison here, using Microsoft Word’s spellcheck doesn’t mean Word owns your novel, but where exactly does “AI assistance” cross into “AI authorship”? Courts haven’t fully answered that yet.
What This Means If You’re a Student, Writer, or Freelancer in India
Most of the landmark rulings shaping global AI copyright law are happening in US courts, but the ripple effects reach Indian creators, students, and AI learners fast, from the tools you use daily to the freelance writing gigs increasingly competing with AI output.
It’s tempting to assume this is purely a Silicon Valley legal drama with no bearing on Bhubaneswar or Bengaluru, but that’s not quite right. Every major AI tool used by Indian students and professionals, from chatbots to coding assistants to AI writing tools, was trained under the same AI copyright law questions being fought out in US courtrooms. How those cases resolve directly shapes what these tools are allowed to do, how they’re priced, and how safely you can build a business on top of them.
A few practical takeaways:
- AI training on books isn’t automatically illegal, courts are focused on how the data was obtained and what the resulting product competes with, not simply whether copyrighted material was involved.
- Piracy is the real legal risk, not training itself, Anthropic’s fine was for pirated sources, not for the act of training an LLM on published books.
- “Transformative” is the word that matters most in nearly every AI copyright law ruling, a product that does something meaningfully different from the original tends to survive fair use scrutiny; one that directly competes tends not to.
- Litigation is still ongoing across most major AI companies, so today’s rulings could be overturned or reinforced by higher courts, nothing here is final law yet.
- India has its own copyright framework (the Copyright Act, 1957) built around “fair dealing” rather than the US’s broader “fair use” doctrine, and Indian courts haven’t yet issued a landmark AI-training ruling of their own, this is a space to watch closely if you’re building AI products or content for the Indian market.
- If you’re a freelance writer or content creator, understanding this distinction, training vs. piracy, transformation vs. competition, helps you make informed decisions about how your own work might be used, and how to think about AI tools in your own workflow.
Frequently Asked Questions on AI Copyright Law
Is it illegal to train AI models on copyrighted books? Not automatically. US courts, including in the Anthropic case, have found that training an AI model on legally acquired copyrighted books can qualify as fair use because it’s considered “transformative.” What’s illegal is acquiring those books through piracy or shadow libraries.
Why did Anthropic have to pay $1.5 billion if AI training was ruled legal? The $1.5 billion settlement wasn’t a penalty for training its models on books, it was for pirating those books from illegal shadow libraries rather than acquiring them legitimately. The judge separated the legality of training from the legality of how the training data was sourced.
What makes an AI’s use of copyrighted material “fair use” versus a violation? Courts look at whether the AI’s output is “transformative”, meaning it serves a different purpose than the original, and whether it directly competes with the original work’s market. Anthropic’s training was found transformative; Ross Intelligence’s competing legal platform was not.
Can content generated entirely by AI be copyrighted? No. In Thaler v. Perlmutter, US courts ruled that a work with no human authorship, meaning it was 100% AI-generated, cannot be copyrighted. How much human involvement is required to make a hybrid work copyrightable is still an open legal question.
Does India have the same AI copyright law as the United States? Not exactly. India’s Copyright Act, 1957, relies on the narrower “fair dealing” doctrine rather than the US’s broader “fair use” standard, and Indian courts have not yet issued a landmark ruling specifically on AI training. Indian creators and AI companies are largely watching how US and global cases unfold.
Is this area of AI copyright law settled yet? No. Most major AI companies are still facing pending litigation, and early rulings like Anthropic’s could be reinforced or overturned as higher courts weigh in. Legal experts describe the current landscape as the “opening volleys” of a much longer legal battle, meaning the rules AI companies follow today could shift meaningfully over the next few years.
Should students and freelancers in India worry about using AI tools that were trained on copyrighted books? Using an AI tool as a reader or consumer is a different legal question from how that tool’s maker trained it, the AI copyright law rulings discussed here are about the companies building the models, not the individuals using them. Still, it’s worth understanding this landscape if you’re building products, content, or a career around AI, since it shapes what these tools can legally do and how the industry is likely to evolve.
The bigger picture is this: AI copyright law is being written in real time, one lawsuit at a time, by judges applying decades-old statutes to technology that didn’t exist when those statutes were passed. Anthropic’s case suggests training on books can be legal fair use. Ross Intelligence’s case shows that direct competition can tip the scale the other way. Thaler v. Perlmutter adds a further wrinkle by questioning who, if anyone, owns what AI creates. None of these are final answers; they’re early data points in a legal fight that will likely take years to fully resolve.
For anyone in India studying AI, building AI-powered products, or simply trying to use these tools responsibly, understanding this shifting landscape isn’t optional background reading, it’s becoming a basic form of AI literacy.
Trying to make sense of how fast AI regulation, tools, and opportunities are moving? That’s exactly what we break down at Kalinga.ai, explore our AI learning resources and workshops built for students and young professionals across Odisha and India.