kalinga.ai

Can AI Companies Legally Train LLMs on Copyrighted Material? The US Government Just Weighed In

Training LLMs on copyrighted material and the US AI copyright debate
The US government’s OpenAI brief could reshape how AI companies use copyrighted content for LLM training.

Imagine writing a book, watching it get scanned into a database without your permission, and then finding out a chatbot learned to write partly by “reading” it,  with zero payment and zero credit. That’s the exact fight playing out in a New York courtroom right now, and on September 2, 2026, the Trump administration formally sided with OpenAI, arguing that training LLMs (large language models,  the AI systems behind tools like ChatGPT) on copyrighted material should be legal. If you’re a student, aspiring AI professional, or content creator in India trying to make sense of what this means for you, here’s the direct answer: the US government isn’t ruling on the case, but its intervention could tilt how courts,  and eventually the whole world’s AI industry,  treat training LLMs on copyrighted material.

This single filing touches nearly every question people ask about AI and copyright: Is it legal? What is “fair use”? Why did Anthropic pay $1.5 billion if training itself isn’t the problem? And what does any of this mean if you’re building your career around AI in Odisha or anywhere else in India? Let’s break it down.

What Exactly Did the US Government Do?

The core fact: in a fresh escalation over training LLMs on copyrighted material, the Trump administration filed a 20-page legal brief in the New York Times v. OpenAI copyright lawsuit, arguing in favor of OpenAI’s right to train its models on copyrighted text without a license. This brief was submitted in a lawsuit that the New York Times filed against OpenAI, with the Trump administration contributing a 20-page brief in defense of the ChatGPT maker’s unlicensed use of copyrighted material to train its LLMs.

Question: Is this brief a legal ruling? No. This is not a ruling,  the case is being tried in the U.S. District Court for the Southern District of New York, and the authors of the brief don’t have jurisdiction over the outcome. A brief like this is called an amicus brief (a “friend of the court” filing),  a formal document where an outside party who isn’t directly suing or being sued shares its opinion to try to influence how a judge thinks about the case. The government doesn’t decide the case, but courts often take government input seriously, especially on questions with national economic or strategic stakes.

Why would the government bother getting involved at all? Because AI policy in the US right now is tied directly to national competitiveness. The brief states that the United States has a strong interest in continuing to develop a robust and competitive artificial intelligence industry that sets the standard for the practice and procedure of AI use globally, and that it is critical for the country to retain global leadership in artificial intelligence. That language echoes an executive order that President Donald Trump signed the previous year focused on removing regulatory barriers to American AI leadership. In plain terms: the administration is treating AI training data policy as a matter of economic and geopolitical strategy, not just a private legal dispute between a newspaper and a tech company.

Why Is the New York Times Suing OpenAI in the First Place?

Definition: Training data is the massive collection of text, images, or other content an AI model learns patterns from before it can generate new content. For LLMs like GPT, Claude, or Gemini, this training data includes books, news articles, code, and websites scraped from across the internet,  often without the original creators’ permission.

LLMs powering chatbots like ChatGPT, Claude, and Gemini are trained on incomprehensibly massive databases of published works, including copyrighted books, articles, and other media that AI companies feed into these systems without permission. The New York Times, among many other publishers, has argued that this practice is illegal when applied to their copyrighted material. The Times’ concern is straightforward: if a chatbot can answer questions using knowledge drawn from thousands of paywalled articles it never paid to license, that potentially undercuts the very business model that funds journalism in the first place.

This isn’t a one-off dispute,  it’s one of the defining legal battles of the current AI era, with outcomes likely to shape how every AI company, from Silicon Valley giants to smaller AI education platforms building tools in India, sources and licenses training content going forward.

Question: Is the New York Times the only publisher fighting this battle? No. The New York Times case is simply the most closely watched example of a much larger pattern: publishers, authors, artists, and news organizations across the US have filed similar lawsuits arguing that training LLMs on copyrighted material without payment or permission amounts to infringement. Some publishers have instead chosen to strike licensing deals with AI companies rather than sue, which shows that even within the industry, there’s no single agreed-upon path forward for how training LLMs on copyrighted material should legally or commercially work.

What Is “Fair Use” and Why Does It Decide Everything Here?

Question → Direct Answer: Is training LLMs on copyrighted material automatically illegal? Not automatically,  it depends on whether the use qualifies as fair use, a legal carve-out that permits certain uses of copyrighted material without the owner’s permission under specific conditions.

Definition + Expansion: Fair use is a doctrine in US copyright law that allows limited use of copyrighted material without needing permission from the rights holder, provided the use meets certain tests,  most importantly, whether the new use is “transformative” (meaning it creates something meaningfully different from the original rather than just reproducing or replacing it). This question,  whether you can legally use copyrighted material to train an AI,  isn’t black and white, which is exactly why there’s so much extensive legal debate around the subject, and these conversations typically center on the fair use doctrine. In cases like this one, the fair use debate comes down to whether AI companies’ use of copyrighted work is “transformative” enough for a judge to rule it legal.

The Trump administration’s brief leans hard into the “transformative” argument, warning against a narrow reading of the law. The brief argues that constraining LLM development under a misunderstanding of fair use doctrine would thwart creative and scientific progress while hindering American prosperity and economic mobility. That’s a strong economic framing of a legal question,  essentially saying that if courts interpret fair use too strictly, it could slow down US AI innovation relative to competitors elsewhere.

Wait,  Didn’t Anthropic Already Lose a Case Like This?

This is where a lot of confusion happens, so let’s be precise. So far, cases about AI training and copyright infringement have largely gone in favor of AI companies. Last year, Judge William Alsup ordered Anthropic to pay a $1.5 billion copyright settlement to a group of writers whose work was used to train the company’s AI models,  but crucially, Anthropic wasn’t penalized for the act of AI training itself. Instead, the company was fined specifically for using illegal shadow libraries,  pirated book repositories,  to source the books it used for training.

Question: So if piracy was the problem, was the training itself considered legal? Yes, largely. Judge Alsup wrote that, like any reader aspiring to be a writer, Anthropic’s LLMs trained upon existing works not to race ahead and replicate or supplant them, but to turn a hard corner and create something different,  comparing the AI’s training process to a human reading a book before writing their own. That’s a big deal: it suggests American courts are already leaning toward treating training LLMs on copyrighted material as transformative fair use, as long as the underlying data wasn’t obtained illegally (i.e., pirated). The Anthropic case punished how the books were acquired, not the concept of learning from copyrighted text,  a distinction that matters enormously for how the OpenAI case might play out.

What Happens Next in the OpenAI Copyright Case?

Question → Direct Answer: When will we know if training LLMs on copyrighted material is officially legal? Not soon. The New York Times v. OpenAI case is still in the discovery and briefing stage in the Southern District of New York, and a final trial verdict,  let alone any appeals,  could take months or even years to play out.

Here’s roughly how the process unfolds from here:

  • Briefing continues: Both sides, plus outside parties like the US government, keep filing arguments (like this amicus brief) before the judge rules.
  • A trial or summary judgment: The court will eventually decide whether OpenAI’s use of NYT content for training LLMs on copyrighted material qualifies as fair use, similar to how Judge Alsup ruled in the Anthropic case.
  • Possible appeal: Whichever side loses is highly likely to appeal, which could push the final word on this all the way to a federal appeals court,  or even the US Supreme Court.
  • Ripple effects on other lawsuits: Because dozens of similar copyright cases against AI companies are working through US courts simultaneously, a strong ruling in this case could influence,  or be influenced by,  those parallel cases.
  • Policy vs. courtroom outcomes: Even if courts side with publishers in some cases, government bodies (as seen in this brief) may keep pushing for AI-friendly policy through executive orders, trade rules, or future legislation.

Definition + Expansion: A precedent is a legal decision that serves as a guide or authority for deciding future cases with similar facts. If a US court eventually rules decisively on whether training LLMs on copyrighted material counts as fair use, that ruling would become a precedent that shapes not just the OpenAI–NYT dispute, but essentially every AI copyright case that follows it,  in the US and, indirectly, worldwide.

AI Training and Copyright: Where the Major Players Stand

AspectOpenAI / NYT Case (2026)Anthropic Settlement (2026)US Government Position
Core disputeTraining ChatGPT on NYT’s copyrighted articles without a licenseTraining Claude on books sourced via pirated “shadow libraries”Amicus brief supporting broad fair-use protection for AI training
Legal statusOngoing litigation, no ruling yetSettled,  $1.5 billion paidNot a ruling; advisory brief only
Was training itself ruled illegal?Not yet decidedNo,  the piracy, not the training, was penalizedArgues training should be considered fair use
Key legal doctrine at stakeFair use / “transformative use”Fair use, plus illegal data acquisitionFair use, framed around US competitiveness
Court/BodyU.S. District Court, Southern District of New YorkRuled by Judge William AlsupFiled in the same SDNY court, no jurisdiction to decide

Training LLMs on Copyrighted Material: Why This Matters If You’re Building an AI Career in India

For students and young professionals in Odisha and across India learning to build, fine-tune, or deploy LLMs, this case isn’t just US courtroom drama,  it directly shapes the tools, datasets, and business models you’ll work with.

  • Dataset licensing is becoming a real skill. As courts scrutinize how training data is sourced, companies (including Indian AI startups) will need people who understand licensed vs. unlicensed data pipelines.
  • “Fair use” thinking will influence global AI policy, including how India and other countries eventually regulate their own AI training practices.
  • Content creators and publishers need to pay attention too,  if courts affirm that training on copyrighted material is broadly fair use, publishers may pivot toward licensing deals instead of lawsuits (something already happening with several US publishers).
  • AI companies you might work for or build may face similar questions eventually, especially as India develops its own AI and data protection frameworks.
  • Understanding legal risk is now part of AI literacy,  not just knowing how to prompt a model, but knowing where its knowledge legally came from.

Kalinga.ai regularly covers these intersections of AI policy, ethics, and hands-on skill-building, because understanding the legal terrain is just as important as understanding the technology itself if you want a durable career in this field.

Question: Has India already seen a similar AI copyright case? Yes,  India got its own preview of this fight in 2026. News agency Asian News International (ANI) sued OpenAI in the Delhi High Court, and on July 24, 2026, Justice Amit Bansal refused to grant ANI an interim injunction, holding,  at least at this preliminary stage,  that OpenAI’s use of ANI’s articles to train ChatGPT fell within the “fair dealing” research exception under Section 52 of India’s Copyright Act. That’s the closest thing India has to its own answer on training LLMs on copyrighted material, and notably, the reasoning tracks the same “transformative use” logic being argued in the US courts and in the Trump administration’s brief. However, this was only an interim ruling,  the main ANI v. OpenAI suit is still pending, so India’s copyright law on AI training remains, like America’s, unsettled rather than finally decided.

For anyone hoping to work in AI policy, compliance, or data strategy in India, following how both the US and Indian courts resolve questions around training LLMs on copyrighted material isn’t optional reading,  it’s a live preview of the rules your future employer, client, or startup will eventually have to follow.

Frequently Asked Questions

Is it currently illegal to train LLMs on copyrighted material in the US? There’s no final ruling yet. Courts have so far leaned toward treating AI training as potentially legal under fair use,  as seen in the Anthropic case,  but the New York Times v. OpenAI case is still ongoing, and the Trump administration’s brief is only advisory, not binding law.

What did the Trump administration actually argue in its OpenAI brief? It argued that training LLMs on copyrighted material should be protected under fair use doctrine, framing restrictive interpretations of copyright law as a threat to US AI competitiveness and economic growth, referencing an existing executive order on AI leadership.

Did Anthropic get fined for training LLMs on copyrighted material? Not exactly. Anthropic was fined $1.5 billion, but the penalty was for sourcing books from illegal pirated “shadow libraries,” not for the act of training LLMs on copyrighted material itself.

What is fair use in the context of AI training? Fair use is a US copyright law doctrine that permits limited, unlicensed use of copyrighted material if the new use is sufficiently “transformative”,  meaning it creates something new rather than simply reproducing or replacing the original work.

Does this US legal battle affect AI companies or students in India? Indirectly, yes. Global AI companies operating in India follow US and international precedent closely, and the outcome could influence licensing practices, data sourcing standards, and even future Indian AI regulation,  all relevant to anyone building an AI career today.

Is this case fully resolved now that the government has filed its brief? No. The brief is one filing in an ongoing lawsuit in the Southern District of New York. No final judgment has been made, and the case could still take months or longer to resolve, with appeals likely regardless of the outcome.

Has an Indian court ruled on training LLMs on copyrighted material? Yes, in a preliminary way. In July 2026, the Delhi High Court declined to grant news agency ANI an interim injunction against OpenAI, finding that training ChatGPT on ANI’s articles fell within India’s “fair dealing” exception at this early stage,  but the main lawsuit is still pending, so this isn’t India’s final word on the issue either.


Want to understand how AI policy, copyright law, and training LLMs on copyrighted material actually connect to the tools you use every day? Explore Kalinga.ai’s AI education workshops and stay tuned for more breakdowns of the legal and technical shifts shaping the AI industry.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top