All articles

Court News

AI and Copyright: Navigating Legal Challenges in the Digital Age

Updated 26 August 2026
AI and Copyright: Navigating Legal Challenges in the Digital Age

Authors vs. Algorithms: The Mounting Legal War Over AI Training and Copyright Law

A 50-Year-Old Legal Framework Faces Next-Gen Intelligence

Licensing Demands, Judicial Split, and the Shadow Library Paradox

By Legal Editor

New Delhi: August 24, 2026:

The global artificial intelligence race is currently colliding with a foundational legal question: Is it permissible to ingest copyrighted human expression to construct generative AI models? As large language models (LLMs) such as ChatGPT, Gemini, and Claude devour hundreds of millions of books, online articles, academic papers, and creative works, the legal systems governing intellectual property are experiencing unprecedented friction. Authors and publishers argue that tech giants are engaging in systemic copyright infringement to build multi-billion-dollar products, while AI developers maintain that ingesting text to learn structural patterns constitute non-infringement or transformative fair use.

 

Navigating this legal labyrinth requires dissecting the statutory foundations, recent precedent-setting judicial rulings, the core tenets of fair use doctrine, and the fundamental tension between technical input acquisition and synthetic output creation.

 

The Statutory Paradox: Applying the Copyright Act of 1976 to Modern AI

At the heart of the current legal battle lies a severe chronological mismatch. The primary governing legislation for intellectual property rights in the United States remains the Copyright Act of 1976 (17 U.S.C. § 101 et seq.). Drafted half a century ago, the statute was designed to address traditional physical formats and early computing applicationslong before neural networks, vector embeddings, or machine learning datasets were conceived.

 

Under 17 U.S.C. § 106, copyright owners hold exclusive rights over their works, including:

The right of reproduction (making physical or digital copies).

The right to prepare derivative works based upon the copyrighted material.

 

The right of public distribution, performance, and display.

When an AI developer ingests copyrighted literature, two primary technical actions take place. First, the data is downloaded, stored, and processed onto physical server infrastructure (creating digital copies). Second, the model processes the text through numerical representations to learn linguistic structures, syntax, and conceptual associations.

 

Under strict traditional interpretations of Section 106, unauthorized initial copying of copyrighted works into a database constitutes prima facie copyright infringement. However, copyright law distinguishes sharply between consuming a work to extract uncopyrightable concepts or statistical patterns and copying a work to replicate its protected expression. 17 U.S.C. § 102(b) explicitly specifies that copyright protection extends only to expression, never to "any idea, procedure, process, system, method of operation, concept, principle, or discovery."

 

Because LLMs analyze text to map token probabilities rather than storing verbatim copies within their parameter architecture, developers contend that machine learning mirrors human learning—an activity that copyright law has never prohibited.

Section 107 and the Fair Use Defense: Transformative Purpose vs. Market Harm

When unauthorized copying occurs, AI entities overwhelmingly rely on the affirmative defence of Fair Use under 17 U.S.C. § 107. Section 107 provides that unauthorized use of copyrighted material for purposes such as criticism, comment, news reporting, teaching, scholarship, or research is non-infringing.

 

Courts determine whether an unauthorized use qualifies as fair use by evaluating four statutory factors:

 

The legal battlefield predominantly turns on Factor 1 (transformative use) and Factor 4 (market effect). A use is deemed transformative if it adds something new, with a further purpose or different character, altering the original work with new expression, meaning, or message.

 

In past digital piracy litigations—such as Authors Guild, Inc. v. Google LLC (2015)—the Second Circuit ruled that Google's scanning of millions of books to build a searchable text index constituted transformative fair use because it provided a novel informational tool without providing a substitute for the books themselves. AI developers rely heavily on the Google Books precedent, asserting that using text to train a general-intelligence algorithm is fundamentally transformative.

 

However, rights-holders contend that unlike Google Books, which displayed only snippet samples to guide users toward purchasing full books, generative AI produces synthetic prose that directly threatens the economic livelihood of human creators.

 

Judicial Splits and Landmark Precedents: Anthropic, Thomson Reuters, and OpenAI

 

Recent rulings highlight a judicial landscape divided over how fair use and digital acquisition standards apply to AI platforms.

The Shadow Library Fall: Bartz v. Anthropic

 

In Bartz v. Anthropic, Judge William Alsup of the U.S. District Court for the Northern District of California established a crucial bifurcated precedent. The court recognized that training an LLM on published works to learn linguistic nuance and generate entirely novel content is inherently transformative fair use under 17 U.S.C. § 107. Judge Alsup analogized model training to an aspiring writer studying classic literature to master the craft.

 

However, the court distinguished how the training data was acquired. Anthropic obtained hundreds of thousands of books by downloading them from unauthorized pirate repositories ("shadow libraries"). The court ruled that downloading pirated files constitutes independent, illegal digital acquisition separate from the training act itself. This led to a historic $1.5 billion settlement covering past acquisition claims while keeping the broader legal defence of fair use intact for properly acquired inputs.

┌─────────────────────────────────────────────────────────────────────────┐

│ THE TWO-PRONGED JUDICIAL STANDARD

┌──────────────────────────────┐

│ DATA ACQUISITION │ │ MODEL TRAINING │

├──────────────────────────────┤ ├──────────────────────────────┤

Ingesting text from pirated │ │ Processing text to learn │

│ shadow libraries or illicit │ │ statistical language patterns│

│ scrapers. │ │ & structure. │

│ │ │ │

│ ❌ IMPERMISSIBLE TRANSFORMATIVE FAIR USE

│ (Breaches Copyright & Terms) │ │ (Protected under § 107) │

└──────────────────────────────┘

Commercial Direct Competition: Thomson Reuters v. Ross Intelligence

In Thomson Reuters v. Ross Intelligence, Judge Stephanos Bibas analyzed fair use in the context of commercial competition. Ross Intelligence copied Thomson Reuters’ proprietary legal summaries (Westlaw headnotes) to train an AI platform designed to directly compete with Westlaw. The court held that training an AI model on a competitor's content to replicate its primary function fails the transformative test because the secondary work shares the same essential purpose and character as the original.

 

This established a key principle: if an AI tool uses copyrighted inputs to build a product that serves as a direct market substitute for the original content, Factor 4 heavily favours the copyright holder.

 

Synthetic Derivative Outputs: Authors Guild v. OpenAI

In Authors Guild v. OpenAI, authors including George R.R. Martin and David Baldacci filed class-action claims asserting that OpenAI infringed their copyrights by both ingesting their novels and generating detailed plot summaries, sequels, or character iterations. In late 2025, Judge Sidney Stein of the Southern District of New York denied OpenAI’s motion to dismiss claims regarding output infringement.

 

The court noted that when an AI system generates concise derivative structural summaries or character continuations, those outputs may violate the exclusive right to create derivative works under 17 U.S.C. § 106(2), independent of whether the underlying training process was lawful.

 

Inputs vs. Outputs: Regulating Copyrightability and Protection

A fundamental legal distinction exists between using copyrighted material as an input for training and establishing copyright ownership over AI-generated output.

+-------------------+

| Copyrighted Books |

+---------+---------+

|

v (Input Phase)

[ 17 U.S.C. § 107 Fair Use Analysis ]

* Transformative learning vs. Piracy

* Substantiality & Market Substitution

|

v

+-------------------+

| AI LLM Engine |

+---------+---------+

|

v (Output Phase)

[ Human Authorship Requirement ]

* Thaler v. Perlmutter Standard

* De minimis assistance vs. Pure AI

|

v

+---------------------------------+

| Is Synthetic Content Protected? |

+---------------------------------+

Human Authorship Thresholds

Under U.S. copyright law, copyright protection applies exclusively to works created by human beings. In Thaler v. Perlmutter, the D.C. Circuit reaffirmed that fully autonomous AI-generated outputs cannot receive copyright protection due to the lack of human authorship.

 

The U.S. Copyright Office guidelines state that while mechanical text correction or standard grammar tools (like Microsoft Word spell check) do not strip a human author of their copyright, works generated primarily through natural language text prompts lack the requisite creative control to warrant protection.

 

Infringement via Output Extraction

Even if model training is deemed transformative, an AI model that emits near-verbatim excerpts of copyrighted input text poses substantial legal liability. Rights-holders can allege structural copyright infringement if the model demonstrates "memorization"—reproducing verbatim passages of protected works upon specific user prompting.

Searchable Legal Index & Frequently Asked Questions (FAQ)

================================================================================

AI & COPYRIGHT LAW: SEARCHABLE FAQ INDEX

================================================================================

[FAQ-01] Is training an AI model on published books legal under U.S. law?

[FAQ-02] What legal distinction did Judge Alsup draw in the Anthropic case?

[FAQ-03] Why is the Copyright Act of 1976 struggling with modern AI litigation?

[FAQ-04] What is the difference between AI training inputs and AI outputs?

[FAQ-05] Can an AI-generated book or script be copyrighted by its prompter?

[FAQ-06] How does the Thomson Reuters v. Ross Intelligence ruling affect AI startups?

================================================================================

[FAQ-01] Is training an AI model on published books legal under U.S. law?

Answer: It depends on how the data was acquired and how the model operates. Courts have ruled that analyzing published text to learn statistical language patterns is inherently transformative and protected under the Fair Use doctrine (17 U.S.C. § 107). However, acquiring that text via illicit shadow libraries or pirated websites remains illegal and subject to massive statutory penalties.

[FAQ-02] What legal distinction did Judge Alsup draw in the Anthropic case?

Answer: In Bartz v. Anthropic, Judge William Alsup distinguished between the method of data acquisition and the act of machine training. He ruled that training LLMs to generate new text is lawful transformative fair use. However, downloading nearly half a million books from illicit pirate repositories breached copyright laws, resulting in a $1.5 billion settlement.

[FAQ-03] Why is the Copyright Act of 1976 struggling with modern AI litigation?

Answer: The Copyright Act of 1976 was enacted long before neural networks, machine learning, or generative algorithms existed. As a result, federal judges are forced to extrapolate 50-year-old statutory provisions—designed for traditional physical media and basic broadcasting—to resolve complex technical issues surrounding parameter weights, vectorized text tokens, and algorithmic memorization.

[FAQ-04] What is the difference between AI training inputs and AI outputs?

Answer:

Inputs: Refers to the existing dataset (books, articles, media) ingested by developers to train an AI model. Input disputes center on whether unauthorized copying during dataset preparation constitutes fair use or piracy.

Outputs: Refers to the prose, code, or images produced by the AI after training. Output disputes center on whether generated text infringes on original authors' rights by reproducing protected expression or creating unauthorized derivative works.

[FAQ-05] Can an AI-generated book or script be copyrighted by its prompter?

Answer: No, purely AI-generated works cannot be copyrighted. Under Thaler v. Perlmutter and official U.S. Copyright Office guidelines, human authorship is a mandatory prerequisite for copyright protection. If a work is created entirely via AI prompts without substantial human creative arrangement or modification, it enters the public domain upon creation.

[FAQ-06] How does the Thomson Reuters v. Ross Intelligence ruling affect AI startups?

Answer: In Thomson Reuters v. Ross Intelligence, the court ruled that training an AI platform on a competitor's proprietary content to build a directly competing product is not transformative fair use. This creates significant risk for AI startups that scrape specialized domain data from established competitors to launch substitute products within the same commercial market.

Emerging Legal Frameworks and Future Trajectory

 

As courts continue to issue conflicting decisions across different federal jurisdictions, regulatory interventions and statutory updates appear inevitable.

The evolving legal landscape points toward three major industry shifts:

 

Mandatory Statutory Licensing Models: Similar to mechanical licenses in the music industry, future legislative amendments may require AI developers to pay standardized royalty fees to blanket collective rights organizations for text ingestion rights.

Provenance and Dataset Auditing: AI firms are increasingly shifting toward documented, legal data pipelines—pursuing official licensing deals with major publishers and media conglomerates to mitigate shadow library liability.

 

Technological Guardrails Against Output Memorization: To defend against derivative work claims under 17 U.S.C. § 106(2), AI labs are building post-processing filters designed to prevent chatbots from outputting verbatim passages or near-identical plot summaries of copyrighted books.

 

Until appellate courts or Congress establish unified national guidelines, the boundary between algorithmic learning and technological piracy will remain one of the most hotly contested frontiers in intellectual property law.

 

Statutory Factor — Legal Focus — Practical Application in AI Training

 

1. Purpose & Character of Use — Commercial vs. non-commercial status; whether the use is transformative. — Courts examine whether training an LLM creates an entirely new utility or merely repackages content.

 

2. Nature of Copyrighted Work — Factual/informational vs. highly creative/fictional material. — Highly creative novels receive stronger traditional protections than raw factual data.

 

3. Amount & Substantiality Used — Quantitative proportion and qualitative "heart of the work" ingested. — AI models ingest 100% of full books, which traditionally weighs against fair use.

 

4. Effect upon Potential Market — Market substitution; potential financial harm to the original work's value. — Examines whether AI output directly competes with original books or undercuts licensing markets.