top of page

Training Intelligence, Testing Copyright: Analysing the Future of AI Copyright Law in India

Writer: Niharika Puri
Niharika Puri
14 hours ago
5 min read

Introduction


For the better part of a year and a half, a good chunk of India's copyright bar and most of its nascent AI industry sat through hearing after hearing in a single Delhi High Court courtroom, waiting to find out whether the country's news agencies could stop OpenAI from reading their websites. On 24 July 2026, the Judge gave them an answer and, more importantly, gave Indian law its first sustained judicial engagement with what generative AI actually does to a copyrighted work. The order itself only refuses an interim injunction; the underlying suit still goes to trial. But the 135-page reasoning behind it does something an interlocutory order rarely manages: it builds, almost from first principles, a framework for thinking about training data, memorisation, and fair dealing that every AI copyright dispute in India will now have to reckon with.


Background and Facts


ANI, India's largest news wire service, sued OpenAI in November 2024, alleging that ChatGPT had been trained on ANI's copyrighted news articles and interviews without a licence, and that it reproduced them- sometimes near-verbatim- in its answers to users. ANI split this into two distinct grievances: a "training claim" (unauthorised copying and storage of its content to build OpenAI's language models) and an "output" or "reproduction claim" (the models regurgitating that content back to users). OpenAI contested jurisdiction, denied that either activity infringed, and pleaded fair dealing under Section 52(1)(a) of the Copyright Act, 1957 in the alternative. Six intervenors joined in- publisher’s bodies and the music industry backing ANI, AI developers and policy think-tanks backing OpenAI and the Court appointed two amici curiae to help it navigate genuinely uncharted technical territory. What followed was more than a year of hearings, closing submissions filed as late as April 2026, and a judgment that reads less like a routine interim order and more like a treatise.


Analysis of the Judgment


The Court cleared jurisdiction first, and briskly. Since ANI's registered office sits in Delhi and OpenAI actively courts Indian subscribers, Section 62(2) of the Copyright Act and Section 20 CPC gave the Court territorial jurisdiction over the reproduction claim without much trouble. The training claim was trickier, since the actual copying happens on servers in the United States. The Judge's answer was that storage abroad is merely the "terminal step in the chain of events" that begins with data being accessed and transmitted from India- a defendant cannot outrun Indian copyright law simply by locating the last link in that chain offshore. Leaning on the Neetu Singh v. Telegram[1] line of reasoning on offshore servers, the Court effectively refused to let jurisdiction turn on where a hard drive happens to sit, and usefully for future litigants, declined to treat the training and output claims as hermetically separate causes of action.


Before touching copyright at all, the judgment spends several pages explaining, in plain terms, how a large language model is built: tokenisation, embeddings, the predict-and-correct training loop, and retrieval-augmented generation. This is not throat-clearing; it does real work later. On the output claim, ANI's strongest evidence was a ChatGPT response quoting Neeraj Chopra's mother that closely tracked ANI's own reporting. The Court picked the example apart: the training cut-off for the relevant models predated publication of the very articles ANI relied on, so the response could not have come from memorised training data at all- it had to be RAG, live retrieval from ANI's own, unblocked website. Applying the classic R.G. Anand test of comparing works as a whole rather than isolated fragments[2], the Court found ChatGPT's answers added independent commentary and differed materially in expression, and held that no substantial reproduction was made out.


The fair dealing analysis is where this judgment earns its "landmark" label. The Court read Section 52(1)(a) as an integral, user-facing right rather than a grudging exception to be construed narrowly- a position borrowed partly from Canada's CCH Canadian and partly from Indian precedent on photocopying for education. It then built a two-step test of its own, a purpose test and a fairness test, having found no single formula consistently applied across Indian judgments, and, tellingly, held that the American four-factor test has no direct application here even as it borrowed heavily from American reasoning on transformative use, including Bartz v. Anthropic[3]. On purpose, the Court rejected the argument that commercial use automatically forfeits the defence, noting that Parliament excludes commercial use expressly wherever it intends to, and Section 52(1)(a) contains no such exclusion. It read "private" expansively enough to cover a closed, machine-only training pipeline, and gave "research" an "updating construction" to include machine learning- reasoning that if an AI can replace a human teacher, there is no principled reason research should remain a human monopoly. On fairness, absent proof that ChatGPT had displaced ANI's subscription revenue and given that ANI itself had offered OpenAI a licence for USD 7.5 million, rather undercutting its own claim of irreparable harm, the Court found the balance tilted toward OpenAI, reinforced by the public interest in AI access, education, and accessibility.


This is a broad, almost American-inflected reading of fair dealing dressed in Indian statutory language, and it will not please rightsholders. It gives AI developers real comfort that training on freely accessible, unpaywalled content is unlikely to be enjoined pending trial, and will embolden Indian LLM projects wary of licensing costs, a live concern given the government's own push toward a domestic AI stack. It also leaves publishers with a clear roadmap: block crawlers, evidence lost subscribers, and prove actual memorisation, rather than relying on generalised allegations. But this comfort rests on facts that will not always recur, the output-claim analysis turned on the fortunate accident that ANI's chosen examples predated the training cut-off. A plaintiff who can show genuine verbatim regurgitation of works the model was actually trained on, as happened in Germany's GEMA v. OpenAI[4], would face a very different Court.


Conclusion


Precisely because it is an interim order and not a final verdict, this judgment's power lies in its reasoning rather than its result. It gives Indian AI developers a workable, if generous, roadmap for the fair dealing defence, and gives publishers a clear list of what they will need to prove at trial to succeed- real memorisation, real market substitution, real economic harm. Whether the Division Bench or the Supreme Court eventually narrows this reading remains to be seen, and important questions - RAG-based infringement chief among them - have been left expressly open for trial. For now, though, ANI v. OpenAI is the closest thing Indian copyright law has to a founding text for the generative AI era, and every lawyer drafting a training-data licence, a scraping notice, or the next AI copyright suit will be reading it for years to come.






Niharika Puri

Associate








[1] Neetu Singh v. Telegram, CS (COMM) 282/2020

[2]R.G. Anand v. Delux Films, (1978) 4 SCC 118.

[3]Bartz v. Anthropic PBC, No. 3:24-cv-05417 (N.D. Cal.).

[4] GEMA v. OpenAI, Case No. 42 O 14139/24

Comments


Search By Tags
bottom of page