Blaze Legal
Artificial intelligence has compelled courts worldwide to confront questions that copyright legislation, much of which predates the digital age, was never designed to answer. Among the most significant is whether the use of copyrighted material to train large language models (“LLMs”) constitutes copyright infringement or whether such use represents a technologically transformative activity that copyright law ought to accommodate in the broader public interest.
The Delhi High Court’s recent interim order in ANI Media Pvt. Ltd. v. OpenAI marks India’s first substantial judicial engagement with this issue. Declining to grant an interim injunction against OpenAI, the Court prima facie held that the use of ANI’s news content for training ChatGPT may fall within the ambit of fair dealing under Section 52(1)(a) of the Copyright Act, 1957. The Court further observed that ChatGPT’s outputs, including those generated using Retrieval-Augmented Generation (“RAG”), did not reproduce ANI’s copyrighted expression and that granting an injunction at the interlocutory stage would cause disproportionate prejudice to innovation and the wider public interest.
Although the order does not finally determine the rights of the parties, its importance extends well beyond the immediate dispute. It is the first Indian decision to directly engage with the legality of AI training under copyright law and, in doing so, begins to define the legal framework within which artificial intelligence may develop in India. More importantly, it raises a broader jurisprudential question:
Should copyright regulate the process by which machines learn from copyrighted works, or should it concern itself only with whether those machines subsequently reproduce protected expression?
The dispute arose from ANI’s allegation that OpenAI had used its copyrighted news articles to train ChatGPT without authorisation, thereby infringing its exclusive rights under Sections 14 and 51 of the Copyright Act. ANI’s grievance was not confined to isolated responses generated by ChatGPT. Rather, it challenged the very process of model training, arguing that an LLM cannot be developed without copying, processing and analysing enormous volumes of copyrighted material.
Since OpenAI commercially exploits ChatGPT through subscriptions, enterprise licensing and API services, ANI contended that the unauthorised use of its journalistic content amounted to commercial exploitation of protected works without compensation.
OpenAI, on the other hand, advanced a fundamentally different understanding of machine learning. It argued that an LLM is not a searchable database containing copies of the works on which it has been trained. During training, textual material is tokenised, converted into numerical representations and repeatedly processed through optimisation algorithms that adjust billions of model parameters. The completed model does not preserve readable copies of articles but instead captures statistical relationships between words, phrases and concepts. Accordingly, ChatGPT generates responses by predicting language rather than retrieving copyrighted works, making AI training qualitatively different from conventional acts of reproduction.
At the interlocutory stage, the Delhi High Court declined to restrain OpenAI’s continued operation. Three aspects of the Court’s reasoning are particularly noteworthy.
The Court’s reasoning suggests that AI training is technologically distinct from the traditional acts of copying with which copyright law has historically been concerned. Conventional infringement generally involves reproducing protected expression for communication to the public. Machine learning, by contrast, analyses text to develop statistical representations of language rather than disseminating copyrighted works themselves.
The Court also found no prima facie evidence that ChatGPT reproduced ANI’s original expression. This reflects the well-established principle that copyright protects the expression of ideas—not ideas, facts or information themselves. While news articles enjoy copyright protection in their literary expression, the underlying facts remain within the public domain.
Finally, the order reflects that the Court has considered the broader consequences of granting an injunction. Restricting one of the world’s most widely used AI systems before a final adjudication could have affected innovation, education, research and India’s emerging AI ecosystem. The balance of convenience, therefore, favoured refusing interim relief.
Perhaps the most significant implication of the decision lies in its apparent emphasis on outputs rather than inputs.
Traditional copyright disputes ask whether protected works have been copied. AI complicates this inquiry because the alleged copying occurs during model training, whereas users interact only with the outputs subsequently generated by the system.
Unlike conventional databases, LLMs generally do not store complete documents in a retrievable form. During training, text is converted into tokens and numerical vectors before being processed through repeated optimisation. The resulting model consists of statistical parameters that enable language prediction rather than the storage of articles capable of retrieval.
By focusing on whether ChatGPT reproduced ANI’s protected expression instead of merely asking whether ANI’s articles formed part of the training corpus, the Court appears to have shifted the legal inquiry from the mechanics of learning to the consequences of that learning.
The Court’s apparent reliance on Section 52(1)(a) introduces one of the most important statutory questions arising from the litigation.
Unlike the United States, where courts apply the flexible doctrine of fair use, Indian copyright law recognises a structured framework of fair dealing exceptions. Section 52 identifies specific circumstances in which certain acts do not constitute infringement, including private study, research, criticism, review and reporting current events.
Supporters of the decision argue that AI training resembles research. Since machine learning analyses text to identify patterns rather than communicate protected expression, extending fair dealing may promote scientific advancement without materially prejudicing copyright owners.
Critics, however, contend that Parliament never contemplated commercial AI systems when enacting Section 52. The provision was intended for traditional research and scholarship rather than the large-scale ingestion of copyrighted works by commercial technology companies. Whether such an expansive interpretation is justified is likely to become the central issue as the litigation progresses.
The strongest aspect of the Delhi High Court’s order is its recognition that copyright law should not be mechanically applied to emerging technologies without appreciating their underlying functionality.
Copyright jurisprudence has consistently adapted to technological innovation, from photocopying and search engines to digital indexing, while preserving the economic incentives underlying authors’ rights. Courts have adapted established principles to accommodate search engines, cloud computing and internet intermediaries without abandoning the core objectives of copyright protection. Artificial intelligence presents a comparable challenge.
Supporters argue that copyright protects creative expression – not technological learning. Human beings routinely read books, newspapers and judgments before creating original works. AI developers contend that machine learning performs a computational analogue of this process. While the comparison is imperfect, it raises a legitimate question: if human learning is lawful, should machine learning necessarily be treated differently solely because it occurs at scale?
The Court’s emphasis on public interest also deserves recognition. Generative AI now underpins legal research, healthcare, education, accessibility technologies and software development. An injunction at the interim stage would likely have had far-reaching consequences extending beyond the immediate dispute.
Despite its pragmatic appeal, the order is not without criticism.
The principal criticism concerns Section 52 itself. India’s Copyright Act contains limited statutory exceptions rather than an open-ended fair use doctrine. Whether commercial AI training falls within “research” remains highly debatable.
Although an LLM ultimately stores statistical parameters rather than complete articles, the training process necessarily requires copyrighted material to be copied, tokenised and processed. Critics therefore argue that infringement, if any, occurs during these intermediate acts rather than when outputs are generated.
The commercial implications are equally significant. News organisations invest heavily in investigative journalism and editorial infrastructure. Increasingly, publishers have begun licensing archives to AI developers as a new revenue stream. If AI companies may freely train on copyrighted works, publishers may lose both bargaining power and valuable licensing opportunities.
India is not addressing these questions in isolation.
In the United States, ongoing litigation involving The New York Times, authors and AI developers focuses on whether AI training constitutes transformative use under the doctrine of fair use.
Decisions such as Authors Guild v. Google recognised that large-scale digitisation could be lawful where it created a socially beneficial use without replacing the market for original works, while Andy Warhol Foundation v. Goldsmith reaffirmed the importance of commercial purpose and market substitution.
The European Union has taken a legislative approach by introducing text-and-data mining exceptions under the Digital Single Market Directive while preserving the ability of rights holders to reserve certain uses.
The United Kingdom continues to evaluate possible legislative reforms, reflecting the absence of international consensus.
The litigation has exposed significant gaps within the Copyright Act, 1957. As the matter progresses, Indian courts may ultimately need to determine:
These questions are likely to shape India’s AI copyright jurisprudence for years to come and may ultimately require legislative intervention.
The Delhi High Court's interim decision in ANI Media Pvt. Ltd. v. OpenAI is significant not because it conclusively determines the legality of AI training, but because it reframes the legal inquiry. By recognising that machine learning presents technological and policy considerations distinct from conventional copyright infringement, the Court has signalled a willingness to interpret the Copyright Act in light of contemporary technological realities rather than historical analogies. At the same time, the order leaves unresolved fundamental questions concerning the scope of fair dealing, the meaning of “reproduction” in the context of AI training, and the continued viability of licensing markets for creators.
More importantly, the litigation exposes the limitations of a copyright framework enacted long before the advent of artificial intelligence. Whether commercial AI training should require licences, whether temporary computational copies constitute infringement, and whether liability should be assessed by reference to training inputs or AI-generated outputs are ultimately questions that extend beyond judicial interpretation and into the realm of legislative policy. As other jurisdictions continue to develop AI-specific copyright frameworks, India may likewise need to consider whether statutory reform is necessary to provide certainty for innovators and rights holders alike.
Ultimately, the challenge facing Indian copyright law is not whether it should favour innovation over creators, or vice versa. Copyright has always sought to balance private incentives with the public interest in the creation and dissemination of knowledge. That balance must now be recalibrated for the age of artificial intelligence. Whether AI ultimately becomes the greatest catalyst for human creativity or its greatest commercial disruption will depend less on technological capability than on the legal framework governing its development. In that sense, ANI Media Pvt. Ltd. v. OpenAI is unlikely to be the final word on AI and copyright in India – it is the opening chapter of what promises to be one of the defining areas of intellectual property jurisprudence in the coming decade.