Delhi High Court Examines Fair Dealing in Generative AI Training and Copyright Protection

Delhi High Court Examines Fair Dealing in Generative AI Training and Copyright ProtectionLarge Language Models (LLMs) require mammoth volumes of content to identify patterns, develop predictive capabilities and generate responses to user prompts. However, publishers and creators contend that the collection and storage of creative content or similar works for AI training interfere with their exclusive rights and contribute to the development of commercial products without compensation. Resultantly, the use of copyrighted material for training generative artificial intelligence (AI) systems has emerged as one of the most contested questions in IP law.

The Delhi High Court looked into these issues in the case of ANI Media Pvt. Ltd. vs OpenAI OpCo LLC [CS(COMM) 1028/2024], passing its decision on July 24, 2026. In an application filed by ANI for an interim injunction against OpenAI, ANI contended that OpenAI had copied and stored ANI’s copyrighted material for training the LLMs for ChatGPT. Further, it was argued that ChatGPT reproduced ANI’s works in responses generated for users. OpenAI denied these claims and argued that their actions were protected under Section 52(1)(a) of the Copyright Act, 1957, which allows for fair dealing. The court decided not to grant the interim injunction, stating that ANI had not shown enough evidence of copyright infringement.

Functioning of LLMs

The Court first explained the technical functioning of LLMs. An LLM predicts the next word or sequence of words based on patterns identified from large datasets consisting of licensed material and publicly accessible online content. During training, data is divided into tokens and converted into numerical representations. The model repeatedly analyses these inputs and adjusts its internal parameters to improve its ability to predict language and perform tasks such as answering questions, summarising text and translating content.

The Court separately discussed Retrieval-Augmented Generation (RAG), through which an LLM retrieves current information from external sources in response to a prompt and uses that material to generate an answer. This distinction became relevant as the articles relied upon by ANI were published after the relevant ChatGPT models had completed training. The Court therefore considered whether the responses resulted from memorised training data or from live retrieval through RAG.

Issues Before the Court

The Court considered four issues:

  1. Whether storing ANI’s data for training ChatGPT was copyright infringement.
  2. Whether ChatGPT’s use of ANI’s copyrighted data to generate responses an infringement.
  3. Whether OpenAI’s use of ANI’s material was eligible as fair dealing under Section 52 of the Copyright Act.
  4. Whether Indian courts had jurisdiction when OpenAI’s servers and training operations were located outside India.

Jurisdiction of the Court

It was argued by OpenAI that its models were trained and the relevant data was stored and processed on servers in the United States. Applying the Copyright Act to those activities would amount to giving the legislation extra-territorial effect. The Court rejected the jurisdictional objection and noted that ANI had its principal place of business and registered office within the territorial jurisdiction of the Delhi High Court. Section 62(2) of the Copyright Act provided a jurisdictional basis for the suit.

OpenAI provided its services to users and paid subscribers in India, including Delhi. The outputs mentioned by ANI were generated within India. The Court decided that storing data on servers in the United States was just the last step in a process that started with accessing works from India and sending them abroad.

The Court relied on earlier decisions concerning digital platforms to hold that the location of foreign servers does not, by itself, deprive Indian copyright owners of remedies. It also declined to separate the training and output claims at this stage because the disputed outputs were alleged to result from the training process and had been generated in India. Accordingly, the Court held that it had territorial jurisdiction under Section 62(2) of the Copyright Act and Section 20 of the Code of Civil Procedure, 1908.

Copyright Ownership in ANI’s News Material

ANI relied on a sample professional services agreement providing that copyright in original works created or obtained by personnel on its behalf would vest in ANI. The Court found that the arrangement prima facie supported ANI’s ownership claim under Section 17 of the Copyright Act.

The Court explained that merely because material is available on a website does not mean it is not protected by copyright. Making something public does not mean the copyright is given up or that anyone has the right to reproduce the work. Therefore, ANI has the exclusive rights under Section 14 to reproduce and share its original written works with the public.

However, the Court distinguished copyright in a news report from ownership of the underlying event. Facts, events and information cannot be monopolised through copyright. Protection extends only to the original form, selection, arrangement and expression through which they are communicated.

Originality and Protection Available to News

Relying on Eastern Book Company vs D.B. Modak, the Court reiterated that Indian law does not apply the “sweat of the brow” test. Labour, effort or investment alone does not create copyright. The work must reflect the exercise of skill and judgment with a non-trivial degree of creativity.

The Court also referred to the merger doctrine, under which expression may not be protected where a fact or idea can be expressed only in a limited number of ways. News reports may therefore receive comparatively narrow protection. A claimant must establish copying of a substantial part of the original expression rather than merely the communication of the same facts in different words.

ANI’s Output Claim

ANI argued that ChatGPT memorised material contained in its training data and subsequently reproduced it in responses to user prompts. It relied on several examples which, according to ANI, demonstrated exact or nearly exact reproduction.

OpenAI stated that large language models (LLMs) learn patterns from data instead of memorising and repeating the original data. It recognised that occasionally, these models might accidentally repeat something, but argued that ANI did not prove any specific case of this with its work.

 

The Court found a key issue with ANI’s examples. The models had training cut-off dates in April 2022 and April 2024, while the articles ANI used were published in August and September 2024. This meant those articles could not have been part of the training data, so they could not show memorisation or repetition. The Court believed it was more likely that the responses came from live data retrieval using a method called RAG. It also pointed out that ANI had not specifically included an infringement claim based on RAG, even though it was mentioned during the arguments.

Substantial Reproduction

The Court applied the case of R.G. Anand vs Deluxe Films to determine if infringement occurred. It stated that infringement happens when a significant part of the claimant’s work is taken, including its form, arrangement, and expression. The Court emphasised that works should be compared as a whole, not just in parts.

The Court looked into ANI’s claim about ChatGPT’s summary of an interview with an Indian athlete’s mother after the 2024 Olympics. In the first response, ChatGPT summarised the interview and included its own comments. While some basic facts matched, the Court found the way it was expressed to be different. When ANI asked ChatGPT to repeat the exact words, the Court viewed this as a prompt meant to get a specific answer. Even then, ChatGPT only provided part of the quote and added context and explanations of its own.

The Court further noted that under Section 17(cc), the person delivering a public address or speech is ordinarily the first owner of copyright in it. ANI had not established an assignment from the athlete’s mother and could not prima facie claim ownership over the quotation or its translation merely because it had included the quotation in its report. The Court also observed that ANI had selected only a limited extract from a substantially longer article. When the works were compared as a whole, substantial reproduction was not established.

Latest News Responses and Use of RAG

ANI also relied on a prompt requesting the latest updates from its website. ChatGPT generated brief factual summaries, referred to ANI as the source and directed the user to its website. The Court found that the wording was materially different from ANI’s article titles and that the outputs resembled an AI-enabled search function rather than reproductions of ANI’s protected expression.

Reference to Foreign Decisions

The Court made a distinction from the German case of GEMA vs OpenAI, where exact song lyrics were created using simple prompts. ANI’s articles were published after the model training took place, and even when using detailed and challenging prompts, they did not result in producing a significant amount of the original content.

Associated Press vs Meltwater and the proceedings involving Cohere were also distinguished because they involved pleaded instances of verbatim or extensive copying that were not established in ANI’s case. The Court emphasised that foreign decisions may provide assistance, but the dispute must ultimately be decided under the Indian Copyright Act. The Court therefore held that ANI had not established either memorisation and regurgitation of its works or substantial reproduction through ChatGPT’s outputs.

On the Output Claim

The Court decided that ANI did not prove that its works were memorised or repeated by the model. The examples they used were published after the training of the relevant models, so they could not support the idea that the model memorised information from training. It was uncertain whether ChatGPT keeps raw data and can reproduce it, which may need technical evidence at trial. The Court also found that ChatGPT’s outputs did not show significant reproduction of ANI’s news articles. Thus, ANI did not establish a strong enough case for infringement based on the generated outputs.

Electronic Storage as Reproduction

The Court also considered whether storage of ANI’s works during model training amounted to reproduction. Section 14(a)(i) grants the owner of a literary work the exclusive right to reproduce the work in any material form, including storage by electronic means. The Court interpreted this provision broadly and held that electronic storage of a literary work constitutes reproduction whether the storage is temporary or permanent. The purpose of storage is not relevant when first determining whether the act falls within the copyright owner’s exclusive right.

OpenAI acknowledged that training its language models involved temporarily using literary works by authors. This could be considered copyright infringement unless it is allowed by another part of the law. The Court viewed Section 52 as an essential part of the law that defines legal ways to use copyrighted works. It balances the rights of authors with the public’s interest in sharing knowledge and creativity. To apply Section 52(1)(a), the Court used a two-step test:

  1. The use must be for one of the purposes listed in Section 52(1)(a), referred to as the “purpose test”.
  2. The dealing must be fair, referred to as the “fairness test”.

OpenAI relied on Section 52(1)(a)(i), which protects fair dealing for “private or personal use, including research”.

Commercial Use and Fair Dealing

ANI argued that OpenAI’s use for profit should not be included under Section 52(1)(a)(i). The Court disagreed, stating that while other parts of Section 52 mention non-commercial use, Section 52(1)(a) does not have this limitation. The fact that something is commercial can play a role in deciding fairness, but it does not automatically prevent fair use. The Court also pointed to Canadian cases that show commercial legal research can qualify as fair dealing.

Regarding Source Copy

ANI argued that OpenAI had made an unauthorised copy and could not rely on Section 52. The Court held that the expression “not itself an infringing copy” in the explanation to Section 52(1)(a) was confined to the incidental storage of a computer programme. It did not impose a general lawful-source requirement for every work stored electronically. The Court further noted that OpenAI had accessed material freely available on ANI’s website. There was no allegation that it had bypassed a paywall, used a shadow library or obtained the material from an unlawful source. It therefore found no basis at the interim stage to treat the source copies as infringing.

Meaning of “Private Use”

ANI argued that “private or personal use” was limited to individual human users and could not extend to a commercial corporation. The Court held that while “personal” may relate to an individual, “private” may extend to a closed group, organisation or company. Since the training material was stored in a closed system and was not available to the public for reading or download, the Court treated its use as “private” under Section 52(1)(a)(i).

LLM Training as Research

The Court regarded research as a process of investigation and learning that ordinarily precedes a publicly available output. During LLM training, stored works are analysed and processed iteratively to improve the model’s predictive capabilities. The Court considered this process a form of research directed towards advancing AI systems. Applying the doctrine of updating construction, the Court held that the term “research” must be interpreted in light of technological developments. Research and learning need not be confined to acts directly performed by humans and may extend to machine learning undertaken at the instance of, and for the benefit of, humans. It was concluded that storage and use of ANI’s works for training LLMs satisfied the test under Section 52(1)(a)(i).

Fairness Test by the Court

The Court formulated three factors appropriate to the dispute:

  1. Whether OpenAI’s use of ANI’s works was limited to training its LLMs.
  2. Whether the use economically competed with ANI, prejudiced ANI’s legitimate interests or caused actual or potential commercial harm.
  3. Whether ChatGPT’s functions served public interest.

The Court related these factors to the principles of Article 9 of the Berne Convention, which allows for limitations on reproduction where they do not interfere with normal use or harm the author’s rightful interests.

Training of LLM: ANI had not shown that OpenAI used its works for any purpose other than model training or that the model made exact copies publicly available. The first factor therefore favoured OpenAI.

Commercial Impact: The Court found that ANI was involved in the business of news syndication, whereas ChatGPT was involved in diverse functions that differed from ANI’s. ChatGPT created short summaries and references instead of substitutes for ANI’s complete articles or syndication business. Further, ANI did not provide any evidence regarding loss of subscribers, reduced advertising income or decline in licensing revenue due to ChatGPT.

Public Interest: The Court considered the wider benefits of LLMs in education, scientific research, translation, software development, accessibility and dissemination of information. It held that these public benefits supported the fairness of the dealing.

Training and Fair Dealing

The Court noted that ANI had the technical ability to restrict access by web crawlers but had not exercised the available opt-out mechanism. OpenAI also stated that it had independently blocked ANI’s website for training and RAG purposes. ANI had not demonstrated loss of subscribers or harm to its syndication business. Its offer to license content to OpenAI for USD 7.5 million also indicated that the alleged loss was capable of monetary quantification and could be compensated if ANI ultimately succeeded. By contrast, an injunction requiring deletion of training material or restricting model operation could materially affect OpenAI’s systems, its users and the development of AI models in India.

The Court also considered that requiring licences from every source used for model training could make LLM development economically unviable, particularly for domestic developers. The balance of convenience, irreparable harm and public interest therefore weighed against interim relief.

Court’s Decision

The Court dismissed ANI’s application, concluding that the Delhi High Court had jurisdiction to decide the suit despite the servers being located outside India. Further, it was held that ANI owned copyright in the original literary works published on its platforms, even when such work was publicly accessible. However, the Court opined that copyright did not apply to the underlying facts and events reported by ANI.

The Court further stated that ANI had not established memorisation or regurgitation of its works and ChatGPT responses relied upon were not substantially similar to ANI’s articles. Further, electronic storage during training amounted to reproduction under Section 14(a)(i), but OpenAI’s use was prima facie protected as fair dealing under Section 52(1)(a). Therefore, the Court concluded that the purpose and fairness tests favoured OpenAI, and the balance of convenience, irreparable harm and public interest did not support an interim injunction.

Conclusion

The Delhi High Court’s decision provides the first detailed analysis of how copyright law applies to the training and operation of generative AI systems. The judgment does not confer general immunity on AI systems, as verbatim reproduction, memorisation, market substitution or stronger evidence of commercial harm may produce a different result. In conclusion, the decision recognises that the copying and electronic storage involved in LLM training fall within the copyright owner’s reproduction right under Section 14(a)(i). The legal status of AI training is determined by whether the use is authorised or protected under the copyright law.

Authors: Manisha Singh and Shivi Gupta