AI and copyright: after much sound and fury, do we know anything yet?
Since generative AI hit the mainstream in 2022, there has been a huge amount of litigation about generative AI and copyright infringement. However, there have been very few substantive judgments so far, and we are still waiting for The Big Questions to be answered.
This article looks at what issues have been decided, and what remains uncertain, from the case law in the UK, EU and the US so far.
The vast majority of litigation is in the US, where there have been summary judgments on the 'fair use' defence, none of which have yet addressed liability for training on crawled datasets. There have been no final US decisions on liability, even at first instance. There has also been a smattering of cases across the rest of the world, where several first instance decisions on liability have been reached. Appeals to the highest courts seem inevitable, so in the absence of further legislative intervention, it will be several years before we have certainty.
The main focus of the litigation has been the AI training process. As such, AI training liability is also the main focus of this article. In addition, some cases consider whether pre-trained models contain infringing copies or whether a model itself constitutes an infringing copy, and whether outputs from AI models infringe copyright. Each of these categories of infringement is addressed in more detail below.
Although most commentators and courts have so far considered (A.) training, (B.) the status of the trained model, and (C.) the outputs from models/tools as separate acts which need to be analysed separately under copyright law, some courts have entertained submissions treating all steps from training to outputs as a single, unitary process. As the 'unitary process' approach complicates or muddies the legal analyses, this article approaches each category of liability separately.
Finally, the status of infringement by users when inputting copyright works as prompts into Chatbots, or when using RAG, raises further copyright questions. New licensing models have started to emerge that are aimed at licensing prompt inputs from the user's perspective. However, as the courts are yet to address these issues, they are not addressed in any further detail in this article.
A. Liability for training
Generative AI models are trained on vast amounts of data. Most AI developers keep the exact composition of their training datasets secret, and the transparency obligations under many new legislative frameworks (such as Article 53 of the EU AI Act) remain uncertain in scope.
It is nevertheless clear that many LLMs and diffusion models were trained using copyrighted content without the benefit of a licence. The big open question is whether such unlicensed training is permitted under copyright law.
Whether 'copies' are actually made during an AI training process is a factual question depending on the technical details of the specific training process. If unlicensed copies are made, the legal question is then whether the entity doing the training had the benefit of an exception from copyright infringement liability under the applicable law. If no relevant exception applies, the entity carrying out the AI training process may be liable to pay damages. The precise scope of each applicable exception, particularly under US copyright law, is a multi-billion dollar question.
Many countries have enacted specific statutory exceptions for 'text and data mining' (TDM), which vary significantly in scope between jurisdictions. Other countries (most importantly the US) apply general 'fair use' doctrines, which can be extended to new technologies on a case-by-case basis. These jurisdictional variations make the question of applicable law very important, which will typically be determined by the physical location of the training process. This will often - but not always - be the location of the applicable servers or data centres (hence the volume of litigation in the US courts).
US 'fair use' decisions
The application of the US doctrine of 'fair use' to AI training is the most important IP issue globally relating to AI. The legal analysis requires a fact-specific assessment of four 'fair use' factors, namely:
1. The purpose and character of the use, with the question of whether the use is 'transformative' being particularly relevant to AI;
2. The nature of the copyrighted work. For example, whether it is a highly creative work or more technical in nature;
3. The amount and substantiality of the portion used in relation the copyrighted work as a whole; and
4. The effect of the use upon the potential market for or value of the copyrighted work.
There have so far been three summary judgments on fair use at District Court level relating to the training of AI. The next AI and copyright 'fair use' hearing is due to be Concord v Anthropic, which is currently scheduled for October 2026.
In the first summary judgment, Thomson Reuters v Ross Intelligence, which considered an AI-powered legal research tool trained on Westlaw headnotes that did not employ a generative AI model, the fair use defence failed. In the other two decisions, Bartz v Anthropic and Kadrey v Meta, which considered Anthropic's Claude and Meta's Llama generative AI models, the fair use defence succeeded.
Each decision employed very different reasoning. In Ross, the judge held that the defendant's use of Westlaw headnotes as AI data to create a competing legal research tool was not transformative, because it did not have a "further purpose or different character" from Thomson Reuters' use (applying the Supreme Court's decision in Warhol). In doing so, the judge emphasised that Ross's AI tool was not a form of generative AI that created new content.
By contrast, the courts in Bartz and Kadrey both held that the training of general purpose generative AI models was 'highly, 'exceedingly', 'spectacularly' or 'quintessentially' transformative. The judge in Kadrey noted that Llama could "generate diverse text and perform a wide range of functions", whereas purpose of the plaintiffs' books was only "to be read for entertainment or education". This suggests that models intended (at least at the time of training) to be applied to a broad range of general purpose tasks are more likely to be 'transformative' than those aimed at a much narrower, specialised domain.
Fair use factors 2 and 3 had only limited impact on the outcome of the analysis in both Bartz and Kadrey, despite the fact that the entirety of each applicable copyright work had been copied, and that the works were highly creative. Instead, factor 4 – the effect on the market - is likely to be the most important factor in generative AI cases. Judge Alsup in Bartz, who issued his judgment first, took a very different approach to Judge Chhabria in Kadrey regarding factor 4.
Whilst good analogies can be incredibly helpful in simplifying legal cases involving frontier technologies, AI and copyright cases and commentary are plagued by bad analogies. Judge Alsup employed the following analogy in his decision on the fourth fair use factor:
"[the copyright owners] contend generically that training LLMs will result in an explosion of works competing with their works — such as by creating alternative summaries of factual events, alternative examples of compelling writing about fictional events, and so on. This order assumes that is so […] But Authors’ complaint is no different than it would be if they complained that training schoolchildren to write well would result in an explosion of competing works. This is not the kind of competitive or creative displacement that concerns the Copyright Act. The Act seeks to advance original works of authorship, not to protect authors against competition."
Judge Chhabria was not shy in offering his views on the merits of that analogy:
"Judge Alsup focused heavily on the transformative nature of generative AI while brushing aside concerns about the harm it can inflict on the market for the works it gets trained on. […] But when it comes to market effects, using books to teach children to write is not remotely like using books to create a product that a single individual could employ to generate countless competing works with a miniscule fraction of the time and creativity it would otherwise take. This inapt analogy is not a basis for blowing off the most important factor in the fair use analysis."
However, in finding for Meta, Judge Chhabria was similarly critical of the arguments led by the copyright owners:
"In connection with these fair use arguments, the plaintiffs offer two primary theories for how the markets for their works are affected by Meta’s copying. They contend that Llama is capable of reproducing small snippets of text from their books. And they contend that Meta, by using their works for training without permission, has diminished the authors’ ability to license their works for the purpose of training large language models. As explained below, both of these arguments are clear losers. Llama is not capable of generating enough text from the plaintiffs’ books to matter, and the plaintiffs are not entitled to the market for licensing their works as AI training data. As for the potentially winning argument—that Meta has copied their works to create a product that will likely flood the market with similar works, causing market dilution—the plaintiffs barely give this issue lip service, and they present no evidence about how the current or expected outputs from Meta’s models would dilute the market for their own works."
Although he found in favour of Meta, Judge Chhabria set out some guidance for future plaintiffs regarding market dilution as to how a successful infringement case could be framed by reference to the downstream uses to which the trained LLMs are applied. However, the largely untested concept of copyright dilution has been the subject of detailed, highly sceptical academic commentary, that even questions its constitutionality. Given that most AI and copyright trials have yet to reach a first instance decision, let alone Circuit appeals or the Supreme Court, the position on fair use is unlikely to be settled for a long time.
One of the most noteworthy aspects is the very different approaches they took to the training of models on so-called 'shadow libraries' of books – essentially online repositories of pirated books. Whilst Judge Alsup found that Anthropic could rely on the fair use defence in respect of digitised copies of physical booked which it had purchases and scanned, he found the use of online pirated repositories in training could not benefit from the fair use defence. The proceedings then moved quickly to settlement, with Anthropic agreeing to pay the plaintiff class at least USD 1.5 billion in damages.
By contrast, Judge Chhabria was clear that the sourcing of books from shadow libraries did not give the copyright owners an "automatic win". Rather, he outlined points where the provenance of the training materials could be relevant to the fair use analysis, but ultimately made no finding on these points because the copyright owners had not submitted any relevant evidence.
The Bartz and Kadrey decisions contain a lot of additional 'food for thought', which for reasons of brevity can't be covered in full detail in this article.
EU law on statutory training exceptions
The EU's Copyright in the Digital Single Market Directive of 2019 introduced two statutory 'text and data mining' (TDM) exceptions, which have since gained greater prominence following the passing of the EU AI Act. There has been substantial academic commentary as to whether the EU TDM exceptions, which were enacted before generative AI entered mainstream awareness, actually cover the training of most AI models. However, the EU AI Act and its associated codes of practice assume that the TDM exceptions do cover the training of generative AI models, and this view is supported by the early case law.
The first TDM exception permits 'research organisations and cultural heritage institutions' to carry out TDM for the purposes of scientific research. The second TDM exception is much broader in scope, and includes TDM carried out by commercial companies for commercial purposes. Copyright owners can 'opt out' of the commercial TDM exception, but not the research TDM exception.
To opt out, rightsholders must "expressly reserve" their rights "in an appropriate manner, such as machine-readable means in the case of content made available online". The creates uncertainty for rightsholders and AI developers alike; rightsholders are not told where or in what form to reserve their rights and, in turn, AI developers don't know for certain where they need to look for opt-outs, and how they can build a reliable automated process to read opt-outs in the absence of any legally-mandated standard or format. Impressive work has been done by International Standardization bodies to create workable and technically-sound means of opting out, and the Google-Extended product token has been available for several years, but the law and market knowledge/market practice currently lag the technical capabilities.
The national case law has so far addressed very few aspects of the EU opt-out regime, and there have been no EU-level decisions. The first member state-level decision was the Hamburg District Court's judgment in Kneschke v LAION in September 2024. The Court held that LAION, a non-profit organisation had created its AI dataset for AI training for non-commercial purposes, and that it was entitled to rely on the TDM research exception. The exception was held to cover preparatory activities when compiling the dataset (including web-scraping), and was not limited solely to the model training process itself. The Court also held that, even if downstream uses of the compile dataset by third parties were commercial in nature, that did not affect the non-commercial status of LAION's own activities.
Although strictly obiter, the Hamburg court also stated that a reservation of rights expressed in natural language (as opposed, for example, to the robots.txt file) would be effective.
Shortly after the first LAION decision, the Amsterdam District Court ruled in the HowardsHome case that, even though the claimant attempted to prevent scraping via the robots.txt file on its website, the optout was not effective because the protocol only targeted and blocked specific well-known AI bots only and was not a general reservation of rights effective against all crawlers. In adopting an extremely narrow interpretation of the opt-out requirements, the judgment stated that opt-out mechanisms must explicitly identify their targets, and noted that implied or generic reservations would be insufficient.
The Danish Maritime and Commercial High Court then found, in a preliminary injunction decision in October 2025, that an opt-out written in natural language within a website data and privacy policy was effective in BoligPortal v ReData (similar to the first instance decision in LAION).
However, on appeal to the Higer Regional Court of Hamburg and the Eastern High Court in the LAION and BoligPortal cases respectively, the natural language reservations were held not to constitute effective opt-outs. The Danish Court had held that, to be "machine-readable", the reservation also had to be capable of being acted upon, so that an automated scraping system leaves the content un-used. The Hamburg Court held that, based on the state of technology available at the time of the relevant scraping activity in late 2021, a natural language reservation could not be considered machine-readable. However, a Court might not make the same decision based on the state of technology available today. Machine 'actionability', and not just 'readability', is likely to remain important, however.
The European Commission launched a consultation in December 2025 on the implementation of the opt-out for the purposes of the EU Act obligations. The consultation closed in January 2026, but no outcome or findings have yet been published.
UK law
Current UK copyright legislation provides a TDM exception that is limited to training carried out as part of research for a non-commercial purpose. The UK Government has carried out two consultations since 2021 regarding potential reforms to the TDM exception (the Report resulting from the most recent consultation having been published in March 2026), but there are currently no active proposals for reform.
The Getty Images v Stability AI case has generated enormous amounts of commentary, but the November 2025 judgment does not address liability for AI training at all, in any jurisdiction.
What has not been addressed so far?
Several LLMs are believed to have been trained on datasets provided by Common Crawl, and web crawling in general is one of the primary sources of content for AI training. Rather surprisingly, no substantive judgment of any US or European court has yet addressed liability for training based on content scraped from the web, but given the number of pending cases, it is only a matter of time before we start seeing judgments on this key topic.
B. Do the models contain or constitute infringing copies?
Few courts have made direct findings on this point. There is much technical debate on the degree to which models "memorise" content, and the steps can be taken during training to minimise or prevent memorisation from occurring. The legal significance of any such memorisation also remains in dispute.
In the case law decided to date, the clearest judgment is from the English High Court in Getty Images v Stability AI. On the evidence, it was found that the defendants' pre-trained model did not itself store the data on which it had been trained, and the Court recorded that there was no evidence of any specific copyright work having been memorised. Whilst those findings would likely have precluded a finding of primary infringement under UK law, the claimants still asserted secondary infringement on the basis of an esoteric section of the UK's Copyright, Designs & Patents Act.
In the UK, it is an infringement to import, possess in the course of business, sell or distribute an "article" which is known to be an "infringing copy" of a copyright work. An "article" is an infringing copy if (a) it has been imported from outside the UK; and (b) its making in the UK would have constituted an infringement of the copyright in the work in question. Although this was a clever argument by the claimant's lawyers, which was aimed at creating liability in the UK for a model that had been trained abroad, the court ultimately held that if the model itself was not (or did not contain) a copy of a work, it could not be an infringing copy.
In the US:
- The summary judgment in Bartz v Anthropic took as an assumed fact (i.e. the court did not itself make such a finding) that a fully trained Large Language Model (LLM) retained a "compressed" copy of each work it had trained upon. However, the court also considered the significance of filters applied on top of the model when accessed by users, stating that, "[w]hen each LLM was put into a public-facing version of Claude, it was complemented by other software that filtered user inputs to the LLM and filtered outputs from the LLM back to the user […]. As a result, [the copyright owners] do not allege that any infringing copy of their works was or would ever be provided to users by the Claude service." The Judge appeared to regard the possibility of memorisation within the model as irrelevant for infringement purposes if effective filters prevented users from ever accessing those works.
- The judgment in Kadrey v Meta stated as follows: "[Meta] also post-trained its models to prevent them from “memorizing” and outputting certain text from their training data, including copyrighted material. These training efforts, which Meta calls “mitigations,” appear to have been successful. Meta’s expert witness tested them using a method designed to get LLMs to regurgitate material from its training data (which Meta calls “adversarial prompting”). Even using that method, the expert could get no model to generate more than 50 words and punctuation marks (that is, “tokens”) from the plaintiffs’ books."
By contrast, the Munich District court in GEMA v OpenAI held that OpenAI's models memorised the works on which it was trained. It further held that, because the copyright works in suit (song lyrics) could be reproduced almost verbatim, those songs were stored in the model; even though the works were not saved in the same manner as files in conventional computer memory, in the Court's view the fact that the song lyrics could be reconstituted from the model demonstrated that the model was functioning in a manner analogous to "lossy" compression.
The judgment goes into a surprising amount of detail about the differences between, on the one hand, what the model itself does and produces and, on the other hand, the steps performed by the software layers built on top of the model. The Court held that the fact that the model (as opposed to the ChatGPT service which implemented various filters) was capable of reproducing content was evidence that the training process had resulted in memorisation of works, which fell outside the scope of the training exception. Remarkably, the Munich Court did not address the location of the training process, or how that location might affect the law applicable to the training process.
The wider legal consequences of those findings were not clear from the GEMA v OpenAI judgment, which did not clearly distinguish between liability for training, the use of the model or the generation of outputs by the operator of a managed AI service. These matters will hopefully be clarified on appeal. The forthcoming judgment from the Munich Regional Court in GEMA v Suno, due on 31 July 2026, may also provide clarification.
Memorisation is also at issue in several other pending cases, none of which have resulted in substantive judgments at the time of writing.
C. Liability for infringing outputs
The assessment of whether a specific AI output infringes third-party copyright does not raise new legal questions; a given output needs to be compared against the work or works that are alleged to have been copied, and the existing national law tests for determining copyright infringement can be applied.
The challenges come when linking any specific work within a training dataset to a given output, and assessing whether there has been actionable copying of that specific work, rather than other similar works or works which depict or describe the same subject matter. When dealing with training datasets that may run to billions of individual input works, this can be very difficult, although where a specific subject matter appears only rarely or uniquely within a single document in the training dataset, the chances of being able to demonstrate infringement based on the degree of similarity to a single input work are higher.
In the UK Getty Images v Stability AI case, the claimants initially alleged that several output images that had been generated using the Stable Diffusion AI image generation tool infringed copyright in the specific photographs contained in the training dataset. However, the copyright infringement claims based on outputs were abandoned before the end of trial, and they are not addressed in the judgment as a result.
The highest risk outputs are those depicting fictional characters, such as video game or cartoon characters. In those cases, even if the output cannot be linked to a single image in the training dataset, the depiction of the character in a recognizable form may be sufficient for infringement purposes (i.e. because the character itself is arguably a copyright work, belonging to a single identified owner). Many commercial AI tools now apply filters to prompts and to outputs to reduce the highest risk categories of outputs.
AI coding tools also raise new risks. In particular, because open-source code is readily available, and a particular piece of open-source code may appear numerous times in a training dataset, an AI tool might return recognizable open-source code in response to a user's prompt. Because of this, many AI code editors now have built in open-source scanning functionality.
Identifying who is responsible for the infringement is also important. Infringement liability attaches to the person who carries out the act(s) restricted by copyright. In the case of a local copy of an AI model run on a user's device, this analysis will often be straightforward. Where an AI application is part of a managed service hosted in the cloud, where the outputs are generated by a provider in another country, the analysis becomes more complicated as to who is liable for infringement and which law applies to the infringing acts.
The question of whether an AI model or tool provider can be liable for "authorising" infringements committed by their users has yet to be addressed in any substantive judgments. In the analogue era, products which were capable of both infringing and non-infringing uses (such as double tape-decks) typically did not give rise to liability on the part of manufacturers for 'authorising' infringement carried out by users. By contrast, in the UK cases on filesharing and torrent sites, the operators of websites which provided searching and indexing facilities, and guides for accessing pirated content on other sites, were held to have authorized infringements committed by users. The key criterion is likely to be the degree of control which the tool provider retains and whether he has taken any steps to prevent infringement (such as the implementation of effective prompt and output filters).
Any finding that AI tool providers are liable for contributory infringement for acts committed by their users in the US has now been made less likely by the recent Supreme Court decision in Cox v Sony. The Court held that a service provider must either 'actively encourage' infringement or have 'tailored its service to infringement' before liability for contributory infringement would arise, which is likely to be difficult to show in the case of general purpose AI chatbots and tools.