Artificial Intelligence, News, Report
New European Audiovisual Observatory report examines the copyright challenges of AI training
Artificial intelligence is reshaping the creative landscape at an unprecedented pace, raising important questions for European copyright law about the use of copyright-protected content to train AI systems and the protection of creators and their works. This evolving issue concerns creators, AI systems, AI platforms and their users, as the European legal framework seeks to balance technological development with copyright protection. In July 2026, the European Audiovisual Observatory addressed these questions in its IRIS report Copyright and AI Training, which examines the role of text and data mining (TDM), the relevance of user prompts, and the relationship between AI platforms and their users.
Since its launch , AI has increasingly affected Europe’s creative and business sectors. In this regard, the first chapter provides an overview of the technologies behind generative AI and agentic AI systems, focusing on the challenges arising from the use of copyright-protected works to train AI models. These issues are considered alongside key European regulatory developments, including the Council of Europe Framework Convention on Artificial Intelligence and the EU Artificial Intelligence Act.
One of the central issues in the copyright debate is the text and data mining (TDM), an automated process where software scans and analyzes massive digital datasets to extract hidden patterns, trends, and correlations. In AI development, companies use TDM to scrape billions of human-created works so their algorithms can learn language structures and artistic styles, triggering a massive copyright debate. Moreover,this section examines the European legal framework governing TDM exceptions, tracing their origins to the Copyright in the Digital Single Market Directive (CDSMD) and exploring how these provisions are implemented and applied under national copyright laws. Moreover, it also looks at opt-out mechanisms, transparency obligations, and the ongoing uncertainty over whether TDM exceptions adequately cover the broad scope of AI training activities.
The report further highlights the differences between the approaches adopted by Europe, the UK, and other jurisdictions, particularly in relation to licensing, enforcement, and the balance between the interests of rightsholders and AI developers. Across the various legal systems, policy options vary from strong rights preservation frameworks to extensive exceptions for commercial and non-commercial research.
The third chapter turns to the practical and technical stages of AI training and the copyright issues that these processes entail. In particular, it addresses cases such as the German GEMA v. OpenAI ruling, in which the Munich Regional Court found that copyrighted song lyrics had been memorised in OpenAI’s models and could be reproduced through ChatGPT outputs, raising questions about the copyright implications of using protected works in AI training. Finally, it considers the practice of prompting from a legal perspective, particularly in relation to their role in AI training and the copyright implications that may arise from their use.
The following chapter analyses how AI platforms distribute copyright liability between providers and users in their terms of service. The report provides case studies of major platforms, including Adobe, ChatGPT, Claude, Copilot and Midjourney, to emphasise the significance of transparency, opt-out options and contractual provisions affecting creators and end users alike.
The final chapter brings these findings together and shows that Europe’s copyright framework is struggling to keep pace with the rapid development of AI technologies. In particular, the use of copyrighted data to train AI models remains highly controversial. This is especially evident in the field of text and data mining, which lies at the heart of the current legal uncertainty and, at the same time, raises new questions about human authorship and machine-generated works. Therefore, the need for greater clarity, transparency, and fairness has become increasingly pressing. Ultimately, future debates and legislation will need to strike a careful balance between ensuring access to the datasets required for AI development and safeguarding the rights of copyright holders.
To go further:
European Audiovisual Observatory, “Copyright and AI training”, 2026: