Last updated July 27, 2026
Author: Gerben Nerinckx, Commercial Lawyer & Fractional General Counsel
You have found the perfect AI engine to build into your product. It is fast, it is clever, the demo won the room. But there is just one thing you cannot do: look inside and check how it was built. And that blind spot may very well become your Achilles’ heel.
Every AI has a past
TADAAAA… not quite. Generative AI engines and large language models do not spring fully formed into the world. They are trained. Before a model can draft your morning brief, generate a cartoon or answer your customers’ questions, it must be fed enormous quantities of data, scraped from the open internet, licensed from third parties, or otherwise gathered. The quality of the finished model is a direct function of the quality and breadth of that data.
Which raises the one question every AI vendor would rather its customers did not ask: where did the training data come from, and was the vendor entitled to use it?
A fight that is far from over with copyright leading the way
That question is currently being fought out in courtrooms and regulators’ offices on both sides of the Atlantic, and it is nowhere near settled. The sources and methods used to train AI systems are being litigated, regulated and argued over, and the outcomes are far from settled. One of the principal battlegrounds is copyright law.
Two recent US cases illustrate the shape of the fight.
In Thomson Reuters v. ROSS Intelligence (District of Delaware), a legal research start-up built a competing product by training on editorial content drawn from Westlaw, the West headnotes and the West Key Number System. The court held that this copying was not protected by fair use, siding with the copyright owner. Note that an appeal is pending.
In UMG Recordings v. Suno (District of Massachusetts), a group of major record labels sued a generative music service, alleging that it had copied their sound recordings wholesale in order to train its model. Suno effectively conceded that its training data included the labels’ recordings, but defended on fair use and on a feature of US law specific to sound recordings. That litigation is still ongoing. Warner settled and took a license, while Universal and Sony press on (for now).
For those who would like to read up, some of the most relevant court documents can be found by following the links below. A word of caution though: These are only two examples. They sit within a much larger wave of AI-training disputes now moving through the courts, and any of them may shift the landscape. Also, it is tempting to read these decisions and extract general principles. Resist the temptation at least outside the United States. Both cases were argued and decided within the US legal system, and their reasoning does not transpose cleanly to a European jurisdiction (as an example, most European jurisdictions are not familiar with the fair use argument as it exists in the US legal system, and vice versa you’ll not hear about database protection as a sui generis right being cited in US courts). But that’s beside the point. Building the next super-duper piece of software relying on a third party AI engine or LLM, are you the one wanting to talk legalese in court?
But other ugly animals are surfacing from the woods as well
Copyright is only half the story. Scraping the open web almost always sweeps up personal data as well, which brings data privacy laws (to us European lawyer geeks known as GDPR) into play. The European Data Protection Board just opened a public consultation into draft guidelines aimed squarely at web scraping to train AI, and they are in no mood to treat “it was publicly available” as an answer.
Why this is your problem, not just your vendor’s
Now think about YOUR product. When you license someone else’s AI engine and embed it in your own product, you inherit all of this uncertainty with none of the visibility.
You did not build the model. You did not assemble the training data, and you have no practical way to audit it. You cannot verify that the copyrights were cleared, or that the personal data was processed lawfully. You are taking the vendor’s word for it, if the vendor offers a word at all. And yet, if a rightsholder or a supervisory authority comes knocking, you may be the one left holding the product they are pointing at.
That uncertainty is real, and factually you cannot resolve it. So do not try to solve it with facts. Solve it with the contract.
The tool for the job: the indemnity clause
Uncertainty you cannot resolve factually, you can allocate contractually. That is exactly what an indemnity clause is for.
An indemnity clause is a promise by one party (in this case the licensor of the AI engine) to defend, or to reimburse (in legalese ‘to hold harmless’) the other (this is you) against/for certain claims or losses (think the damages, fines, the legal costs etc.) associated with third party claims (in this case for infringing copyright laws, violating data protection laws or more generally ‘the law’). In plain language: ‘You promise to defend me against any claim associated with your wrongdoing’.
For an AI licence, that promise belongs with the vendor. They built the model, they chose the training data, and they are the only party in a position to assess the risk, and to price it into the fee.
In practice, it looks something like this:
‘The Licensor shall defend and indemnify the Licensee against any third-party claim or regulatory action alleging that the AI engine, or the data used to train it, infringes any intellectual property or database right, breaches data-protection law including the GDPR, or otherwise fails to comply with applicable law, and shall bear all resulting damages, fines and reasonable legal costs.’
But don’t do this at home
Negotiating a proper indemnity is not boilerplate housekeeping; for a start-up or SME it is one of the more important terms of the deal, and it comes with razor-sharp hooks, such as
- Scope. What are you looking to be protected against, and what is your licensor prepared to offer? Only protection against IP infringement, also data-protection claims, or general breaches of ‘the law’ and/or the for the answer to the ultimate question of life, the universe and everything (disclaimer: you’ll notice that, for training purposes, part of my brain read a popular book, so attribution for this phrase goes to Adams, Douglas.The Hitchhiker’s Guide to the Galaxy. Pan Books, 1979).
- Defense. Do you entrust your licensor to run the defense for you, or do you want a say in any settlement so that the licensor cannot get rid of the claim by leaving liability sitting with you?
- The cap. Are you OK with the maximum amount of coverage being limited to the fee you paid for licensing the engine? Prepare to get wet when it starts raining.
The short version
AI has to be trained, and how it is trained is a legal minefield that is nowhere near mapped. If you are licensing someone else’s AI engine to build into your product, you need to buy some sort of insurance, and that comes through a properly scoped indemnity. This is how a start-up or SME hands the obvious risk back to the party that created it…the vendor of the AI engine that powers your product.
Links:
Thomson Reuters v. ROSS Intelligence: Complaint, Defense, Ruling
UMG Recordings v. Suno: Complaint, Defense
EDPB Draft Guidelines

