Image: gguy / shutterstock.com

Artificial intelligence needs data—a great deal of it. Books, newspaper articles, websites, and academic texts help make language models increasingly powerful. But this is precisely where one of the biggest legal debates of our time begins: Is AI allowed to use copyrighted works for training—and if so, under what conditions?

A legal case in the U.S. is now drawing global attention. A court has definitively approved a settlement of approximately 1.5 billion U.S. dollars in the dispute between the AI company Anthropic and a large group of authors and publishers. This is one of the largest settlements ever reached in a copyright case and sends a message that extends far beyond the United States.

The real problem wasn't the AI's learning process

At first glance, one might think the court ruled that training an AI is, in principle, illegal. However, that is not exactly what happened.

Rather, the proceedings centered on the question of where the training data came from. According to the court’s findings, Anthropic is alleged to have obtained millions of copyrighted books from illegal sources and stored them in a central digital library. It was precisely this procurement process that played a decisive role in the proceedings.

What is interesting here is the court’s legal classification: In an earlier ruling, the actual training of a language model was deemed to constitute a fundamentally new and independent use. The storage and use of copies obtained unlawfully, on the other hand, were viewed much more critically. It is precisely this distinction that is likely to set the tone for many other cases.

Why Anthropic Agreed to a Billion-Dollar Settlement

The settlement, which has now been confirmed, amounts to approximately 1.5 billion U.S. dollars. This brings to a close one of the first major copyright lawsuits against a company in the AI industry before a protracted trial over potential damages could take place.

For Anthropic, this settlement was likely a predictable financial outcome. Without a settlement, the company could have faced claims for damages in the worst-case scenario that would have been many times higher. This illustrates the financial risks that can arise when copyrighted content is used without the necessary rights.

A large number of eligible authors and publishers have accepted the settlement and applied for compensation. However, some rights holders consider the payments to be too low and are continuing with their own legal proceedings. This means the legal debate surrounding AI and copyright is far from over.

The lawyers, too, suddenly found themselves in the spotlight

It wasn't just the amount of the settlement that sparked debate. There was also controversy over the fees paid to the attorneys involved.

While significantly higher fees had originally been requested, the court substantially reduced the amount. In the judges’ view, a larger portion of the settlement amount should benefit the affected authors and publishers. In doing so, the court made it clear that even in cases involving exceptionally high settlements, the interests of the actual entitled parties must take precedence.

Why This Ruling Is Also Important for German Companies

Although the case was heard in the United States, companies in Europe are following developments very closely.

More and more companies are developing their own AI applications or using existing language models for internal processes. This regularly raises the same question: Where does the data used to train or enhance these systems come from?

These issues are becoming even more important, particularly in Europe. In addition to copyright law, new requirements under the European AI Act now apply, calling for greater transparency and accountability in the use of artificial intelligence. Companies should therefore not only focus on the performance of an AI system, but also on whether the data used was obtained lawfully.

The Anthropic case clearly demonstrates that not every technical innovation is automatically legally permissible. Anyone who uses large amounts of data should be able to trace at any time where the data comes from and whether the necessary rights of use exist.

The real message lies between the lines

The billion-dollar settlement is far more than just a high-profile legal dispute. It underscores the fact that courts are increasingly distinguishing between technological innovation and the protection of intellectual property. AI should be able to continue to evolve—but not at any cost.

For companies, this means that proper compliance doesn’t begin with the finished AI product—it begins much earlier, with the selection of the data on which the artificial intelligence is built. Companies that handle this process properly not only reduce legal risks but also strengthen the trust of customers, business partners, and investors.

Innovation requires speed—but not shortcuts

This case highlights a fundamental problem facing the entire AI industry. Many companies wanted to develop the most powerful artificial intelligence as quickly as possible, apparently hoping that legal issues could be resolved later on. This approach is proving less and less effective.

Anyone who wants to make billions with AI should also respect the billions of intellectual effort that went into it. Books are not a free resource just because they are available in digital form. At the same time, however, the pendulum must not swing too far in the other direction. If every use of content becomes practically impossible, it will stifle innovation. The future, therefore, does not lie in gray areas or class-action lawsuits, but in fair licensing models, clear rules, and a modern copyright law that protects creators without stifling technological developments.

Subscribe to the newsletter

and always up to date on data protection.