AI Energy

$1.5 Billion in Content Cost: Anthropic and its resolution redefine the boundaries of responsibility for AI training materials.

By Kevin Guo
20 min

GFM Editor's Note

On July 20, 2026, the U.S. District Court for the Northern District of California formally approved the $1.5 billion copyright settlement reached between Anthropic and a group of authors and publishers. This is the largest known copyright class-action lawsuit settlement in the United States, and the first training material copyright case to be substantially resolved with such a huge sum of money since the rise of generative artificial intelligence.

The court's final approval marks a significant procedural milestone in this nearly two-year-long case, and for the first time, the legal source of artificial intelligence training data has come into the public eye at a cost of billions of dollars.

This case could easily be interpreted as a conflict of interest between a tech company and an author. However, it delves deeper: how can an AI company obtain training data, and how should it prove its right to retain, copy, and use that data?

The court did not completely deny the use of copyrighted works for model training in artificial intelligence, nor did it rule that fees must be paid to the authors for every model training session. The line drawn by the court lies in the process of data entering the enterprise's system. The use of the work may be transformative, and the method of acquiring the work still needs to be examined under copyright law.

In recent years, the artificial intelligence industry has focused primarily on model capabilities, computing power supply, and product competition. Training materials are often viewed as resources naturally existing on the internet. Books, news, photographs, recordings, and research findings are collectively referred to as "materials," and the specific people behind these works—authors, publishers, photographers, journalists, and researchers—have gradually lost their clear identities in this technical terminology.

The Anthropic case serves as a reminder to the entire industry that once content enters a business model, its source can no longer be ambiguous. Artificial intelligence can learn from long-accumulated human knowledge, and companies operating these models must account for how data enters their systems.

Today, this responsibility has a specific figure: $1.5 billion.

(Image caption) U.S. District Court for the Northern District of California. The $1.5 billion Anthropic copyright settlement has been granted final judicial approval, defining new legal boundaries for the legitimate source and rights and responsibilities of artificial intelligence training materials.


What did the court approve on July 20?

In August 2024, authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson filed a class-action lawsuit against Anthropic, accusing the developer of the Claude AI model of downloading, copying, and storing a large number of copyrighted books without permission.

The case took a crucial turn in 2025. The presiding judge, William Alsup, ruled that Anthropic's use of books to train its large language model was transformative in the specific circumstances presented and thus protected under the fair use doctrine. However, the court also held that the company's acquisition of millions of pirated ebooks from online shadow libraries like LibGen and their storage in a central digital library did not automatically acquire legality simply because some of these books were later used for model training.

The case could have gone to a damages trial. If the jury had found Anthropic guilty of willful tort, the potential liability could have far exceeded the final settlement amount. Faced with enormous litigation risks, the two parties reached a $1.5 billion settlement in 2025, which the court preliminarily approved.

The settlement did not formally take legal effect until July 20, 2026, when Federal District Court Judge Araceli Martínez-Olguín signed the final approval order. Final approval means that key issues such as the scope of the work, the claims process, the distribution of proceeds, and legal fees have been confirmed by the court, thus ushering this landmark case into the enforcement and payment phase.

This is also the reason for re-examining the Anthropic case today. The 2025 ruling set legal boundaries, while the final approval on July 20, 2026, translated those boundaries into a real, enforceable corporate cost.

One book, two origins

To train its large language model, Anthropic built a massive e-book database. During the court proceedings, the books were divided into two distinct categories based on how they were acquired.

One type of book involved Anthropic purchasing physical copies through legitimate channels. After acquiring the books, the company disassembled them, scanned the pages into digital files, stored them in an internal database, and used them for model training. The court held that, under the specific circumstances presented in this case, converting legally purchased books into digital versions and using them for training might be protected under the fair use doctrine.

Another type of book came from unauthorized online databases. Anthropic did not acquire these electronic files through normal purchasing channels, nor did it obtain permission from the authors or publishers. Some of them were used in the model training process, while others were stored permanently in the company's central digital library, awaiting potential future use.

The same training system connects on one end to paper books that the company paid for and manually removed the bindings from, and on the other end to electronic archives downloaded in bulk from a shadow library. They ultimately enter the same database, are read by the same model, but have completely different legal origins.

In its 2025 ruling, Judge Alsup held that the transformative nature of model training cannot retroactively alter how the data was originally obtained. No matter how advanced the technology a work is later applied to, it cannot erase the history of how it entered the enterprise's systems. The court allowed Anthropic to claim fair use for some of its model training activities but required the company to continue to face liability for infringement arising from the establishment of a pirated central library.

This ruling examines the different stages of AI training separately. How a company uses a work, and how it acquires that work, each has its own legal boundaries. The level of technological innovation in a model cannot provide an unrestricted license for the data collection process.

(Image caption) A company disassembled and scanned legally purchased physical books into electronic archives for use in training artificial intelligence models. The court held that such behavior may be protected as fair use under certain conditions.


How was $1.5 billion formed?

The final settlement involved approximately 482,000 eligible works, with a nominal value of approximately $3,000 per work. More than 90% of the eligible works had already been claimed before the court issued its final approval. This is a remarkably high claim rate for a class-action lawsuit involving hundreds of thousands of works.

Behind these numbers is the fact that hundreds of thousands of works have been placed into the same judicial liquidation process for the first time. The titles, authors, publishers, copyright registrations, and ownership of rights need to be checked one by one. Rights that were originally scattered among different authors and publishing houses have gained common enforcement power due to the class-action lawsuit.

The $1.5 billion award is not the result of a court-assessed market value of each book in question after the full trial. Anthropic also did not acquit on all charges through the settlement. This amount comprehensively reflects the risks of statutory damages, litigation costs, the difficulty of proving the case, the uncertainty of a jury verdict, and the potential commercial impact of the case's continued development.

It cannot be regarded as a uniform price for all AI training content. More accurately, it is a judicial price signal formed after a concentrated settlement of historical responsibilities.

In the past, the unclear source of training materials was merely an abstract compliance risk. Many companies knew that certain materials might have copyright issues, but they struggled to determine when and how much this risk would appear on their balance sheets.

The Anthropic case transformed this uncertainty into a figure visible to boards of directors, investors, insurance companies, and finance departments. $1.5 billion is enough to alter a company's understanding of data management costs and will influence how capital markets will value the data assets of AI companies in the future.

Before the book entered the model, whose hands did it pass through?

The training of artificial intelligence, as commonly imagined, resembles a massive reading process. The model reads books, learns language, knowledge, and structure, and then generates new content.

This screen omits a significant amount of work that must be done before reading begins.

Data needs to be found, downloaded, copied, cleaned, deduplicated, and indexed before it can be handed over to the training system. It may flow between different departments, cloud servers, and outsourcing agencies, or it may continue to be stored after the first training, waiting to be incorporated into the next model, the next test, or the next commercial product.

Every transfer and reuse may encounter different boundaries of rights. Buying a physical book does not mean you can arbitrarily create and distribute electronic copies; obtaining the right to publish a photograph in the press does not mean you can use the photograph for long-term training of image models; obtaining the right to translate an article does not necessarily include the right to permanently place the translation in a commercial database.

Throughout this process, the engineering team typically focuses on whether the data is sufficient and in a usable format, while management is concerned with whether the model can be completed on time and the product can be launched as soon as possible. The legality of the data sources and the scope of authorization are often left to be addressed last.

As the scale of the model continues to expand, the risks accumulated in this approach also increase. A claim for a single work may be insignificant to a large company, but when hundreds of thousands of works are accumulated, it can create legal liabilities that the company cannot bear.

Large datasets can enhance model capabilities, but they can also aggregate previously scattered copyright disputes into massive debts. This is a less-discussed aspect of the scale effect in artificial intelligence.

(Image caption) The act of bulk downloading pirated ebooks from online shadow libraries and establishing a central digital library became a key point of contention in the Anthropic case, where the court determined that it could not be automatically legalized by subsequent training.


Why couldn't the author follow up in the past?

Writers, journalists, and photographers have long faced a real dilemma: they find it difficult to know whether their work has been included in a particular model.

Artificial intelligence companies typically do not publicly disclose complete lists of training materials. Even if an author suspects that their work has been used, it is difficult to prove when the company downloaded it, in which database it was stored, and in which model version it was used.

For an independent author, tracing the usage of a book often involves more time, legal fees, and technical evidence-gathering costs than the potential compensation. Many small but real rights remain unenforceable.

Class action lawsuits have altered this power dynamic.

When authors and publishers compile lists of works, copyright registrations, publication materials, and download evidence, previously scattered individual rights are transformed into a single claim that can be adjudicated by the courts. The settlement management process requires identifying each work, confirming the rights of the author and publisher, and handling situations where multiple rights holders claim the same work.

This work itself also provides an institutional model worth observing for the journalism industry.

Content needs a clear identity to enter the authorization process; the rights holder needs to be identified so that the revenue should be distributed to whoever benefits from it. When the work, the rights holder, and the usage record remain unclear for a long time, even if the content is used extensively, it is difficult to convert it into income that the author can actually obtain.

AI companies will usher in a new era of data auditing.

Large AI companies often emphasize that models do not save every training book word for word like a hard drive; the training process is, in some ways, closer to how humans learn language and knowledge.

This claim has a place in the legal analysis of fair use, but it cannot replace a company's due diligence on the source of its data.

In the future, AI companies will need to answer very specific questions. They will need to know who provides the training materials and whether the provider is authorized; whether the original download records are still preserved; how many model versions a single license can be used for; whether the materials can be transferred to cloud service providers, outsourcing agencies, and affiliated companies; whether the company can locate and handle the relevant content if the author withdraws the license; and whether the old materials remain in the system after the model is updated.

These issues will gradually permeate investment and mergers and acquisitions, insurance underwriting, corporate auditing, and board risk management.

Having a vast amount of data does not equate to possessing high-quality data assets. When a large amount of content lacks clear sources and authorization certificates, the sheer size of the data can become a reason for valuation discounts. In contrast, a smaller, more reliable, clearly defined, and continuously updated database often has more stable commercial value.

The artificial intelligence industry has traditionally compared the number of parameters, computing power costs, and model performance. Following the Anthropic case, the legal quality of data will gradually become part of a company's competitive advantage.

(Image caption) In the artificial intelligence data center, books and text materials enter the model training system in the form of data streams, highlighting the importance of the legality of the training materials and the company's compliance responsibility.


The power dynamics within journalism are more difficult to dismantle than those in books.

Book publishing has a relatively mature system of copyright registration, ISBN, and publishing contracts. News content is usually more complex in composition.

A news report may simultaneously contain text written by journalists, news agency materials, photographs by photographers, documents provided by interviewees, social media screenshots, publicly available information, interview recordings, translations, and data charts. What readers see is a complete article, but multiple sets of power dynamics may exist within it.

A simple interview recording can illustrate the difficulties involved.

The interviewee agrees to speak publicly, the journalist compiles the conversation into a report, and the media obtains the right to publish the article based on an employment or cooperation relationship. At this stage, the relevant rights are usually quite clear.

The situation changes if the recording is later edited into a podcast and used to train a speech model. The interviewee initially agreed to have the conversation published, but this didn't necessarily include converting their voice into a synthetic model capable of generating new sentences.

With each additional use, the boundaries of existing consent may be pushed outward. The person involved may not even be aware that their voice, image, or text has entered a new technological system.

Traditional news production primarily serves a one-time release. The editorial department verifies facts, confirms byline, and handles image sources; once the article is published, most of the work is complete.

Artificial intelligence has extended the lifecycle of content. A report can be translated into multiple languages, made into audio and video, entered into search databases, or licensed to research platforms or modeling companies. When content is repeatedly used, licensing relationships that were initially maintained by an email, a chat log, or a verbal promise are no longer sufficient to support long-term and complex commercial use.

Why does GFM need a more complete content asset ledger?

There is a direct and specific relationship between the Anthropic case and the GFM content asset ledger.

The risks revealed by this case stem from the fact that companies lack a clear system to prove where their data came from, who owns it, and to what extent it is permitted to be used. The ledger established by GFM aims to record these issues before content enters translation, audio recording, database creation, commercial licensing, or model training.

The author, title, and publication date only prove when and by whom an article was published; they cannot fully explain whether the text, images, audio recordings, and data within the article can be incorporated into a new product. Therefore, the record book needs to preserve the original source of the content, including original articles by journalists, contributions from partner organizations, publicly available documents, information provided by interviewees, purchases from commercial databases, or AI-assisted generation, and retain the original archives, acquisition dates, provider identities, and necessary version records.

Ownership of rights also needs to be recorded separately for each content element. The text, images, videos, audio recordings, and translations of a report may belong to different people. Media outlets that acquire the right to publish an article may not simultaneously acquire the right to sublicense images, audio recordings, or third-party materials.

The scope of authorization needs to be more specific than before. Website publication, social media dissemination, commercial reproduction, translation, voice production, database inclusion, and model training all have different impacts on the rights holder. A vague "agree to use" statement is unlikely to cover all new technological applications.

Contracts, authorization letters, emails, submission agreements, payment records, and platform terms should also be traceable to the specific work. Years later, when content is incorporated into a new product that did not exist at the time of creation, the editorial department should still be able to find out what rights were initially obtained, whether the authorization has expired, and whether further consent is required for any new uses.

This ledger is not an additional piece of corporate publicity attached to the article. It is an institutional response to the legal issues revealed in the Anthropic case, a response that can be implemented within a news organization.

(Image caption) News media have established a complete content asset ledger and rights management system to record the source, scope of authorization and usage restrictions of each piece of content in order to meet the new requirements for content rights confirmation in the era of artificial intelligence.


Property rights confirmation cannot rely solely on timestamps.

In recent years, content ownership verification has sometimes been simplified to adding a timestamp to a work to prove that someone submitted or published the content at a certain time.

Timestamps have evidentiary value, but they can only prove part of the facts.

An individual can register the date of a photo they don't own; an organization might publish an article provided by a partner but lack the right to re-authorize it to an AI company. The initial publication date can establish a point of evidence, but it cannot solely resolve issues of ownership, scope of authorization, and revenue distribution.

For a content asset to be viable in the long term, it is necessary to explain who created the work, who holds the relevant rights, how the media acquired these rights, how long the rights can be used, and which products they are allowed to be used in.

If these questions cannot be answered, "content assets" will easily remain at the level of concepts and promotion.

Traditional media contracts often used general terms such as "publish," "reprint," and "disseminate." In the context of traditional newspapers and websites, these expressions were generally sufficient for everyday use; however, in the era of artificial intelligence, their boundaries have become blurred.

GFM can establish tiered licensing rules for original content, specifying which articles are allowed to be publicly searched and which content can only be included in paid research databases; which images are limited to news reporting and which can be included in commercial models; whether translation rights, audio rights, and model training rights need to be priced separately; and whether the other party still has the right to retain existing copies after the cooperation ends.

The clearer the rules, the easier it is for content to be properly authorized, and the lower the chance of future disputes.

News needs to be read, disseminated, and used to generate public value. Artificial intelligence will also open up new ways to use news content. What needs to be corrected is the old system that relies on implicit consent and vague authorization, and the reality that creators and media have long been excluded from revenue sharing.

After $1.5 billion

The Anthropic settlement will not end the copyright disputes surrounding artificial intelligence. US courts will continue to determine the scope of fair use in various cases, and new licensing models will gradually emerge among news organizations, publishers, authors, artists, and technology companies.

The significance of this case lies in the fact that it makes the entire industry clearly see just how tangible the cost of failing to manage data sources can be.

In the past, some tech companies believed that as long as the model's output didn't extensively copy the author's original text, issues with the training data wouldn't be a decisive risk. The Anthropic case illustrates that legal liability can arise even before the model generates content. The very act of downloading, copying, and permanently storing works is enough to generate substantial claims.

GFM is pushing for news to become a content asset that can be managed long-term, legally authorized, and with revenue fairly distributed.

These types of assets require high-quality content and significant dissemination impact, as well as a clear and traceable legal basis. Media outlets should know the origin of each piece of content, the extent to which it can be used, and how future revenue generated should be returned to those who contributed to its creation.

The journalism industry has traditionally viewed a report as a product whose primary mission is accomplished on the day of publication. Artificial intelligence is extending its lifespan and enabling it to enter many previously unseen business scenarios. A report can continue to be searched, translated, cited, adapted, and licensed years later, provided the media outlet can prove its right to manage the content.

The $1.5 billion paid by Anthropic doesn't answer the question of how much a book or a news article is actually worth. It shows the entire industry that ignoring content sources and licensing records will ultimately determine its value.

This price can be transformed into revenue shared by the author, media, and technology company through normal licensing; however, it may also become a judicial settlement, legal fees, and damage to corporate reputation after the rights relationship is out of control.

The system GFM aims to establish is one that ensures content value is realized along the aforementioned path as much as possible. Before artificial intelligence extensively utilizes human knowledge, the source, rights, and distribution methods of works should be clearly recorded. This protects creators and establishes a sustainable data order for the artificial intelligence industry.

Disclaimer:

This article is compiled based on publicly available court documents, settlement materials, and related news reports. Its purpose is to analyze the development trends of artificial intelligence training data, copyright liability, and content asset systems. It does not constitute legal advice, investment recommendations, or a final determination of the liability of any party involved. Specific rights, eligibility for compensation, scope of settlement, and legal effects involved in the case should be based on official court documents, settlement agreements, and relevant judicial procedures.