The Company That Wanted to Feed Your Books to AI

The Company That Wanted to Feed Your Books to AI

How millions of books became raw material in the race to build artificial intelligence


In early 2024, someone started buying books. Not rare first editions or the latest bestsellers, but ordinary used books, ordered in batches of tens of thousands from secondhand sellers such as Better World Books and World of Books.

What happened next was stranger. The books were taken to scanning facilities, where hydraulic cutters sliced off their spines. The loose pages were fed through high-speed scanners, and the paper was sent off for recycling. Nothing was kept except the words, now stored as digital files.

The buyer was Anthropic, the San Francisco company behind the Claude chatbot, a rival to ChatGPT. Internally, the operation had a name: Project Panama. A planning document from April 2024, later unsealed in court, put the goal plainly: the project was “our effort to destructively scan all the books in the world.”

The public only learned about it in January 2026, when The Washington Post reviewed more than 4,000 pages of documents released in a copyright lawsuit. They showed a company that had kept the project quiet, and one that understood how it might look.

The real question was never just what Anthropic did to the books. It was whether the company had the right to do it.

Why an AI company wants your books

Chatbots like Claude learn by reading. They are trained on vast amounts of text, and from it they pick up vocabulary, grammar, facts, arguments, and styles of writing. The more good text they see, the better they get.

Much of the internet is short, messy, and repetitive. Books are different. A history book can hold decades of research; a novel holds thousands of careful choices about language and character. According to the court documents, Anthropic’s executives wanted books precisely because they could teach Claude to write well, rather than imitate the low-quality chatter of the web.

To lead the effort, Anthropic hired Tom Turvey in February 2024. He had helped build Google Books, the project that scanned tens of millions of volumes two decades earlier. One vendor’s proposal described a job to convert between 500,000 and two million books in six months. Unlike Google, which scanned books without destroying them, Anthropic chose the faster and cheaper route of cutting them apart.

Buying a book and scanning it for your own use turned out to be the legally safer part of the story. The trouble was the books Anthropic had gathered before Project Panama began.

Years earlier, in 2021, co-founder Ben Mann had downloaded millions of books from so-called shadow libraries: websites such as Library Genesis (LibGen) and the Pirate Library Mirror that distribute copyrighted books without permission. In total, the court found, Anthropic had downloaded more than seven million pirated copies and kept them as a permanent library. Project Panama was, in part, a response to the legal risk that collection created.

The lawsuit and the $1.5 billion bill

In 2024, three authors sued: novelist Andrea Bartz and nonfiction writers Charles Graeber and Kirk Wallace Johnson. They argued that Anthropic had used their books, and millions of others, to build its AI without permission or payment.

At the heart of the case was an old question in a new form. A person can buy a book, read it and learn from it, and no one expects them to pay the author every time they remember something. But is it the same when the “reader” is a company’s computer system, trained on millions of books and sold as a product?

In June 2025, Judge William Alsup gave a split answer. Training an AI on books that had been bought legally, he ruled, was “fair use”: a legal exception that allows copyrighted work to be used without permission when the use transforms it into something new. Scanning purchased books to save shelf space was fine, too. But downloading pirated copies to build a permanent library was not protected, and that claim would go to a jury.

The ruling was a partial win for Anthropic, but the risk that remained was enormous. Alsup allowed the authors to sue on behalf of everyone whose books were in the pirate libraries. US law allows damages of up to $150,000 per work for willful infringement, and roughly 500,000 books were covered.

In August 2025, months before trial, Anthropic agreed to settle for $1.5 billion, about $3,000 per book, without admitting wrongdoing. After Alsup retired, Judge Araceli Martínez-Olguín granted final approval on July 20, 2026, calling it the largest copyright settlement in US history. More than 91% of eligible authors and publishers had filed claims.

The deal also requires Anthropic to destroy the pirated files and any copies made from them. The company has said those pirated datasets were not used to train the Claude models it released commercially. The scanned books from Project Panama, bought and paid for, are not affected.

What is a book worth to a machine?

The settlement closed one case, but it left the bigger question open. Because Anthropic settled, Alsup’s fair-use ruling will never be reviewed by an appeals court, so it binds no one else. Some authors and publishers opted out of the deal and have filed their own lawsuits, which are still pending. Writers, artists, photographers and newspapers have brought similar cases against other AI companies.

For publishers, the worry is economic. Producing a book means paying an author, an editor, a designer and a printer. An AI system never needs to sell a copy of that book to compete with it. It only needs to have learned what is inside, so it can summarize the argument or explain the history when someone asks.

That is why the line Alsup drew matters so much. He decided that how a company obtains a book counts, and that buying it is different from taking it. He did not decide whether authors deserve a share of the value their work creates once a machine has learned from it. That question belongs to future courts, lawmakers, and contracts.

When a person reads a hundred books and writes something new, we call it education. When a machine reads millions and becomes a product worth billions, we are still deciding what to call it.

For five centuries, the book was the product. Project Panama suggests a future in which it is something else: the raw material, cut from its spine, scanned and recycled, its words living on inside a machine.

0 comments

Leave a comment

Please note, comments need to be approved before they are published.