AI Companies Are Buying and Destroying Thousands of Books to Train Their Models

 


For many artificial intelligence companies, books are no longer viewed only as sources of knowledge or creative works. They have also become valuable raw materials for training large AI models.

Some companies have reportedly used massive digital book collections, including copies allegedly obtained from pirated databases. Copyright disputes involving companies such as Meta and Anthropic have raised serious questions about whether authors and publishers gave permission for their works to be included in AI training datasets.

Anthropic later agreed to pay authors approximately $1.5 billion to settle a major copyright lawsuit involving the use of books for AI development.

But when digital versions of certain titles were unavailable, Anthropic reportedly turned to physical secondhand books.

According to court documents, the company purchased large quantities of used books from resellers. It then used hydraulic cutting machines to remove the bindings, scanned the loose pages using industrial equipment, and converted the text into digital data for AI training.

The physical books were reportedly destroyed during the process.

Anthropic argued that it legally purchased the books and therefore had the right to alter or destroy them under the first-sale doctrine. This legal principle generally allows the owner of a lawfully purchased item to resell, modify, or dispose of it.

A judge also ruled that converting legally acquired books into digital training data was transformative and could qualify as fair use.

The growing demand for physical books has also changed the role of ISBNdb, a database originally created to help libraries, bookstores, and distributors locate and sell titles.

According to 404 Media, ISBNdb is now reportedly assisting AI companies in sourcing anywhere from 1,000 to as many as one million books in a single purchasing operation.

The practice highlights a growing conflict at the center of artificial intelligence development: AI companies need enormous amounts of human-created content, but authors, publishers, and copyright holders continue to question how that content is obtained, processed, and used.