india employmentnews

Explainer: AI is reading thousands of books, but why are companies buying them and cutting them up?

 | 
vv

Recently, numerous booksellers across Australia and Europe reported bulk purchases of books—ranging from rare editions and vintage copies to inexpensive paperbacks. It later emerged that many of these books were being acquired to train AI models. In recent years, major AI companies—including the creators of ChatGPT, Gemini, and Claude—have trained their models on billions of words available on the internet. However, the landscape is shifting, and copyright-related lawsuits are on the rise. Consequently, AI companies are now seeking trustworthy, high-quality content written by humans, which is why they are turning their attention to books.

According to a report by *The Guardian*, Australian booksellers received requests from intermediaries for a wide range of titles, from rare editions to cheap paperbacks. Similarly, *Fortune* reported a comparable incident involving a Dutch bookseller. Many of these books are unique, with no other copies available.

Why books, specifically?
The more high-quality data AI Large Language Models (LLMs) ingest, the better and more accurate their responses to our queries become. By utilizing vast amounts of text data, LLMs are capable of understanding and predicting human language, as well as generating new text that mimics human writing styles.

Books often feature superior editing compared to most online text. Furthermore, books excel at providing in-depth coverage of subjects and exploring their various facets, offering profound insights across diverse topics. Crucially, the information and prose found in older books are often truly original—content that is scarce on the internet.

Experts suggest that books published prior to 2022 are particularly valuable for this purpose, as they contain genuinely original content. In essence, AI companies now view books as an excellent resource for learning authentic human writing styles. What is the primary reason for the opposition?
The core of this entire controversy lies not in the purchase of the books, but in the method used to scan them. A copyright infringement lawsuit filed against Anthropic revealed that "destructive scanning" was the primary trigger for the legal action.

Destructive scanning is a process in which books are destroyed after being scanned. The process begins by removing the book's binding and separating the pages. These pages are then fed individually into a high-speed scanner. Once scanned, the content is converted into digital text using Optical Character Recognition (OCR) technology. OCR technology identifies printed or handwritten characters on paper, photographs, or scanned documents and converts them into digital text that a computer can read, search, and even edit.

This text is subsequently used to train AI models. In many instances, the books are recycled or destroyed after the process is complete. Essentially, new books are used throughout this entire procedure.


Disclaimer: This content has been sourced and edited from NDTV India. While we have made modifications for clarity and presentation, the original content belongs to its respective authors and website. We do not claim ownership of the content.