AI Giants Are Shredding Millions of Books for Training Data
Artificial intelligence companies are buying physical books by the millions, destroying them, and feeding the scanned pages into their training datasets.
coinbeat.newsTech companies are quietly purchasing physical books in massive quantities to feed hungry language models. Workers slice the spines off these books and scan every single page into digital datasets. Once the scanning process finishes, the original physical copies go straight into the recycling bin.
This aggressive hunt for training material highlights the growing shortage of high quality human text on the internet. As artificial intelligence models demand more data to improve their reasoning skills, developers are turning to the physical world to find fresh material. This trend raises major questions about copyright law and how authors will be compensated when their printed works are fed into machine learning systems.
Traders and investors should watch how legal battles unfold around copyright infringement in the tech sector. If courts rule against these scraping practices, artificial intelligence companies might face major hurdles and higher costs to acquire legal training data. Keep an eye on upcoming policy updates and lawsuits involving data privacy and intellectual property rights.
Market sentiment
Be the first to react
▍Comments (0)
No comments yet. Start the conversation!


