AI labs are getting desperate

AI companies are running out of internet.

Because the internet is now chock-full of AI-made content and because AI models that train on AI slop get the AI version of mad cow disease, AI labs are turning to books.

They’re buying old books from shops around the globe, carving them up and digitizing them to extract the fresh, non-AI data they need to train their new models.

The move reeks of desperation and the numbers prove it.

●The total number of books published in human history is around 170 million.

●The average book yields around 80,000 text tokens for model training.

●Do the math and the sum total of tokens contained in all the books ever published is around 14 trillion.

Now here’s the kicker: the tokens needed for a single frontier-model training run is 15 trillion to 20 trillion.

Yup, one training run requires more text tokens than are contained in all books ever.

Which means the AI labs’ new book-scanning campaign is nowhere near a long-term solution to the problem of model collapse, i.e. the mad cow disease mentioned at the top.

Because LLMs need way more data than the history of print can offer, AI labs will have to find another way out of collapse (and fast).

Is there one? Well, yeah. See if you can guess what it is. (Hint: it’s humans.)

Previous
Previous

Your AI-generated content is worthless

Next
Next

AI is a tool, yes, but don’t use it as a shovel