AI grave diggers

Spirit Airlines died and Google swallowed its carcass.

That’s the simplest summary of last week’s event when Google won the auction for data left behind by a bankrupted airline.What did Google get for $10 million? The trove included 100 million emails, 500 million Microsoft Teams chats, data about revenue, aircraft operations and more. What Google didn’t get? Passenger profiles and profiles from the loyalty program. It’s also worth mentioning that the data which Google receives will be scrubbed from any personal identifiable information to make you feel better.

It sounds comforting. Until you realize that many of the Spirit Airlines customers are using Gmail and finding matches between these two data sets can’t be simpler.

In the Copyright Wars post I mentioned that Anthropic was buying physical books in bulk, cutting the binding and feeding them to scanners to be digitized. Anthropic is not alone in this type of activity. Amazon is doing the same. Imagine a company, which started as the biggest bookstore on the planet, gets into the business of buying books from others and destroys them in the process. #chokingonirony

These two stories remind me of the post Oil and Data. There will be blood which I penned in April 2024. I reread the post with a certain level of satisfaction.

The headlines at that time called the data well dry sometime between 2026 (today) and 2030 (soon enough) and the report cited to support the story was based on the study Will we run out of data? Limits of LLM scaling based on human-generated data.

The abstract contained this: ‘Our findings indicate that if current LLM development trends continue, models will be trained on datasets roughly equal in size to the available stock of public human text data between 2026 and 2032, or slightly earlier if models are overtrained.

That’s as far as the journalist read.

At the end of the study, the researchers mentioned that: It’s likely that alternative sources of data will likely be adopted before then, allowing ML systems to continue scaling.

When in doubt, report the bad news only.

The perceived data scarcity was invoked based on the parallel with oil. The first prediction of peak oil is from 1880 and the next one from 1919. Somehow there is still an abundance of this non-renewable energy source.

One of the reasons for buying old books — old means anything before 2022, is that these books are not tainted by AI. The worst you can get is plagiarized content or complete falsification. Yes, these were the simple days. Just imagine that the library in your parents or grandparents house is no longer a source of family dispute over who has to take these books to the dump, but the fight over who will become a millionaire. Maybe not a millionaire, but next time you have Amazon delivery, they take your books away for free.

The point in all of this?

It is nonsense that we are running out of data to use for training. Just because we used one book for training AI, doesn’t mean that we are done with that book. We are still learning what the learning part should look like and we will come up with novel ways to learn and importantly how to forget. Remember that context is everything. You might digitize millions of books but without the context it is just a text.

And we are talking only about text. Another data point of interest. The Rubin Observatory generates 10 terabytes of data every 24 hours, same as if you watch Netflix for two years. Most of the data processing is focused on finding differences e.g. to identify asteroids hurling towards the Earth to reinstate dinosaurs back to their glory.

Compare that to the Large Hadron Collider (LHC) at CERN where scientists collected 1 exabyte of data from running all the experiments. In movie units, that’s almost 50,000 years of watching TV. It will take decades to process the data. Decades!

The recurrent pattern? We have more data than we know what to do with it. The limitation is not the volume. The limitation is not what we can train the AI with. The limitation is that we don’t even know what learning is and what we want to learn.

Next
Next

Anthropic chip. Why?