The launch of ChatGPT polluted the world forever(theregister.com)
theregister.com
The launch of ChatGPT polluted the world forever
https://www.theregister.com/2025/06/15/ai_model_collapse_pollution/
7 comments
Someday maybe we’ll have a term similar to “low-background steel” for information and web content.
Large discussion earlier this week: https://news.ycombinator.com/item?id=44239481
The root of it is deterioration in trust. Even before LLMs hit the scene there was suspicion of narrative manipulation by social media sites. ChatGPT only changed how popular this take is, but not its measure.
This is a great analogy.
The article keeps making it sound as if it's a problem for humans. e.g.:
> Now here the date is more flexible, let's say 2022. But if you're collecting data before 2022 you're fairly confident that it has minimal, if any, contamination from generative AI. Everything before the date is 'safe, fine, clean,' everything after that is 'dirty.'"
Though what it seems to actually mean is that it's a problem for (future) generative AI (the "genAI collapse"). To which I say;
> Now here the date is more flexible, let's say 2022. But if you're collecting data before 2022 you're fairly confident that it has minimal, if any, contamination from generative AI. Everything before the date is 'safe, fine, clean,' everything after that is 'dirty.'"
Though what it seems to actually mean is that it's a problem for (future) generative AI (the "genAI collapse"). To which I say;
This seems like a very badly written article that rambles on in random directions. It proposes incredibly dumb ideas to anyone with half a brain like water marking AI output.
The most damning part for me is mentioning the Apple paper and the refute of the Apple paper, to my knowledge that paper had nothing to do with training on generated data. It was talking about reasoning models, but because they use the word “model collapse”, apparently, the author of this article decided to include it in, which just shows how they don’t know what they’re talking about (unless I’m completely misunderstanding the Apple paper).
The most damning part for me is mentioning the Apple paper and the refute of the Apple paper, to my knowledge that paper had nothing to do with training on generated data. It was talking about reasoning models, but because they use the word “model collapse”, apparently, the author of this article decided to include it in, which just shows how they don’t know what they’re talking about (unless I’m completely misunderstanding the Apple paper).
This! And I’d add, it’s the Register–it has always had a very low bar.
lowbackgroundsteel.ai sounds really promising. I don't really care for it as a clean AI training source, but I'm interested in a curated internet where I know it's not diluted with generative content. I'm not sure what that would look like when it comes to social media. This AI era has made me return to reading physical books as a hobby and engaging with offline/non-anonymous online communities more. Confidence in authenticity is one of the most important things for me these days.
Humanity now lives in a world where any text has most likely been influenced by AI, even if it’s by multiple degrees of separation.