Researchers at Oxford and Cambridge show that ChatGPT has now a 'big problem' and America's biggest investor Michael Burry agrees with them; says: Research on this is ...
Researchers at Oxford and Cambridge have identified a fundamental flaw threatening the future of large language models like ChatGPT, and \"Big Short\"
Researchers at Oxford and Cambridge have identified a fundamental flaw threatening the future of large language models like ChatGPT, and "Big Short" investor Michael Burry says the finding backs up a warning he's been making about the technology's core limitations for a while. A new study from the University of Oxford and University of Cambridge has revealed a major threat to large language models like ChatGPT, calling it “model collapse.” The research shows that when generative AI systems are repeatedly trained on synthetic data produced by earlier models, they begin to lose the “tails” of the original data distribution. In practice, this means models forget the creative, fringe, and unique nuances of human writing, collapsing into repetitive echo chambers.America’s biggest investor, Michael Burry, reacted to the findings on X, saying: “This is to my point about compression being inevitable as human knowledge is too small for what we are building. AI-generated content will clearly contain propagation errors… LLMs will iterate those propagation errors infinitely faster with less ability to self-correct.”The problem: models trained on their own outputBurry drew attention to the research by sharing a post outlining what scientists call "model collapse." The concern centers on a shift already underway across the internet: as generative AI produces more and more of the text online, future AI models risk being trained on data that was itself generated by earlier AI systems, rather than by humans.According to the research, feeding AI-generated content back into training indiscriminately causes what researchers describe as irreversible defects in the resulting model. Specifically, the model begins losing the "tails" of the original data distribution, meaning it gradually forgets the creative, unusual, and distinctly human nuances present in real human writing. Over successive generations, this causes the model to collapse into an increasingly narrow, repetitive echo chamber of its own outputs.Crucially, the researchers found this isn't unique to ChatGPT or language models specifically. Their theoretical framework showed the collapse phenomenon is ubiquitous across learned generative models more broadly, appearing in large language models as well as in variational autoencoders and Gaussian mixture models.The stakes are significant for the entire AI industry, given how heavily tech companies rely on scraping large volumes of internet data to train increasingly capable models. The research warns that if companies want to sustain the benefits of training on web data going forward, model collapse needs to be taken seriously as a genuine constraint rather than a theoretical curiosity. The clear takeaway, as the research frames it: genuine, human-generated data is going to become an increasingly scarce and valuable resource in a web that's rapidly filling up with AI-generated content.Burry's own warning, echoed and expandedWriting on X, Burry tied the research directly to an argument he's made before: that compression of available information is essentially inevitable, because human knowledge itself is too limited relative to the scale of what AI companies are trying to build, and because human wants, needs, and questions are inherently repetitive and redundant.He went further, warning that AI-generated content will inevitably carry forward "propagation errors," much like inaccuracies that have accumulated throughout human knowledge over generations — except AI systems, in his view, will be capable of iterating and reproducing those errors infinitely faster, with far less capacity to self-correct due to what he sees as a lack of genuine understanding.A deeper claim: LLMs may never reach AGIBurry used the moment to make a broader, more philosophical argument about artificial general intelligence. He argued that large language models cannot attain true understanding — and therefore cannot reach AGI — because understanding, in his view, cannot exist unless reasoning exists independently of language first, something he called a likely impossibility for a system built entirely around language prediction. He acknowledged that researchers are already exploring ways to work around this limitation, while noting that many in the field haven't accepted his underlying premise in the first place.Setting up a clash with Jensen HuangBurry's skepticism lands in direct contrast to comments made earlier this month by Nvidia CEO Jensen Huang, who declared that AGI has already arrived, citing OpenAI's GPT-6 Astra model as proof. Huang's claim drew immediate pushback from parts of the AI research community.Stanford's Institute for Human-Centered AI defines AGI as AI capable of learning, reasoning, and applying knowledge across a broad range of tasks at or beyond human-level performance — critically, with the ability to adapt to unfamiliar situations, unlike narrow AI built for specific tasks. The concept remains genuinely unsettled within the field, with no universally accepted definition or standardized test for confirming when AGI has actually been reached, alongside ongoing debates over its safety and ethical implications.You use AI every day. Now get your AI Quotient. Take the AIQ test.
Topics in this story
Gathered from external sources. Rights to this text belong to whoever originally published it.