4 Comments
User's avatar
CM's avatar
Jul 4Edited

Typically informative post. One observation - the LLM community's bias towards considering their apparatus "intelligent" is embedded in the terminology they use, such as "training" models which "hallucinate". They aren't "training" anything - they are loading our collective creative and work product in order to sell it back to us. And the resulting systems don't "hallucinate" - they paste over discontinuities to make the output appear seamless rather than simply marking the holes.

These things aren't intelligent and we need to stop talking about them that way.

Donald Duncan's avatar

You omit one of the main sources of new data - social media. The NYTimes-funded study found that the second and fourth most cited sources were Facebook and Reddit. This certainly isn't going to produce high-quality output. In fact, I read that there's a subreddit (#poisonal?) that's dedicated to feeding LLMs who use these sources off-the-wall "facts". It succeeds at some level; DuckDuckGo's "Search Assist" recently reported that Donald Trump had died of rabies which he caught from J.D. Vance!