Bing AI can't be trusted(dkb.blog)
dkb.blog
Bing AI can't be trusted
https://dkb.blog/p/bing-ai-cant-be-trusted
601 comments
I have come to two conclusions about the GPT technologies after some weeks to chew on this:
1. We are so amazed by its ability to babble in a confident manner that we are asking it to do things that it should not be asked to do. GPT is basically the language portion of your brain. The language portion of your brain does not do logic. It does not do analyses. But if you built something very like it and asked it to try, it might give it a good go.
In its current state, you really shouldn't rely on it for anything. But people will, and as the complement of the Wile E. Coyote effect, I think we're going to see a lot of people not realize they've run off the cliff, crashed into several rocks on the way down, and have burst into flames, until after they do it several dozen times. Only then will they look back to realize what a cockup they've made depending on these GPT-line AIs.
To put it in code assistant terms, I expect people to be increasingly amazed at how well they seem to be coding, until you put the results together at scale and realize that while it kinda, sorta works, it is a new type of never-before-seen crap code that nobody can or will be able to debug short of throwing it away and starting over.
This is not because GPT is broken. It is because what it is is not correctly related to what we are asking it to do.
2. My second conclusion is that this hype train is going to crash and sour people quite badly on "AI", because of the pervasive belief I have seen even here on HN that this GPT line of AIs is AI. Many people believe that this is the beginning and the end of AI, that anything true of interacting with GPT is true of AIs in general, etc.
So people are going to be even more blindsided when someone develops an AI that uses GPT as its language comprehension component, but does this higher level stuff that we actually want sitting on top of it. Because in my opinion, it's pretty clear that GPT is producing an amazing level of comprehension of what a series of words means. The problem is, that's all it is really doing. This accomplishment should not be understated. It just happen to be the fact that we're basically abusing it in its current form.
What it's going to do as a part of an AI, rather than the whole thing, is going to be amazing. This is certainly one of the hard problems of building a "real AI" that is, at least to a first approximation, solved. Holy crap, what times we live in.
But we do not have this AI yet, even though we think we do.
1. We are so amazed by its ability to babble in a confident manner that we are asking it to do things that it should not be asked to do. GPT is basically the language portion of your brain. The language portion of your brain does not do logic. It does not do analyses. But if you built something very like it and asked it to try, it might give it a good go.
In its current state, you really shouldn't rely on it for anything. But people will, and as the complement of the Wile E. Coyote effect, I think we're going to see a lot of people not realize they've run off the cliff, crashed into several rocks on the way down, and have burst into flames, until after they do it several dozen times. Only then will they look back to realize what a cockup they've made depending on these GPT-line AIs.
To put it in code assistant terms, I expect people to be increasingly amazed at how well they seem to be coding, until you put the results together at scale and realize that while it kinda, sorta works, it is a new type of never-before-seen crap code that nobody can or will be able to debug short of throwing it away and starting over.
This is not because GPT is broken. It is because what it is is not correctly related to what we are asking it to do.
2. My second conclusion is that this hype train is going to crash and sour people quite badly on "AI", because of the pervasive belief I have seen even here on HN that this GPT line of AIs is AI. Many people believe that this is the beginning and the end of AI, that anything true of interacting with GPT is true of AIs in general, etc.
So people are going to be even more blindsided when someone develops an AI that uses GPT as its language comprehension component, but does this higher level stuff that we actually want sitting on top of it. Because in my opinion, it's pretty clear that GPT is producing an amazing level of comprehension of what a series of words means. The problem is, that's all it is really doing. This accomplishment should not be understated. It just happen to be the fact that we're basically abusing it in its current form.
What it's going to do as a part of an AI, rather than the whole thing, is going to be amazing. This is certainly one of the hard problems of building a "real AI" that is, at least to a first approximation, solved. Holy crap, what times we live in.
But we do not have this AI yet, even though we think we do.
I've posted this into another thread as well, from Sam Altman, CEO of OpenAI, two months ago, on his Twitter feed:
"ChatGPT is incredibly limited, but good enough at some things to create a misleading impression of greatness. it's a mistake to be relying on it for anything important right now. [...] fun creative inspiration; great! reliance for factual queries; not such a good idea." (Sam Altman)
"ChatGPT is incredibly limited, but good enough at some things to create a misleading impression of greatness. it's a mistake to be relying on it for anything important right now. [...] fun creative inspiration; great! reliance for factual queries; not such a good idea." (Sam Altman)
The amount of trust people are willing to place in AI is far more terrifying than the capabilities of these AI systems. People are too willing to give up their responsibility of critical thought to some kind of omnipotent messiah figure.
Our exposure to smart-sounding chatbots is inducing a novel form of pareidolia: https://en.wikipedia.org/wiki/Pareidolia .
Our brains are pattern-recognition engines and humans are social animals; together that means that our brains are predisposed to anthropomorphizing and interpreting patterns as human-like.
For the whole of human history thus far, the only things that we have commonly encountered that conversed like humans have been other humans. This means that when we observe something like ChatGPT that appears to "speak", we are susceptible to interpreting intelligence where there is none, in the same way that an optical illusion can fool your brain into perceiving something that is not happening.
That's not to say that humans are somehow special or that or human intelligence is impossible to replicate. But these things right here aren't intelligent, y'all. That said, can they be useful? Certainly. Tools don't need to be intelligent to be useful. A chainsaw isn't intelligent, and it can still be highly useful... and highly destructive, if used in the wrong way.
Our brains are pattern-recognition engines and humans are social animals; together that means that our brains are predisposed to anthropomorphizing and interpreting patterns as human-like.
For the whole of human history thus far, the only things that we have commonly encountered that conversed like humans have been other humans. This means that when we observe something like ChatGPT that appears to "speak", we are susceptible to interpreting intelligence where there is none, in the same way that an optical illusion can fool your brain into perceiving something that is not happening.
That's not to say that humans are somehow special or that or human intelligence is impossible to replicate. But these things right here aren't intelligent, y'all. That said, can they be useful? Certainly. Tools don't need to be intelligent to be useful. A chainsaw isn't intelligent, and it can still be highly useful... and highly destructive, if used in the wrong way.
For me the fundamental issue at the moment for ChatGPT and others is the tone it replies in. A large proportion of the information in language is in the tone, so someone might say something like "I'm pretty sure that the highest mountain in Africa is Mount Kenya" whereas ChatGPT instead says "the highest mountain in Africa is Mount Kenya", and it's the "is" in the sentence that's the issue. So many issues in language revolve around "is" - the certainty is very problematic. It reminds me of a tutor at art college who said too many people were producing "thing that look like art". ChatGPT produces sentence that look like language, and because of "is" they read as quite compelling due to the certainty it conveys. Modify that so it says "I think..." or "I'm pretty sure..." or "I reckon..." and the sentence would be much more honest, but the glamour around it collapses.
I had this idea the other day concerning the 'AI obfuscation' of knowledge. The discussion was about how AI image generators are designed to empower everyone to contribute to the design process. But I argued that you can only reasonably contribute to the process if you can actually articulate the reasoning beyond your contributions. If an AI made it for you, you probably can't, because the reasoning is simply "this is the amalgamation of training data that the AI spat out." But, there's a realistic version of reality where this becomes the norm and we increasingly rely on AI to solve for issues that we don't understand ourselves.
And, perhaps more worrying, the more widely adopted AI becomes, the harder it becomes to correct its mistakes. Right now millions of people are being fed information they don't understand, and information that's almost entirely incorrect or inaccurate. What is the long term damage from that?
We've obfuscated the source data and essentially the entire process of learning with LLMs / AIs, and the path this leads down seems pretty obviously a net negative for society (outside of short term profit for the stake holders).
And, perhaps more worrying, the more widely adopted AI becomes, the harder it becomes to correct its mistakes. Right now millions of people are being fed information they don't understand, and information that's almost entirely incorrect or inaccurate. What is the long term damage from that?
We've obfuscated the source data and essentially the entire process of learning with LLMs / AIs, and the path this leads down seems pretty obviously a net negative for society (outside of short term profit for the stake holders).
Was it what, just a week ago I was being called dumb for suggesting there'd be accuracy issues with this? I mean Bing had like a whole three weeks to slap this together after OpenAI first demoed it's ability to make things up.
oh only six days ago:
https://news.ycombinator.com/item?id=34699087
> This is a commonly echoed complaint but it’s largely without merit. ChatGPT spews nonsense because it has no access to information outside of its training set.
> In the context of a search engine, single shot learning with the top search results should mitigate almost all hallucination.
hows that going?
oh only six days ago:
https://news.ycombinator.com/item?id=34699087
> This is a commonly echoed complaint but it’s largely without merit. ChatGPT spews nonsense because it has no access to information outside of its training set.
> In the context of a search engine, single shot learning with the top search results should mitigate almost all hallucination.
hows that going?
There's also the instance of the Bing chatbot insisting that the current year is 2022 and being EXTREMELY passive-aggressive when corrected.
https://libreddit.strongthany.cc/r/bing/comments/110eagl/the...
https://libreddit.strongthany.cc/r/bing/comments/110eagl/the...
I mean, it's in beta and it's not really intelligent despite the cavalier use of the term AI these days
It's just a collage of random text that sorta resembles what someone would say, but it has no commitment to being truthful because it has no actual appreciation for what information it is relaying, parroting or conveying.
But yeah, I agree Google got way more hate for their failed demo than MS... I don't even understand why. Satya Nadella's did a great job conveying the excitement and general bravado on his interview on CBS News[1] but the accompanying demo was littered with mistakes. The reporter called it out, yet coverage on the press has been very one-sided against Google for some reason. First mover advantage, I suppose?
----------
1. https://www.cbsnews.com/news/microsoft-ceo-satya-nadella-new...
It's just a collage of random text that sorta resembles what someone would say, but it has no commitment to being truthful because it has no actual appreciation for what information it is relaying, parroting or conveying.
But yeah, I agree Google got way more hate for their failed demo than MS... I don't even understand why. Satya Nadella's did a great job conveying the excitement and general bravado on his interview on CBS News[1] but the accompanying demo was littered with mistakes. The reporter called it out, yet coverage on the press has been very one-sided against Google for some reason. First mover advantage, I suppose?
----------
1. https://www.cbsnews.com/news/microsoft-ceo-satya-nadella-new...
The potential for being sued for libel is huge. It's one thing to say the height of Everest wrong, another to falsely claim that a vacuum has a short cord, or that a company had 5.9% operating margin instead of 4.6%.
Bing AI gets a pass because it's disruptive. Google doesn't because it is the incumbent. Mystery solved.
I may be an unusual audience but something I've appreciated about these models is their ability to create unusual synthesis from seemingly unrelated sources. It's like if a scientist read up on many unrelated fields, got super high and started thinking of the connections between these fields.
Much of what they would produce might just be hallucinations, but they are sort of hallucinations informed by something that's possible. At least in my case, I would much rather then parse through that and throw out the bullshit, but keep the gems.
Obviously that's a very different use case than asking this thing the score of yesterday's football game.
Much of what they would produce might just be hallucinations, but they are sort of hallucinations informed by something that's possible. At least in my case, I would much rather then parse through that and throw out the bullshit, but keep the gems.
Obviously that's a very different use case than asking this thing the score of yesterday's football game.
Likely going to be a wave of research/innovation "regularizing" LLM output to conform to some semblance of reality or at least existing knowledge (e.g. knowledge graph). Interesting to see how this can be done quickly enough...
I think this is a weird non-issue and it's interesting people are so concerned about it.
- Human curated systems make mistakes.
- Fiction has created the trope of the omniscient AI.
- GPT curated systems also make mistakes.
- People are measuring GPT against the omniscient AI mythology rather than the human systems it could feasibly replace.
- We shouldn't ask "is AI ever wrong" we should ask "is AI wrong more often than the human-curated information? (There are levels of this - min wage truth is less accurate that senior engineer truth.)
- Even if the answer is that AI gets more wrong, surely a system where AI and humans are working together to determine the truth can outperform a system that is only curated by either alone. (for the next decade or so, at least)
- Human curated systems make mistakes.
- Fiction has created the trope of the omniscient AI.
- GPT curated systems also make mistakes.
- People are measuring GPT against the omniscient AI mythology rather than the human systems it could feasibly replace.
- We shouldn't ask "is AI ever wrong" we should ask "is AI wrong more often than the human-curated information? (There are levels of this - min wage truth is less accurate that senior engineer truth.)
- Even if the answer is that AI gets more wrong, surely a system where AI and humans are working together to determine the truth can outperform a system that is only curated by either alone. (for the next decade or so, at least)
Out of curiosity, I searched the pet vacuum mentioned in the first example, and found it on amazon [0]. Just like Bing says, it is a corded model with a 16 feet cord, and searching the reviews for "noise" shows that many people think that it is too loud. At least in this case, it seems that Bing got it right.
[0]: https://www.amazon.com/Bissell-Eraser-Handheld-Vacuum-Corded...
[0]: https://www.amazon.com/Bissell-Eraser-Handheld-Vacuum-Corded...
To follow up on the author's example Bing search doesn't even know when the new Avatar is film is actually out (DECEMBER 17 2021?)
https://www.bing.com/search?q=when+is+the+new+avatar+film+ou...
Bing AI doesn't stand a chance.
https://www.bing.com/search?q=when+is+the+new+avatar+film+ou...
Bing AI doesn't stand a chance.
There is no point in hyping about a 'better search engine' when this continues to hallucinate incorrect and inaccurate results. It is now reduced to a 'intelligent sophist' instead of a search engine. Once many realise that it also frequently hallucinates nonsense, it is essentially no better than Google Bard.
After looking at the limitations of ChatGPT and Bing AI it is now clear that they aren't reliable enough to even begin to challenge search engines or even cite their sources properly. LLMs are just limited to bullshit generators which is what this current AI hype is all about.
Until all of these AI models are open-sourced and transparent enough to be trustworthy or if a competitor does it instead, then there is nothing revolutionary about this AI hype other than a AI SaaS using a creative Clubhouse-like waitlist mania.
After looking at the limitations of ChatGPT and Bing AI it is now clear that they aren't reliable enough to even begin to challenge search engines or even cite their sources properly. LLMs are just limited to bullshit generators which is what this current AI hype is all about.
Until all of these AI models are open-sourced and transparent enough to be trustworthy or if a competitor does it instead, then there is nothing revolutionary about this AI hype other than a AI SaaS using a creative Clubhouse-like waitlist mania.
I already don't trust virtually any search results except grep/rg.
> Bing AI can't be trusted
Of course it can't. No LLM can. They're bullshit generators. Some people have been saying it from the start, and now everyone is saying it.
It's a mystery why Microsoft is going full speed ahead with this. A possible explanation is that they do this to annoy / terrify Google.
But the big mystery is, why is Google falling for it? That's inexplicable, and inexcusable.
Of course it can't. No LLM can. They're bullshit generators. Some people have been saying it from the start, and now everyone is saying it.
It's a mystery why Microsoft is going full speed ahead with this. A possible explanation is that they do this to annoy / terrify Google.
But the big mystery is, why is Google falling for it? That's inexplicable, and inexcusable.
I don't know if it's started to use AI for regular search queries, but I noticed within the past week or two that Bing results got much worse. It seems it doesn't even respect quoting anymore, and the second and subsequent pages of results are almost entirely duplicates of the first. I normally use Bing when Google fails to yield results or decides to hellban me for searching too specifically, and for the past few years it was acceptable or even occasionally better, but now it's much worse. If that's the result of AI, then do not want!!!
I have been trying to help folks understand what the underlying mechanisms of these generative LLM's are so it's not such a surprise when we get wrong answers from them by putting together some youtube videos on the topic.
* [On the question of replacing Engineers](https://www.youtube.com/watch?v=GMmIol4mnLo)
* [On AI Plagiarism](https://www.youtube.com/watch?v=whbNCSZb3c8)
The consensus seems to be building now on HackerNews that there is a huge over-hype. Hopefully these two videos help see some of the nuance behind why it's an over-hype.
That being said, being that language generation is probabilistic, a given language model which is transformer based can either be trained or fine-tuned to have fewer errors in a particular domain - so this is all far from settled.
Long-term, I think we're going to see something closer to human intelligence from CNN's and other forms of neural networks than from transformers, which are really a poor man's NN. As hardware advances and NN's inevitably become cheaper to run, we will continue to see scarier and scarier A.I. -- I'm talking over a 10-20 year timeframe.
* [On the question of replacing Engineers](https://www.youtube.com/watch?v=GMmIol4mnLo)
* [On AI Plagiarism](https://www.youtube.com/watch?v=whbNCSZb3c8)
The consensus seems to be building now on HackerNews that there is a huge over-hype. Hopefully these two videos help see some of the nuance behind why it's an over-hype.
That being said, being that language generation is probabilistic, a given language model which is transformer based can either be trained or fine-tuned to have fewer errors in a particular domain - so this is all far from settled.
Long-term, I think we're going to see something closer to human intelligence from CNN's and other forms of neural networks than from transformers, which are really a poor man's NN. As hardware advances and NN's inevitably become cheaper to run, we will continue to see scarier and scarier A.I. -- I'm talking over a 10-20 year timeframe.
I frequently use ChatGPT to research various topics. I've noticed that eight out of 10 times I ask it to recommend some books about a topic it recommends non-existing books. There's no way I'd trust a search engine built on it.
I think ChatGPT and their lookalikes spell the end of the public internet as we know it. People now have tools to generate pages as they seem fit. Google will not be able to determine what are high quality pages if everything looks the same and is generated by AI bots. Users will be unable to find trustworthy results, and many of these results will be filled with generated garbage that looks great but is ultimately false.
What would be nice is for Microsoft to get hit by a barrage of lawsuits, MS to be ridiculed in the press and punished on Wall Street, and vindication of Google's more responsible introduction of AI methods over the years.
There will still be startups doing reckless things, but large, established companies that can immediately have bigger impact also have a lot more to lose.
There will still be startups doing reckless things, but large, established companies that can immediately have bigger impact also have a lot more to lose.
[deleted]
I wonder how much the upspeak way of typing affects this. People (even the author) often end declarations with question marks. Does this have any influence on the way the LLM parses the prompt?
AI can't be trusted in general, at least not for a long time. It gets basic facts wrong, constantly. The fear is that it will start eating its own dogfood and being more and more wrong since we are putting it in the hands of people that don't know any better and are going to use it to generate tons of online content that will later be used in the models.
It does make some queries much easier to find, for instance I had trouble finding out if the runner ups got the win in the Tour De France after the Armstrong doping scandal and it answered it instantly. The problem is that is offers answers with confidence, I think them adding citation is an improvement over ChatGPT, but it needs more.
Luckily, it's still a beta product and not in the hands of everyone. Unfortunately, ChatGPT is, which I find more problematic.
It does make some queries much easier to find, for instance I had trouble finding out if the runner ups got the win in the Tour De France after the Armstrong doping scandal and it answered it instantly. The problem is that is offers answers with confidence, I think them adding citation is an improvement over ChatGPT, but it needs more.
Luckily, it's still a beta product and not in the hands of everyone. Unfortunately, ChatGPT is, which I find more problematic.
What the hype machine still doesn't understand is that it's a language model, not a knowledge model.
It is optimized to generate information that looks as much like language as possible, not knowledge. It may sometimes regurgitate knowledge if it is simple or well trodden enough knowledge, or if language trivially models that knowledge.
But if that knowledge gets more complex and experiential, it will just generate words without attachment to meaning or truth, because fundamentally it only knows how to generate language, and it doesn't know how to say "I don't know that" or "I don't understand that".
It is optimized to generate information that looks as much like language as possible, not knowledge. It may sometimes regurgitate knowledge if it is simple or well trodden enough knowledge, or if language trivially models that knowledge.
But if that knowledge gets more complex and experiential, it will just generate words without attachment to meaning or truth, because fundamentally it only knows how to generate language, and it doesn't know how to say "I don't know that" or "I don't understand that".
Someone posted on Twitter that chatGPT is like economists - occasionally right but super confident that they are always right
[0]: https://files.catbox.moe/xoagy9.png