I’m always a bit skeptical about these sorts of things. Perhaps I’m just ignorant about the methods used.. but the amount of data we can get from the most distant known galaxy can’t be very much. How confident can we be that the shift in observed light or whatever is actually from the presence of Oxygen and not one of probably countless other causes, both known and unknown.
This is 100% a factor. The internet has some pretty dark and nasty corners; therefore so does the model. Seeing it unfiltered would be a PR nightmare for OpenAI.
Can someone help me understand why it's a problem for companies to train these huge LLM on your copyrighted material? What exactly is the harm that is being done to the copyright holder?
I can understand why the New York Times (for example) wants to claim that a couple billion dollar companies have done it actual harm; but I am struggling to actually identify what it is.
Whether or not this will reduce CO2 is yet to be seen. But it will almost certainly raise the price of beef and dairy products. Taxes on producers inevitably get passed on to consumers.
Isn’t it true that the only thing that LLM’s do is “hallucinate”?
The only way to know if it did “hallucinate” is to already know the correct answer. If you can make a system that knows when an answer is right or not, you no longer need the LLM!
I think we should less worried about people believing deep-fakes are real and more worried that politicians (and others) will be able to claim things which they actually said are deep fakes.
I think the important part of the technique is talking about what you see, not the actual act of seeing things. He talks about creating brain pathways between the visual and linguistic parts of your brain.
Sure, but "were they in the office" is /much/ easier to measure at scale. And I'm sure you'd throw a pretty big fit if your pay was docked because the company didn't think you did a good enough job, even though you were in the office.
And it would only notify someone for human review if a certain threshold was reached; just having one or two violating images would have tripped the system.
It seems to me that something really dangerous or illegal needs to be going on to morally justify publicly releasing a 100GB worth of (presumably unfiltered) private company data.